Self-hosted document intelligence

Understand your documents. On your hardware.

OCR, auto-labeling, search and chat over your documents, on your own hardware.

  • Open source (AGPL)
  • HTTP API for integrations
  • Docker Compose
  • PostgreSQL, Valkey, S3-compatible

Product

One workspace for ingest, structure, and answers

Crisp UI shots from the real app: library, fields, labels, folders, and document chat on infrastructure you operate.

Library and upload

Upload PDFs and scans, filter by status and metadata, and track extraction progress in one place.

  • Central upload and connector import
  • Filter by status, labels, and metadata
  • Per-document extraction progress
Docuvate library document list

Recognized fields

Amounts, dates, senders, and custom fields are suggested automatically. You confirm suggestions with one click.

  • Suggestions for amount, date, and sender
  • Custom fields with one-click confirm
  • Same data for search, chat, and API
Document detail with extracted fields

Labels and folders

Labels group topics; folders mirror how you file. Recommendations assist without forcing automation.

  • Color labels for topics and projects
  • Folder tree that matches how you file
  • Recommendations you accept or ignore
Labels and folder explorer

Document chat

Ask about a document with a locally connected chat model (e.g. Ollama). Answers use the recognized text in your file.

  • Chat per document on extracted text
  • Locally connected model (e.g. Ollama)
  • No sending files to third-party clouds
Chat panel on document view

How it works

1. Ingest

Upload manually or pull from a connector. Docuvate stores the original in object storage.

2. Extract

The worker pulls text and fields. Progress and results appear in the UI.

3. Use

Find documents again with labels, folders, and full-text search. Chat and the API use the same underlying data.

API-first for your stack

Connect scripts, portals, and back-office tools with service keys. Same documents and labels as the web UI, without screen scraping.

Node.js
import { DocuvateClient } from '@docuvate/sdk';

const client = new DocuvateClient({
  baseUrl: 'https://your-host',
  apiKey: process.env.DOCUVATE_API_KEY,
});

const { data } = await client.api.listDocuments();
console.log(data?.items?.[0]?.title);

Integrations

Connections for import and automation, from cloud storage to services on your network.

Amazon S3

Available

Import objects from a bucket into your library.

Paperless-ngx

Available

Bridge an existing DMS archive.

Home Assistant

Available

Trigger automations from document events.

Gmail

Setup required

Mail attachments as a source; OAuth with your client credentials.

Microsoft Outlook

Setup required

Mail attachments as a source. One-time setup in settings.

Self-host on infrastructure you control

No per-seat cloud bill. Run Docuvate on your server, VM, or NAS. You choose when to upgrade, with optional fully local AI. License: AGPL. You run the stack; you own the data.

FAQ

Who is Docuvate for?

Tax advisors, small businesses, and households that want PDFs and scans organized without mandatory cloud storage.

Do I need a GPU?

Not to start. OCR and chat paths are CPU-friendly; heavier models are optional and documented.

How do I connect Docuvate to other systems?

Sign in through the web app for day-to-day use. For scripts and third-party tools, create a service key and use the HTTP API described in the API reference.

How do updates work?

You deploy and upgrade on your schedule. Breaking changes are versioned; the exported OpenAPI document describes the contract.