DocSense
Document intelligence that turns invoices and contracts into structured data.
- Client
- Finance operations team (NDA)
- Role
- AI engineer
- Year
- 2024
- Industry
- Finance operations
Problem
Every document was typed in by hand
Invoices, credit notes and contracts arrived by email as PDFs and scans, in whatever layout each vendor used. Analysts keyed them into the ERP line by line, and mistakes only surfaced at month-end reconciliation.
- Fwd: invoice attached (scan).pdfAccounts inbox, via email
- IMG_4471.jpg: photo of a delivery noteVendor, via email
- Month-end: totals don't match the ERPAnalyst, via spreadsheet
- Which PO was this invoice for?Finance, via chat
Research
Starting from the documents, not the model
Discovery began with a redacted sample of real documents and the ERP fields they had to fill. That showed which fields were easy, which were ambiguous, and where a person would always need to decide.
What we learned
- A small set of fields caused most corrections, so confidence had to be per field, not per document.
- Reviewers trusted extraction far more when they could see the exact region a value came from.
- Some checks (the PO exists, the vendor is active, totals add up) are deterministic and shouldn't be left to a model.
- Classification had to come first, because the right schema depends on the document type.
User journey
From inbox to ERP without retyping
A composite persona from the operations team, following one supplier invoice.
BeforeWith the product
- 1
Documents arrive
Before: Downloaded attachments from a shared inbox one by one.
Now: Emails are ingested automatically and each attachment becomes a tracked document.
- 2
Identify the document
Before: Opened each file to see what it was and who sent it.
Now: Documents are classified by type and vendor before any extraction runs.
- 3
Capture the data
Before: Typed header fields and line items into the ERP.
Now: Fields are extracted with a confidence level for each one.
- 4
Check the values
Before: Cross-checked totals and PO numbers by eye.
Now: Only low-confidence or rule-failing fields are highlighted for review.
- 5
Post to the ERP
Before: Keyed the final values and filed the PDF in a shared drive.
Now: Approved records export through the ERP API with a link to the source document.
Solution
A pipeline that knows when to ask a person
Each document is classified, read and validated in stages. Fields come back with confidence levels, business rules check them against ERP data, and only what's uncertain reaches a reviewer, with the source region highlighted.
How the work flows
Switch between the old process and the one the product runs.
- 1
Inbox ingestion
Attachments are picked up and stored with their email.
AutomatedOwner: Ingestion API - 2
OCR and classification
Text with layout, then document type and vendor.
AutomatedOwner: Pipeline - 3
Extraction with confidence
Typed fields, each with a confidence level.
AutomatedOwner: Pipeline - 4
Rule validation
PO exists, vendor is active, totals add up.
AutomatedOwner: Pipeline - 5
Review flagged fields
Only uncertain fields, with the source highlighted.
Person, in productOwner: Analyst - 6
Export to the ERP
Approved records post through the ERP API.
AutomatedOwner: Ingestion API
Architecture
A staged pipeline with a paper trail
A FastAPI service ingests documents and exposes the review queue. Celery workers run each stage (OCR, classification, extraction, validation) as a separate task, so a failure retries one step rather than the whole document. Every stage writes its output to PostgreSQL, which is what makes every exported value traceable.
- Interface
- Service
- AI
- Data and queues
- External system
Intake and review
API and storage
Pipeline stages
Destination
Extract fields
Extracts typed fields into a type-specific schema, each with a confidence level.
Why it's built this way
A value is only accepted if it can be found in the OCR text, so the model can't invent one.
Connections
UI
Review, don't retype
The review screen is where the product earns trust. The document and the extracted fields sit side by side, uncertain fields are highlighted in both, and every flag explains itself.
Development
Evaluated like a product, not a demo
Extraction quality can quietly regress with any prompt or model change, so the pipeline was built around an evaluation set from day one.
Stack
- API
- PythonFastAPI
- Pipeline
- CeleryOCRLLM
- Data
- PostgreSQLS3
- Review app
- React
- Integrations
- ERP REST APIEmail inbox
AI
OCR plus an LLM, with a person where it counts
Layout-aware OCR turns each page into text with positions. A classifier picks the document type and vendor, and an LLM extracts fields into a type-specific schema. Confidence comes from the model and from checks against the text itself, and validation rules run on top.
Invoice extraction
Globex Supplies GmbH
Invoice
- Invoice no. INV-20417
- Invoice date 12/09/2026
- PO reference PO-77120
- Machined brackets, batch delivery
- Net amount 8,420.00 EUR
- VAT 1,599.80 EUR
- Total due 10,019.80 EUR
Extracted fields
- Invoice number
- Invoice date
- PO reference
- Net amount
- VAT
- Total due
- 1
OCR with layout
Text comes back with coordinates, so every value can be traced to a region of the page.
- 2
Classify
Document type and vendor decide which schema and which rules apply.
- 3
Extract with confidence
Fields come back typed with a confidence level; values not found in the OCR text are rejected.
- 4
Validate and route
Rules check POs, vendors and totals against the ERP; anything uncertain goes to review.
Guardrails
- An extracted value must appear in the OCR text of the page. The model can't invent one.
- Reviewer corrections improve prompts and rules for that vendor, and become evaluation cases.
- Every exported record links back to the document and the region it came from.
Integration
Meets documents where they already arrive
Nothing changed for vendors: they keep emailing invoices to the same address. DocSense sits between that inbox and the ERP.
Email inbox
IMAP
The shared accounts-payable mailbox is polled; attachments become documents linked to their email.
InboundERP
REST API
Reads purchase orders and vendor status for validation, and receives approved records.
Two-wayS3
SDK, signed URLs
Stores originals and page images; the review app reads them through short-lived links.
Two-way
Performance
Throughput without surprises
Month-end brings bursts of documents. The pipeline absorbs them by spreading work across stages and keeping the review app independent of processing load.
Outcome
Reviewers check, they don't type
The team's work shifted from keying documents to checking the few fields the pipeline isn't sure about. Validation now runs at intake, so a wrong PO number is caught the day an invoice arrives rather than at month-end, and every value in the ERP can be traced back to its source.
What's next
Next: line-item matching against goods receipts, reusing the same confidence and review model.
What changed
- Reviewers check highlighted fields instead of typing whole documents
- Validation catches mismatches at intake instead of month-end
Outcomes are described qualitatively. Client figures stay with the client.
Have a similar project?
Whether it's an AI feature that needs to be trustworthy or a system you need to integrate with, tell me what you're building. I reply within one business day with questions and a suggested first step.