Case study
InvoiceLens
Invoice extraction where the model is not allowed to do arithmetic, so the checks mean something.
- Year
- 2026
- Role
- Solo build: architecture, backend, validation layer, frontend, deployment
Deterministic checks, no model involved
Test cases over the validation layer
Money comparison tolerance
Of invoices routed to human review
The problem
Invoices arrive as PDFs and scans, and somebody retypes them into an accounting system. Slow work, and the errors it produces are quiet ones. A transposed digit in a total looks exactly like a correct total.
Hand that job to a multimodal model and you lose the typing but keep the quiet errors. A model reading a smudged 8 as a 3 reports the 3 with the same calm it reports everything else.
So the system had to catch its own reading mistakes without a human checking every field. Which meant finding something inside the document that could disagree with the model.
How it works
The upload is checked by content, not by name
The first bytes of the file are read and compared against known PDF, PNG, and JPEG signatures. A file called invoice.pdf that is actually something else is rejected here, because the extension is a claim by whoever uploaded it and the leading bytes are not.
Size is capped without loading the file
Exactly one byte more than the limit is read. If that extra byte exists the file is too big and is refused, and the server never holds a 500MB upload in memory to find out. Reading the whole file to measure it is how upload endpoints become a way to take a server down.
The model reads the document and reports only what it sees
Fields come back as structured JSON with every single field allowed to be null. The instruction is explicit that a field which is absent or unreadable must be null, because a schema that demands a value is a schema that teaches the model to guess.
The model is forbidden from computing anything
It reports the subtotal printed on the page. It doesn't add up the line items. This is the load-bearing decision in the whole system, and it exists so the validation step has two independent readings to compare.
Seven deterministic checks run over the extraction
Line items are summed in code and compared to the printed subtotal. Tax is recomputed from the rate. The total is checked against subtotal plus tax. Dates are checked for order and sanity, required fields for presence, confidence against a floor. Plain Python, no model involved, same answer every time.
Money is compared with a tolerance
Two amounts count as equal within 0.02. Decimal fractions cannot be represented exactly in binary floating point, so 0.1 plus 0.2 is not 0.3, and an exact-equality check would flag correct invoices as broken. Two cents is wide enough for rounding and narrow enough to catch a real discrepancy.
Every invoice goes to a human
The status function ignores the flags entirely and returns needs_review every time. Clean invoices are faster to approve, and none of them skip the queue.
Corrections and decisions are guarded
A reviewer can only edit fields on an allowlist, so a correction cannot rewrite arbitrary database columns. Approving or rejecting something already decided returns a 409 rather than silently overwriting the first decision. Only approved records reach the CSV export.
What it looks like
Decisions
Every one of these had a cheaper option that would have worked in a demo. What follows is what I picked, and what I turned down.
The model reads, the model never calculates
If the model both reports the subtotal and adds up the line items, the check compares the model against itself and passes on any invoice where it was consistently wrong. Forbidding the arithmetic turns the check into a genuine second opinion.
Turned down
Asking the model for the arithmetic too, which is the intuitive design and produces a validation layer that cannot detect the errors it exists to detect.
Every extracted field is nullable
Null is real information: it says the document did not contain this. A missing purchase order number should surface as missing, not as a plausible invented one.
Turned down
Required fields with defaults. It removes null handling from the code and moves the problem into the data, where a fabricated value is indistinguishable from a read one.
Float comparison with a 0.02 tolerance
Correct for the problem, honest about the constraint, and cheap. The tolerance is a named constant rather than a magic number scattered through the checks.
Turned down
Exact equality, which fails on arithmetic that is actually right. Decimal end to end is the textbook answer for money and is what I would use if this touched real payments, but it is more plumbing than this system needed.
Human review is mandatory, with no auto-approve path
What this saves is typing. Oversight stays. An auto-approve threshold would be the first thing to quietly widen under pressure, and the first thing to let a bad invoice through.
Turned down
Auto-approving invoices that pass all seven checks. It reads as efficient and it hands the failure mode straight back.
Validation is tested, the model isn't mocked
The validation layer is pure functions over plain data, so it can be tested exhaustively at boundaries: one cent under tolerance, one cent over, empty line items, negative amounts, missing dates. That is where the correctness guarantees live.
Turned down
Mocking the model to test the whole pipeline end to end. It produces tests that pass because the mock was written to make them pass.
409 Conflict on a second decision
Two reviewers with the same invoice open is a normal Tuesday. Failing loudly keeps the first decision intact and tells the second reviewer why.
Turned down
Last write wins. Silent, and it destroys the audit story the review queue exists to provide.
What it does not do
Written down because I would rather say these first than have them found.
- Storage is SQLite on a container disk, which is fine for a demo and wrong for anything real. Postgres plus object storage for the files is the correct shape, and the code is structured so that swap is contained.
- Extraction is synchronous inside the request. A large batch would need a queue and a worker rather than a caller holding a connection open.
- Currency handling assumes one currency per invoice and does no conversion.
- There's no OCR fallback. A scan the model cannot read produces nulls and a low confidence score rather than a second attempt through a different pipeline.
Stack
- Backend
- FastAPI, Python 3.12, Pydantic, pytest
- AI
- Gemini multimodal extraction, JSON schema constrained output
- Frontend
- React, Vite, TypeScript
- Infra
- SQLite, Docker multi-stage, Render
Keep reading