Case study
DocSentry
A document Q&A system that would rather say nothing than say something it cannot source.
- Year
- 2026
- Role
- Solo build: architecture, backend, retrieval, frontend, deployment
Context recall of the shipped retrieval config, over 123 labelled questions
Grounding score floor, checked before the model call
Of 43 real injection prompts caught by the screener
Independent paths to a refusal
The problem
A company puts its handbooks, policies, and product docs behind a chatbot so staff stop asking the same questions in Slack. The obvious version of this takes an afternoon to build and is dangerous to use.
Here's why. A language model always produces an answer. Ask it about a refund policy that doesn't exist in the documents and it'll write you one, in the same steady tone it uses for the real ones. Someone reads it, acts on it, and tells a customer something untrue.
So most of the work went into refusal. Making it a real outcome rather than a fallback, and making every answer that does come back traceable to the exact paragraph behind it.
How it works
The question is cleaned and screened
Questions are capped at 500 characters and control characters are stripped first. Forty-five patterns then look for prompt-injection phrasing. A match raises a flag that follows the request through, and it never blocks on its own, because blocking on a pattern match silences real users asking real questions.
The question becomes a vector
The text is turned into a list of numbers positioning it on a map of meaning, so that questions land near passages that mean the same thing even when they share no words. The technical name is an embedding.
Question and document are embedded differently
Documents are embedded as things to be found, questions as things doing the searching. Passing the same text through the same model in the two modes gives two different vectors, and matching improves because a question and its answer are rarely phrased alike. The technical name is asymmetric embedding.
Top five chunks are retrieved
The vector database returns the five nearest passages ranked by closeness. Five is a deliberate ceiling: models reliably read the beginning and end of a prompt and skim the middle, so stuffing in twenty passages buries the answer rather than helping.
Retrieval quality is judged before the model is called
If the best of those five scores below 0.45, the request is refused right there and no model call happens. Nothing in the documents came close enough, so there is no point paying to have that confirmed. This is the cheapest guardrail in the system and it runs first.
The model answers under a strict contract
The five chunks are handed over with labels, and the reply must come back as JSON with three fields: the answer, the list of chunk labels it used, and a boolean saying whether it was confident. Temperature is set to 0.1, so wording stays close to the source instead of drifting into paraphrase.
Citations are verified against what was actually retrieved
Cited labels are checked against the chunks that were sent. A label the system never supplied is discarded. This is the step that catches the model inventing a source, and it's the difference between citations that decorate an answer and citations that prove one.
Anything unproven is refused
No surviving citations, or the model's own confidence flag set false, and the answer is dropped in favour of a refusal. There are three separate paths to refusal, and only one path to an answer.
The answer ships with its sources attached
The user gets the text, the exact document sections behind it, and any injection flags raised on the way in. Nothing is hidden from the person reading it.
What it looks like
Decisions
Every one of these had a cheaper option that would have worked in a demo. What follows is what I picked, and what I turned down.
Chunks of 800 characters with 120 characters of overlap
Big enough to carry a whole idea, small enough that retrieving one doesn't drag in three unrelated topics. The overlap means a sentence sitting on a boundary still appears whole in one of the two neighbours.
Turned down
Whole documents as single units: retrieval gets coarse and the answer drowns in surrounding text. Very small chunks: the matching sentence arrives without the context that made it meaningful.
Refuse before calling the model, based on retrieval score
If nothing relevant was found, the model has nothing to work from and asking it anyway costs money and adds latency for a worse version of the same refusal.
Turned down
Always calling the model and instructing it to refuse when unsure. Models refuse inconsistently under pressure from a persuasive question, and every off-topic request would be billed.
Verify every citation against the retrieved set
A citation is only evidence if it can be checked. Checking is a dictionary lookup, so it costs nothing.
Turned down
Trusting the model's citation strings. That turns citations into decoration: they look like proof while being generated by the same process that might have invented the answer.
Flag injection attempts, never block on them
The downstream design already makes a hijacked model harmless, since an answer without verified citations is refused regardless of what the model was told to do.
Turned down
Hard-blocking on a pattern match. Regex on natural language produces false positives, and the cost of a false positive here is a real employee being told their real question is an attack.
Store the embedding model name inside the collection metadata
Embeddings from two different models are two different maps of meaning, and querying one with the other returns confident nonsense with no error anywhere. Recording the model makes that mismatch loud instead of silent.
Turned down
Assuming the model never changes. It changes, quietly, and the failure that follows looks like a retrieval quality problem rather than a configuration one.
Serve the frontend from the same FastAPI app
Same-origin means no CORS surface in production, one container, one deploy, one thing to keep alive on a free tier.
Turned down
Split frontend and backend hosting. More moving parts and a CORS configuration to get wrong, for a demo that gains nothing from independent scaling.
Compose the pipeline with LangChain, keep the verification outside it
The chain is a framework concern: retrieval, prompt assembly, structured output, tracing. The grounding floor and the citation check are not. They live in plain Python inside the chain, because they are the two decisions that must hold when the model misbehaves, and a framework abstraction is one more thing between me and a guarantee.
Turned down
Expressing the guardrails as prompt instructions or as framework config. Both make the safety property depend on something I cannot read as a straight line of code.
Ship section-aware chunking, and publish the two configurations that lost
Four retrieval configurations ran over the same 123 labelled questions with no model in the loop, so the numbers reproduce exactly on any machine. Section-aware chunking won on ranking rather than on recall: MRR 0.890 against naive chunking’s 0.833, with recall almost unchanged at 0.943 against 0.935. Hybrid BM25 with reciprocal rank fusion came out at 0.927 recall and 0.849 MRR, and adding a lexical rerank on top made it clearly worse at 0.809 and 0.730.
Turned down
Reporting only the configuration that shipped. It reads better, and it hides the finding: the two upgrades everyone recommends for this problem both lost on this corpus.
Grow the corpus before running the ablation, not after
At 26 chunks with a top-five retrieval, every query was returning close to a fifth of the whole corpus, so all four configurations would have landed within noise of each other. At 235 chunks a query returns about 2%, which is the range where the comparison says anything at all.
Turned down
Running the ablation on the corpus that already existed. It produces four rows of numbers on schedule, and the numbers measure the corpus being small rather than the retrieval being good.
Pin structured output to JSON schema mode, not tool calling
The library's default returns null roughly two times in three on this model, and null is not an error. It flows into the verifier, fails the citation check, and comes back as a polite refusal. So a broken model call wears the costume of a correct grounding decision, and the system looks like it is working carefully while quietly refusing valid questions.
Turned down
Leaving the default and trusting the null guard. The guard is still there, but a guard that silently absorbs a two-in-three failure rate is hiding the bug, not handling it.
What it does not do
Written down because I would rather say these first than have them found.
- The ablation measures retrieval, not answers. It checks whether the passage carrying the answer was fetched, and says nothing about the quality of what was written from it. The answer-quality half is judged by a model, and it has been verified test by test on clean runs rather than in one full-suite pass, because the daily generation quota kept running out.
- The rerank that lost was lexical. A cross-encoder is the version worth testing and it has not been run, so the reranking result is a result about one cheap reranker rather than about reranking.
- Every number here comes from one corpus: 25 documents of one company's policies. Nothing measured here says the same configuration wins on a different shape of document.
- The grounding floor of 0.45 was set by hand from observed behaviour, and it predates the labelled set. There are now 150 labelled questions to tune it against and it has not been tuned against them.
- The screener catches 88% of injection prompts, which means it misses some. That is survivable only because it is a flag rather than a gate, and the grounding and citation checks are what actually stop an injected instruction from being obeyed.
Stack
- Backend
- FastAPI, Python 3.12, Pydantic
- Orchestration
- LangChain 1.0, LCEL chain, structured output, 71-case pytest suite
- Retrieval
- Postgres with pgvector via langchain-postgres, Gemini embeddings, LangChain splitters
- Frontend
- React, Vite, TypeScript
- Evaluation
- 150 labelled questions, 4-config retrieval ablation, 50-prompt injection suite over 26 techniques
- Observability
- LangSmith tracing: per-step latency, tokens, retrieved context
- Infra
- Docker multi-stage, Celery and Redis for index rebuilds, Alembic migrations, Supabase Postgres, Render
Keep reading