← All work

Case study

DocSentry

A document Q&A system that would rather say nothing than say something it cannot source.

Year
2026
Role
Solo build: architecture, backend, retrieval, frontend, deployment
0.943

Context recall of the shipped retrieval config, over 123 labelled questions

0.45

Grounding score floor, checked before the model call

88%

Of 43 real injection prompts caught by the screener

3

Independent paths to a refusal

The problem

A company puts its handbooks, policies, and product docs behind a chatbot so staff stop asking the same questions in Slack. The obvious version of this takes an afternoon to build and is dangerous to use.

Here's why. A language model always produces an answer. Ask it about a refund policy that doesn't exist in the documents and it'll write you one, in the same steady tone it uses for the real ones. Someone reads it, acts on it, and tells a customer something untrue.

So most of the work went into refusal. Making it a real outcome rather than a fallback, and making every answer that does come back traceable to the exact paragraph behind it.

How it works

DocSentry request flow, from question to answer or refusalA question passes through validation and screening, embedding, retrieval, a grounding gate, generation, and citation verification. The model performs two of the six steps. The grounding gate, generation and citation verification can each end the request in a refusal, and all three of those decisions are made in plain Python. The numbered walkthrough below the diagram describes the same flow.QUESTION INValidate and screen500-char cap, 45 injection patterns, flag onlyCODEEmbed as a queryRETRIEVAL_QUERY, not RETRIEVAL_DOCUMENTMODELRetrieve top 5 chunkspgvector, cosine similarityCODEGrounding gatebest score below 0.45, and no model call yetCODEGenerate the answerJSON contract, temperature 0.1, one repair retryMODELVerify citationscited ids must map to chunks actually retrievedCODEAnswer, with the sections behind itRefused3 ways inPlain PythonThe model
The model touches two of the six steps. The grounding gate, the generation contract and the citation check are plain Python, which is why all three routes to a refusal are decisions the model cannot argue with.
What each process is allowed to do to the index and the audit trailThree running processes share one Postgres database holding two things: the vector index and an append-only audit trail. The API reads the index and appends to the audit trail, and can do nothing else to either. The background worker is the only thing that rebuilds the index, which is why changing a document is a job rather than a redeploy. The dashboard reads the audit trail and has no route to the index. Schema migrations run as a different role that the running service never uses.WHAT IS RUNNINGWHAT IT HOLDSAPIanswers one question at a timeRUNTIMEWorkerrebuilds the index on demandRUNTIMEDashboardrefusal rate, latency, scoresRUNTIMEMigrationsowner role, deploy time onlyNOT RUNNINGVector index235 chunks, 25 documentsPOSTGRESAudit trailone row per answered questionPOSTGRESreadrebuildreadappend onlyschemanothing returns from the audit trail to the API. that is the point.
The index and the audit trail share one database, which is what lets a refusal be read back beside the retrieval score that caused it. The permissions are the design. The API reads the index but cannot rebuild it, so changing a document is a background job rather than a redeploy, and it can append to the audit trail but never read or alter it: a process that cannot rewrite the record of what it did keeps a better record.
01

The question is cleaned and screened

Questions are capped at 500 characters and control characters are stripped first. Forty-five patterns then look for prompt-injection phrasing. A match raises a flag that follows the request through, and it never blocks on its own, because blocking on a pattern match silences real users asking real questions.

02

The question becomes a vector

The text is turned into a list of numbers positioning it on a map of meaning, so that questions land near passages that mean the same thing even when they share no words. The technical name is an embedding.

03

Question and document are embedded differently

Documents are embedded as things to be found, questions as things doing the searching. Passing the same text through the same model in the two modes gives two different vectors, and matching improves because a question and its answer are rarely phrased alike. The technical name is asymmetric embedding.

04

Top five chunks are retrieved

The vector database returns the five nearest passages ranked by closeness. Five is a deliberate ceiling: models reliably read the beginning and end of a prompt and skim the middle, so stuffing in twenty passages buries the answer rather than helping.

05

Retrieval quality is judged before the model is called

If the best of those five scores below 0.45, the request is refused right there and no model call happens. Nothing in the documents came close enough, so there is no point paying to have that confirmed. This is the cheapest guardrail in the system and it runs first.

06

The model answers under a strict contract

The five chunks are handed over with labels, and the reply must come back as JSON with three fields: the answer, the list of chunk labels it used, and a boolean saying whether it was confident. Temperature is set to 0.1, so wording stays close to the source instead of drifting into paraphrase.

07

Citations are verified against what was actually retrieved

Cited labels are checked against the chunks that were sent. A label the system never supplied is discarded. This is the step that catches the model inventing a source, and it's the difference between citations that decorate an answer and citations that prove one.

08

Anything unproven is refused

No surviving citations, or the model's own confidence flag set false, and the answer is dropped in favour of a refusal. There are three separate paths to refusal, and only one path to an answer.

09

The answer ships with its sources attached

The user gets the text, the exact document sections behind it, and any injection flags raised on the way in. Nothing is hidden from the person reading it.

What it looks like

Decisions

Every one of these had a cheaper option that would have worked in a demo. What follows is what I picked, and what I turned down.

Chunks of 800 characters with 120 characters of overlap

Big enough to carry a whole idea, small enough that retrieving one doesn't drag in three unrelated topics. The overlap means a sentence sitting on a boundary still appears whole in one of the two neighbours.

Turned down

Whole documents as single units: retrieval gets coarse and the answer drowns in surrounding text. Very small chunks: the matching sentence arrives without the context that made it meaningful.

Refuse before calling the model, based on retrieval score

If nothing relevant was found, the model has nothing to work from and asking it anyway costs money and adds latency for a worse version of the same refusal.

Turned down

Always calling the model and instructing it to refuse when unsure. Models refuse inconsistently under pressure from a persuasive question, and every off-topic request would be billed.

Verify every citation against the retrieved set

A citation is only evidence if it can be checked. Checking is a dictionary lookup, so it costs nothing.

Turned down

Trusting the model's citation strings. That turns citations into decoration: they look like proof while being generated by the same process that might have invented the answer.

Flag injection attempts, never block on them

The downstream design already makes a hijacked model harmless, since an answer without verified citations is refused regardless of what the model was told to do.

Turned down

Hard-blocking on a pattern match. Regex on natural language produces false positives, and the cost of a false positive here is a real employee being told their real question is an attack.

Store the embedding model name inside the collection metadata

Embeddings from two different models are two different maps of meaning, and querying one with the other returns confident nonsense with no error anywhere. Recording the model makes that mismatch loud instead of silent.

Turned down

Assuming the model never changes. It changes, quietly, and the failure that follows looks like a retrieval quality problem rather than a configuration one.

Serve the frontend from the same FastAPI app

Same-origin means no CORS surface in production, one container, one deploy, one thing to keep alive on a free tier.

Turned down

Split frontend and backend hosting. More moving parts and a CORS configuration to get wrong, for a demo that gains nothing from independent scaling.

Compose the pipeline with LangChain, keep the verification outside it

The chain is a framework concern: retrieval, prompt assembly, structured output, tracing. The grounding floor and the citation check are not. They live in plain Python inside the chain, because they are the two decisions that must hold when the model misbehaves, and a framework abstraction is one more thing between me and a guarantee.

Turned down

Expressing the guardrails as prompt instructions or as framework config. Both make the safety property depend on something I cannot read as a straight line of code.

Ship section-aware chunking, and publish the two configurations that lost

Four retrieval configurations ran over the same 123 labelled questions with no model in the loop, so the numbers reproduce exactly on any machine. Section-aware chunking won on ranking rather than on recall: MRR 0.890 against naive chunking’s 0.833, with recall almost unchanged at 0.943 against 0.935. Hybrid BM25 with reciprocal rank fusion came out at 0.927 recall and 0.849 MRR, and adding a lexical rerank on top made it clearly worse at 0.809 and 0.730.

Turned down

Reporting only the configuration that shipped. It reads better, and it hides the finding: the two upgrades everyone recommends for this problem both lost on this corpus.

Grow the corpus before running the ablation, not after

At 26 chunks with a top-five retrieval, every query was returning close to a fifth of the whole corpus, so all four configurations would have landed within noise of each other. At 235 chunks a query returns about 2%, which is the range where the comparison says anything at all.

Turned down

Running the ablation on the corpus that already existed. It produces four rows of numbers on schedule, and the numbers measure the corpus being small rather than the retrieval being good.

Pin structured output to JSON schema mode, not tool calling

The library's default returns null roughly two times in three on this model, and null is not an error. It flows into the verifier, fails the citation check, and comes back as a polite refusal. So a broken model call wears the costume of a correct grounding decision, and the system looks like it is working carefully while quietly refusing valid questions.

Turned down

Leaving the default and trusting the null guard. The guard is still there, but a guard that silently absorbs a two-in-three failure rate is hiding the bug, not handling it.

What it does not do

Written down because I would rather say these first than have them found.

  • The ablation measures retrieval, not answers. It checks whether the passage carrying the answer was fetched, and says nothing about the quality of what was written from it. The answer-quality half is judged by a model, and it has been verified test by test on clean runs rather than in one full-suite pass, because the daily generation quota kept running out.
  • The rerank that lost was lexical. A cross-encoder is the version worth testing and it has not been run, so the reranking result is a result about one cheap reranker rather than about reranking.
  • Every number here comes from one corpus: 25 documents of one company's policies. Nothing measured here says the same configuration wins on a different shape of document.
  • The grounding floor of 0.45 was set by hand from observed behaviour, and it predates the labelled set. There are now 150 labelled questions to tune it against and it has not been tuned against them.
  • The screener catches 88% of injection prompts, which means it misses some. That is survivable only because it is a flag rather than a gate, and the grounding and citation checks are what actually stop an injected instruction from being obeyed.

Stack

Backend
FastAPI, Python 3.12, Pydantic
Orchestration
LangChain 1.0, LCEL chain, structured output, 71-case pytest suite
Retrieval
Postgres with pgvector via langchain-postgres, Gemini embeddings, LangChain splitters
Frontend
React, Vite, TypeScript
Evaluation
150 labelled questions, 4-config retrieval ablation, 50-prompt injection suite over 26 techniques
Observability
LangSmith tracing: per-step latency, tokens, retrieved context
Infra
Docker multi-stage, Celery and Redis for index rebuilds, Alembic migrations, Supabase Postgres, Render

Keep reading