The problem
A small manufacturer keeps its rules in dozens of controlled documents: quality procedures, work instructions, IT policies, HR handbook sections, safety and export rules. People ask the same questions every day, like how long MRB has to disposition an NCR, or whether a contractor can sign off an FAI. Most "chat with your docs" prototypes fail in three ways. They answer from an old revision. They make up an answer when the documents are silent. And they show HR or IT material to people who should never see it.
What I built
I built a document assistant for Tallgrass Precision Components, a fictional AS9100 machine shop. I wrote 39 internally consistent documents for it: numbered sections, form numbers, retention periods, and three superseded revisions that act as traps. The assistant:
- answers with inline citations like
[QP-7.5 §5.4]that open the exact section in a document viewer; - says "I can't find that in the documents available to you" instead of guessing;
- drops documents the user can't read before anything is ranked, so restricted text never reaches the model;
- answers from the current revision and sets superseded ones aside.
I also built the measurement: a 40-question eval set covering five failure modes, deterministic retrieval metrics, and an LLM judge. Every model call is recorded verbatim.
How it works
Chunks follow the document headings, so every chunk maps to a section number a citation can point at. Retrieval is hybrid. BM25 handles exact identifiers like F-750-04. A small embedding model running locally on CPU handles paraphrase. Reciprocal-rank fusion merges the two rankings. Live mode calls Claude through the Anthropic SDK, or through headless Claude Code for local runs. The public demo replays recorded runs and does real BM25 search in the browser, scoped to whichever persona you pick.
Security and control
Permission trimming happens before ranking, and the context is checked again before the prompt is built. In Graph mode, a user's groups come from Microsoft Graph. Their identity comes from the authenticating proxy, never from the request body; with no identity header the API returns 401. The model gets no tools, and document text is fenced as content, not instructions. The static demo ships the fixture documents to the browser so the persona switcher works offline. It shows trimming but does not enforce it, and the README says so.
Results
All answers were recorded on 2026-10-07: generated by claude-sonnet-5-5, judged by claude-opus-5-5.
- Corpus: 39 documents, 32,467 words, 532 chunks.
- Retrieval MRR@10 (document level): BM25 0.908, embeddings 0.948, hybrid 0.966.
- Restricted passages retrieved, across all 40 questions and 3 retrievers: 0.
- Refusals on unanswerable and permission-restricted questions: 11/11 correct, with 0 false refusals on the 29 answerable ones.
- Superseded-revision traps: 6/6 correct.
- Citations pointing at passages the model was actually given: 69/69.
- Strictly correct answers: 36/40 in the first run, 38/40 after one change (8 context passages instead of 6). The judge found every answer grounded in its excerpts.
The four first-run misses were all retrieval misses. In each one, the model said a fact was missing rather than guessing. Two still fail, because the answer sits in a document the question never names. Fixing that needs multi-hop retrieval, which I have not built.
I also had to fix my own judge. The first rubric graded answers only against the retrieved text, and it scored those retrieval misses as correct.
Stack
TypeScript, Next.js 16, React 19, transformers.js (bge-small-en-v1.5), hand-written BM25 and RRF, Anthropic SDK, headless Claude Code, MSAL + Microsoft Graph, Vitest, Playwright.
What I'd do for your company
I'd start with your real document set and permission model: SharePoint sites, Entra groups, revision control. Then I'd write an eval set from the questions your people actually ask, before tuning anything, and set it to run on every change so you can see the effect of each one. You get an assistant that cites, refuses and respects access, plus the numbers that show it does.






