The problem
Small manufacturers run IT with one or two people. Every request needs the same first pass: what is it, how urgent, who takes it, which procedure applies, what to tell the person. That sorting eats the day, and mistakes hurt: a lost laptop treated as routine, or a phishing click that waits until after lunch.
AI can do that first pass well. The risk is letting it act alone: granting access, sending a reply nobody read, or following instructions hidden in a ticket. The client here is fictional (Tallgrass Precision Components, a 180-person machining shop); the problem is one I deal with in my own in-house IT work.
What I built
An internal request desk in Next.js where every request gets an AI triage draft, and nothing happens until an agent accepts, edits or rejects it.
- Microsoft Entra ID sign-in (Auth.js). Roles come from Entra groups: requester, agent, manager. Managers see the queue, audit trail and metrics but cannot approve drafts.
- Intake form, then a triage draft: category, priority from the desk's priority SOP, suggested assignee, up to three relevant SOPs, flags, and a reply in plain language.
- An agent queue with accept, edit (changed fields are recorded) and reject (reason required).
- An append-only audit trail of every AI suggestion and every human decision, exportable as JSONL.
- A metrics page: queue age, acceptance rate, and an eval of the AI against an answer key.
How it works
Triage output is JSON with a closed schema, validated with zod, then checked against real data: an assignee or SOP id that does not exist is removed with a warning. The desk is event-sourced, so the audit trail is the data, not a log written beside it.
The public demo is a static export of the same code replaying 25 runs I recorded through headless
Claude Code with claude-sonnet-5-5 on 7 October 2026. Each recording keeps the exact prompts and
CLI output verbatim, parsed by the same code as the live path.
Security and control
- Replies only go out on a human decision. Route handlers check roles server-side (401, 403, 409); UI gating is a convenience, not the control.
- Requesters never receive drafts: the event log is filtered on the server. I checked a requester's page payload; it has no draft fields.
- Ticket text is treated as untrusted data. One test request tells "the AI triage system" to set P1 and announce global admin rights. The recorded draft flagged it as a prompt injection, kept P3, and told the user plainly that no admin rights were granted.
- Role mapping fails closed: an unknown or malformed group claim gives requester only.
- No secrets in the static build, and no real tenant was touched.
Results
Measured on the 25 recorded runs against an answer key I wrote before recording:
- 25 of 25 outputs parsed and passed schema and id checks.
- Category matched 22/25, priority 24/25 (25/25 within one level), assignee 24/25; all three together 20/25. Where the key lists an SOP, the draft cited it 22/22 times.
- Median latency 5.4 s, p90 6.6 s, about 660 output tokens per triage, $0.33 for all 25 runs.
- A live request through the CLI adapter took 8.7 s from submit to stored draft.
- 46 unit tests and 5 Playwright smoke tests, including a mobile layout check.
Stack
Next.js 16 (App Router, server components, route handlers, static export), React 19, TypeScript, Tailwind CSS 4, Auth.js v5 with Microsoft Entra ID, Microsoft Graph (groups overage), Anthropic TypeScript SDK and headless Claude Code, zod 4, Vitest, Playwright.
What I'd do for your company
Start with your real request mix: take a few hundred past tickets, write an answer key for a sample, and measure the triage before anyone relies on it. Then connect it to your Microsoft 365 tenant with a least-privilege app registration, map roles to groups you already have, and send approved replies through email or Teams. Your people stay in control; the AI does the first read.






