The problem
I use Claude Code, Codex and Grok daily, each with its own memory and its own idea of what it may touch. I wanted one place to write a task in plain words and get a finished report back, with agents that plan, build and check on whichever model fits. That only works if agents cannot send email in my name or run destructive commands, and stop asking me what another agent could answer.
What I built
Four parts, all local:
- An agent runtime (Python). Bots run on Claude Code, Codex or Grok. Each turn gets a throwaway home with only that CLI's login and the runtime's own MCP tool server. Built-in shells are off, so every tool call passes one approval gate.
- A company app (Python, React 19, Vite, Tailwind). A CEO agent routes tasks to departments. Questions climb the reporting line before reaching me. A quality agent signs off before a report is published.
- A memory ledger and model router (TypeScript, Node's built-in SQLite). Append-only, hash-chained, every record sealed under its own key.
- A phone app (Node, installable PWA on my private tailnet) with a passcode and push notifications for questions and reports.
How it works
Each agent has a model chain, typically Claude, then Codex, then Grok. A usage cap marks the run "limited" and requeues it on the next engine.
The gate runs in a fixed order, and the top two rungs cannot be overridden:
if floor(cmd) or matches(deny, cmd): refuse
elif yes_flag or matches(allow, cmd): run
else: apply mode # deny | allow | ask | auto; cron never gets auto
One task forced a redesign: a small product-design job took about 20 hours and 79 agent runs while agents re-derived checks and replanned every lap. Now every choice comes from a closed list settled before work starts, with one fix pass. Of 1,293 recorded decisions, 89% were settled by a rule rather than a model.
Security and control
- Drafts only. Agents never send email or ticket replies. A separate guard checks every proxied MCP call: Graph paths are normalised before matching, batch requests are checked one by one, and send tools are blocked by name. An independent model review found three bypasses; all three were closed. One known gap is written down: calendar invites with attendees.
- Secrets are pulled out before anything is stored and replaced with placeholders. Forgetting a record destroys its key.
- Tools are granted per department, never all-to-all. Everything binds to localhost; the phone app is tailnet-only.
- In October I moved my own install to full access for speed: approvals in scope are auto-allowed. The floor, the deny list and the drafts-only guard still apply. Of 516 approval requests logged, 515 were allowed.
Results
- 607 agent runs from 1 to 6 October 2026 across 8 models on 3 engines: 588 succeeded (96.9%), 11 failed, 5 hit a usage cap and were requeued, 3 were cancelled.
- 20 runs ran on a fallback model further down the chain.
- 1,293 decisions logged: 1,152 by rule, 65 by me, 44 by department default, 32 by a lead.
- Memory ledger: 35,697 records, 2,480 typed facts, 106 secrets moved to the vault.
- Code (clean generic copy): about 21,200 lines of runtime, 27,800 lines of company app, 693 test functions.
- 353 commits across the runtime, company app, ledger, phone app and the clean copy (22 Sep to 6 Oct 2026).
- The clean copy ships with no company data. A build of it was installed on two colleagues' laptops in October 2026.
Stack
Python, TypeScript, Node.js 22, React 19, Vite, Tailwind v4, SQLite (FTS5), Model Context Protocol, Claude Code, Codex CLI, Grok CLI, Web Push, Tailscale, systemd, WSL2.
What I'd do for your company
Start with one team and one task you repeat every week. Agents get only the tools that team already uses, every outbound message is a draft for a person to send, and the deny list is agreed with you before anything runs. Reports come with the decisions written down, so you change a rule instead of re-explaining it.




