The problem
A lot of apps now start as a prompt. They demo well, then go live with the same security holes: no row security, one customer reading another's data, secrets in the browser, forgeable sessions. Owners can't tell, because the app works.
I wanted to show the rescue work end to end on a realistic app, with proof instead of a checklist.
What I built
The app is a supplier portal for Tallgrass Precision Components, a fictional machine shop. Suppliers log in, see their purchase orders, upload certificates of conformance, and message buyers.
I generated the first version with headless Claude Code (claude -p, Sonnet) on 2026-10-06, prompting
it as a non-technical founder. The run is recorded in the repo. Its first pass was more careful than
I expected: parameterized queries, bcrypt, signed sessions, same-origin CORS, upload checks.
That isn't what most rescue jobs look like, so I reverted the before/ app to the common
vibe-coded flaw classes and documented each one I injected. The model did not write those flaws.
Then I built the hardened after/ app and an automated test suite that attacks both.
How it works
In the hardened app, every request carries a signed JWT. The middleware verifies it and sets a nonce
CSP. Route handlers connect as a least-privilege portal_app role and set the user id
inside each transaction. Row-Level Security policies then decide which rows that user can see. A supplier
asking for another supplier's purchase order gets a 404, because the row isn't visible to them.
The test runner fires each exploit at both running apps and writes the evidence to
tests/results.json and audit/RESULTS.md.
Security and control
- RLS enabled on all tenant tables, with policies and a non-superuser app role.
- Signed sessions (jose). A forged cookie gets a 401.
- All SQL parameterized. The part-number search no longer dumps the users table.
- Login throttled per IP and account. Failed attempts get a 429.
- Uploads must be real PDFs, size-capped, served as
application/pdfattachments with nosniff. - No secrets in the client bundle.
.envis git-ignored. Boot-time env validation. - Errors logged server-side with a reference id. Clients get a generic message.
- Nonce CSP, HSTS, X-Frame-Options DENY, Referrer-Policy, Permissions-Policy.
- All keys in the repo are obviously fake (
sk-FAKE-...). No real tenant or data was touched.
Results
Measured by tests/run.mjs on 2026-10-07:
- 13 findings: 4 Critical, 4 High, 4 Medium, 1 Low.
- 13/13 exploits succeed against the before app. 0/13 succeed against the after app.
- 8/8 functional tests pass against the after app: login, wrong-password rejection, supplier sees only their own POs, buyer sees all, search, messaging, PDF upload, health check.
- SQL injection in before dumped 6 credential rows. In after it returns nothing.
- A forged session cookie in before exposed 10 POs across 3 suppliers. In after it gets a 401.
- Login brute force: 12 rapid attempts never throttled in before. After throttles at 6 attempts.
- Dashboard N+1: 10 separate certificate queries per list in before, 1 in after.
Stack
Next.js 14 (App Router, JavaScript), Postgres 16 with Row-Level Security, pg, bcryptjs, jose,
zod, Docker. Tests in Node. Screenshots are real Playwright renders. First pass generated with
Claude Code (claude -p, Sonnet).
What I'd do for your company
If you have an app built with Claude, Cursor, Lovable, Bolt or similar and want it in front of real customers, I would:
- Read the code and the database setup, and write up each finding with file and line.
- Write an exploit for each one, so you can see it is real.
- Fix them without rewriting what already works, and add tests that prove the app still does its job.
- Hand back the fixed code, the test suite, and a plain-language report.
Tallgrass Precision Components is a fictional company used for this demo.







