Behind the case studies

The systems I build to work with agents

Agents write most of my code. What remains is the part that decides whether any of it can be trusted: specifying work precisely enough to verify, checking what comes back, and placing my judgment where being wrong is costly. Six systems hold that line when I’m not watching.

Each system below leads with its single most useful idea, so those six lines are the short version. The detail under each is there if you want to check it.

01

Agent-filed issue intake

Agents file their own issues, check them against the backlog, and cross-link the overlaps.

The cost isn’t fixing issues faster – it’s rebuilding context. Three issues in the same module should pay that cost once.

How it works

New issues get checked against open ones before they are filed. Near-duplicates fold in. Real overlaps — same module, same root cause — get a note on each issue describing the relationship.

The guardrail

Auto-close requires 50 human-labeled cases at 95% precision with zero wrong closes. Append requires 20. The numbers come from the confidence bound: with zero errors, 20 cases only pin the true error rate near 14%; 50 bring it near 6%. Filing caps at 15 issues a day.

Why this was the hard part

Deciding what the system may close on its own. A wrong close hides open work. A wrong comment is ignorable. Those are not the same risk – they don’t get the same bar.

02

Weekly roadmap re-sort

Once a week, the roadmap re-ranks itself against the strategy and tells me what to build next.

The sorter is graded on the decision it reached – not on how well it explained itself. The explanation is what you read, which makes it the easiest thing to be fooled by.

How it works

It reads the strategy, the open issues, and the code each one touches, then ranks with reasoning attached. Priorities I approve get applied to the board automatically.

The guardrail

Thin or ambiguous evidence produces recorded uncertainty, not a confident rank. A Claude and GLM panel labels the automatic cases together – one family grading its own call is not a second opinion.

Why this was the hard part

Separating a good decision from a good write-up. Evidence resolves against a pinned commit, so a rank cannot rest on code that has since changed.

03

Parallel worker coordination

Each issue spins up its own worker in its own worktree. A coordinator runs them in parallel, caps concurrency, and only surfaces to me when a decision is required.

My answer binds to the exact question that triggered it. If the worker has moved on, the answer is refused – not silently delivered into the wrong context.

How it works

The coordinator launches one worker per issue, limits how many run simultaneously, and follows an event stream. A worker’s question reaches me verbatim. A finished worker reports its pull request – nothing more.

The guardrail

Every relay carries a token tied to its question – late answers are refused outright. Events are acknowledged in order; replaying one is safe. A send that can’t be confirmed is reported, never retried automatically.

Why this was the hard part

Cheap execution means you can launch more work than you can inspect. This layer’s job is to subtract attention, not add surface area. The coordinator never does a worker’s task, never answers its product questions, never merges. The moment it starts deciding on content, it becomes a shortcut around the gates it exists to feed.

04

Easy-issue auto-fix

Issues with a runnable proof get fixed and returned with their evidence attached.

The check must fail before the fix – for exactly the reason the issue predicted. A test that only passes proves nothing unless you watched it fail first.

How it works

Red run at base. Apply the fix. Survival assertions confirm the check wasn’t weakened – same bytes, no dropped cases. Byte-audit the changed files. Replay the whole sequence in a second fresh worktree.

The guardrail

Three outcome classes – verified, refused, and harness-fault – so a tooling problem never reads as a bad fix. Candidate code runs minimal and never sees my credentials. Results arrive as drafts.

Why this was the hard part

Every expensive failure here has been a verification failure, not an implementation one. A rewrite corrupted encoding while every gate stayed green. A cleanup loop deleted a block of tests and the suite kept running – reporting four fewer, no alarm. No gate may report success without having checked. The harness’s hardest job is watching its own gates.

Where it was wrong

921 tests passed before it ran end to end. The first six real runs found six structural defects anyway – including a ledger whose stage order was impossible in either direction. Same root cause every time: fixtures pre-seeding state the real flow was supposed to create.

05

Unattended worktree cleanup

Parallel work leaves finished worktrees behind. A daily job reclaims the dead ones without losing anything.

Every safety probe can answer unknown, and unknown stops the deletion. Most of the engineering here is refusing to let uncertainty collapse into yes.

How it works

Each tree gets a disposition. Merged into main – removed. Clean but unmerged – the folder goes, the branch stays. Carrying stray docs – those are rescued to a draft pull request first. Untouched long enough – archived with a manifest that can rebuild it. Everything else is held.

The guardrail

Rescued docs are scanned as the exact bytes being pushed, checked for credential shapes, then pushed fast-forward only with per-file readback. Any tree a live session touched in 72 hours is held – read from job leases, not guessed from timestamps.

Why this was the hard part

The only cases that matter are the ones you cannot prove safe. A failed status call, a truncated page of API results, an unreadable session log. None of them is allowed to mean yes.

06

Norm, a curated research library

A research library agents consult autonomously when the work touches it.

“Nothing matched – treat this as ungrounded” is a required output. A tool that always returns something is indistinguishable from one that fabricates.

How it works

Notes live under declared topics, each carrying a summary, usage guidance, caveats, and pointers to whatever corrects them. Retrieval merges keyword and semantic recall. Agents pull at defined triggers – not when asked, but when the work demands it.

The guardrail

Consults from other projects are read-only – retrieval cannot rewrite what it cites. Sponsored sources are flagged in the note and count against independence.

Why this was the hard part

Confident retrieval on a topic the library doesn’t hold is indistinguishable from knowledge. So every answer surfaces sources, caveats, and disagreements. Claims are labeled by whether their sources are genuinely independent. Three notes drawn from one interview cannot pose as corroboration.

Which model does what — and what it may not decide

Routing is not about collecting models. Each one owns work it is suited to, and each one has a boundary that keeps it from quietly becoming the deciding voice.

Claude

Owns

Product intent, plan authorship, orchestration, and final synthesis.

Does not decide

No vote on technical deadlocks – those are settled with code, tests, traces, or a contract, not by polling another model.

Codex

Owns

Repository grounding, implementation, debugging, tests, and most code review.

Does not decide

Its review of its own work counts as verification, never independent corroboration.

Astra

Owns

The senior checkpoint when a result can shift an architectural or release decision.

Does not decide

Capped at one full pass plus one bounded delta pass; not spent on formatting or routine fan-out.

GLM

Owns

Independent premise falsification: what framing or assumption is most likely wrong, asked before the solution locks.

Does not decide

Advisory only – never implements, never authorizes scope expansion, and sees the problem without my proposed answer attached.

Fable

Owns

Bounded design arbitration when two viable options diverge in their consequence for the user.

Does not decide

Same model family as the planner, so it contributes design judgment, not independent corroboration.

Code gets reviewed by a different model family than the one that wrote it – same-family review inherits the author’s blind spots. Every commit is mine. I read the diff first.

Where I overruled a model

Routing work to the right model is useful. Letting one set product scope is not. A design model once recommended an inspectable eligibility checklist before a user applies. I kept the simpler single check for the first release and moved the richer version into its own plan.

Where the evidence comes from

Research and retrieval sit outside the models that do the work. Each one has a job and a boundary.

Norm

My curated research and pre-worked decisions – caveats and disagreements included.

Returns ungrounded rather than filling gaps from memory. Every cross-project consult is read-only.

Manus

Bounded outside research when a decision turns on current evidence the repository cannot supply.

Research context only, on a capped budget. Claims that drive a decision get verified before they are used.

Podscan

Practitioner discussion and the language customers actually use, cited to a specific episode.

One person’s account stays an account – never promoted to market evidence.

Monid

A route to structured data or a real API, checked before anyone writes a scraper that breaks next month.

Consulted to find a path, not used in place of a dedicated integration that already exists.

These are not reviews. They supply evidence; the models still have to judge it. A research tool agreeing with me is not corroboration.

Current as of . These systems are under active development; the engineering described here is what they are built on.

Building something complex? Let's talk.

Looking for my next Product Engineer role on a small AI-native team.

Get in Touch

Designed & Built by Drew Miller

© 2026. Version 3.3