The short version
A fixed set of rules decides which specialists take on a task, each one drafts, and a critic challenges every draft in writing before the author revises. An engagement lead turns it into one plan and hands up the questions only a person can answer. The gate then refuses to release the work until I answer them, and overriding it costs my name and a written reason.
Read the full story · about 2 minutes
The use case
As a consultant, I want a team of specialists who will do the analytical work of an engagement and argue with each other about it, so that what reaches me is a considered position with its disagreements intact rather than a confident average.
The problem
The obvious way to build a multi-agent team produces a machine that agrees with itself. Each agent reads the last one’s output and extends it, because that is what cooperation looks like to a language model, and by the time the work reaches a person the doubt has been smoothed out of it. What arrives is fluent, plausible, and unfalsifiable. I built that bias into this system myself without noticing: for three commits the pipeline instructed every downstream specialist to treat its teammate’s work as “data to build on”, which is an instruction to extend rather than question.
The second failure is quieter. An agent team that hands a person a finished plan has made every judgment call on the way there and shown none of them. The human becomes an approver of conclusions instead of a decider of questions.
How it works
A deterministic router staffs the task, and records both who it called in and what would have called in each specialist it did not, so the first question anyone asks about a team — who is missing — has an answer. Each specialist drafts against a written persona and a house brief that every one of them receives. A critic then challenges that draft under a fixed protocol: state the strongest version of the argument first, run a pre-mortem, then file findings against named dimensions. The critic never edits. The author answers every point in writing and reissues the work, so authorship and accountability stay together. An Engagement Lead reads the finished positions, builds one plan, and hands up the questions no advisor can settle.
Everything the engagement accumulates feeds forward. Discovery notes, transcripts and client material dropped into the engagement folder are read before any advisor drafts, and each run records exactly which documents it read, truncated or dropped, so nothing is silently missed.
Where the human sits
At the gate, and the gate is not decorative. Approval attests to the artifact bytes, not to a summary of them: every artifact is hashed when produced and re-verified before any decision is recorded, so editing the manifest or swapping a file is refused. When the team escalates a question it says only a person can answer, the gate will not release the run at all until it is overridden — and the override costs a name and a written reason, recorded permanently as forced.
The parts that make it usable are the quiet ones. A reservation can be recorded without rejecting the whole run, because a gate with only approve and reject teaches people to use neither. A decision can be reopened, superseding the earlier one without erasing it, because a decision nobody can revisit is a decision people avoid making. The history reads forward: approved, reopened, rejected, with who and why at every step.
Demo
A recorded terminal walkthrough of a real run going through the gate ships in the repository. It shows the first approval being refused because the team escalated three questions, the override going through with a written reason, a reservation recorded without rejecting, and the decision being reopened when the reservation turned out to matter. Real provider output from three separate runs is committed alongside it, so the system can be judged without a key or a spend.
Outcomes
The team catches things a person would have to catch. On a live run designing an AI enablement programme, one specialist cited an OSHA page in support of an override-and-accountability design; its critic challenged the citation as not supporting the claim, and the author replaced it with the NIST AI Risk Management Framework, which does. That is a fabricated citation caught by a peer rather than by me, and it is the failure most likely to embarrass a consultant in front of a client.
On the same run, the Engagement Lead escalated three questions upward, including whether the go-live date had been fixed by a commitment nobody had disclosed, and the observation that no safety, security or operations seat had produced anything even though the override protocol depended on all three. The gate then refused to let the run be approved. That is the whole thesis working in one interaction: the team knew what it did not know, said so, and the machine declined to let a human wave it through.
What broke in production
Four things, all instructive, none hidden.
The repository shipped publicly with a live API key committed in a .env file and a 306 MB virtualenv accounting for 14,526 of its 14,673 tracked files. The keys were revoked and the history scrubbed; .git went from 100 MB to 168 KB.
Adversarial review then found three ways to forge an approval in code I had already reviewed three times. Artifact bytes were never verified, so editing one field in the manifest turned a flagged run into a clean approval with no trace. Model output was printed raw to the terminal, so an artifact containing escape sequences could repaint the reviewer’s screen immediately before they typed approve. And the .env loader would set any variable at all, including one that redirects model calls, key attached, to another host. All three are now refused and regression-tested.
The first run against a real model produced excellent content that stopped mid-sentence, because a token budget I had set for cost control was too small — and the missing sections were reported as a content failure rather than a truncation. An operational fault wearing a quality fault’s clothes is worse than an outright crash, so truncation is now named explicitly and reported first.
And the one I am least comfortable with: in the first two live critique rounds, the authors accepted every single challenge, nine of nine and then thirteen of thirteen. That is either consistently good critique or the same agreeableness the loop exists to fight, pointed in the other direction. I could not tell from two runs, so rather than claim it as a success metric I built the instrument that flags it — total agreement now surfaces as a warning to investigate, not a score to celebrate.


