Tony Bleything

THE CONSULTING TEAM

It refused to let me approve it.

The human gate: it stays shut until I answer.

Product

Decisions that need several specialists to disagree in writing before anyone signs.

In professional services and deployed engineering: Proposals and architecture options drafted by specialists, challenged by a critic, and released only under a named human decision.

A pattern, not a past engagement: it describes how this system would be used, not where it has been.

Staffs the task22 specialists · who, and who notThey arguesteelman, pre-mortem, reviseOne plandisagreements kept, not smoothedonly you can answer thisTHE GATEit will not open until I answerapproveReleased to the clientwith my name on the decisionor back to pending.Changing my mind is cheap.every challenge, every override, every reversal is on the record
The team can hand a question up. The gate stays shut until it is answered.
Twenty-two specialists, every draft challenged in writing.
Twenty-two specialists, every draft challenged in writing.
The gate stays shut until I answer what only a person can.
The gate stays shut until I answer what only a person can.
Released, with my name on the decision.
Released, with my name on the decision.

The stages

  1. A fixed set of rules, not a judgement call, picks which specialists to bring in, and records who it left out as well as who it used
  2. Each specialist writes against a written brief of its own and a shared house style
  3. A critic attacks the draft to a fixed routine: state its strongest version first, imagine how it might fail, then write up the findings
  4. The author answers every point in writing and reissues the work
  5. An engagement lead turns it all into one plan and hands up only what a person can settle

Where the gate sits

A named person, not an account, is what moves a run out of pending. The system re-reads the actual file before accepting it rather than trusting a summary, and an unanswered question blocks approval outright. You can override that, but only with your name and a written reason, recorded permanently as forced.

What moves between stages

The client folder is read before anyone drafts a word, and each run records exactly which documents it read, which it cut short, and which it skipped entirely. Every decision is added to a permanent log.

What broke, and how it surfaced

The code went public with a working password sitting in a settings file, alongside 306 MB of installed libraries that should never have been in there. The keys were cancelled and the history wiped. A deliberately hostile review then found three ways to forge an approval, in code I had already reviewed three times: the system trusted a summary of a file instead of the file itself; it printed AI output straight to the screen, where hidden characters could repaint what the reviewer saw an instant before they typed approve; and its settings loader would accept any instruction at all, including one redirecting the work to someone else's server. All three are now refused, with tests that keep them that way.

Outcomes, with sources

3ways to forge an approval, found by a deliberately hostile review before release: trusting a summary of a file instead of the file, printing AI output raw where hidden characters could fake the screen, and a settings loader that accepted any instructionSource Two deliberately hostile reviews, 21 Aug 2026: one by an AI reviewer set up to attack it, one by Codex. The fixes, and the tests that keep them fixed, are in CHANGELOG.md 0.3.0 and tests/test_gate.py
252automated tests passing, covering how work is assigned, the governance rules, the approval gate, the critique routine, client engagements, the dashboard and the command lineSource The full test suite, run at commit bc52b52 on 10 Sep 2026 (uv run pytest), and passing again on the automated build server (CI run 34518357160)
0runs released without a named human decisionSource How it's built: approve and reject are the only ways out of the waiting queue, both require a named person (--by), and every decision goes into a permanent log (logs/runs.jsonl)
22 / 22specialist roles with a written persona, plus a shared house brief every specialist receivesSource The project's evaluation results (evals/results/baseline.json: persona_coverage 1.0, specialists_receiving_house_brief 22). The automated build fails if coverage ever drops
0 → 2explicit assumption labels in the Strategist's output on an identical task, before and after the persona requiring every figure to carry a source or a labelled assumptionSource Two recorded runs of the same task, before and after the persona (20260821_154143_bea32b and 20260821_154701_58d927). The later output is saved at docs/example-run/strategist.md
13 / 13critique points accepted by authors in a live run, recorded as a warning rather than a success, and now flagged automaticallySource The live run's record, 9 Sep 2026 (docs/example-run/enablement/manifest.json). The automatic warning is built into huminloop/stats.py and covered by tests

A fixed set of rules decides which specialists take on a task, each one drafts, and a critic challenges every draft in writing before the author revises. An engagement lead turns it into one plan and hands up the questions only a person can answer. The gate then refuses to release the work until I answer them, and overriding it costs my name and a written reason.

Read the full story · about 2 minutes

The use case

As a consultant, I want a team of specialists who will do the analytical work of an engagement and argue with each other about it, so that what reaches me is a considered position with its disagreements intact rather than a confident average.

The problem

The obvious way to build a multi-agent team produces a machine that agrees with itself. Each agent reads the last one’s output and extends it, because that is what cooperation looks like to a language model, and by the time the work reaches a person the doubt has been smoothed out of it. What arrives is fluent, plausible, and unfalsifiable. I built that bias into this system myself without noticing: for three commits the pipeline instructed every downstream specialist to treat its teammate’s work as “data to build on”, which is an instruction to extend rather than question.

The second failure is quieter. An agent team that hands a person a finished plan has made every judgment call on the way there and shown none of them. The human becomes an approver of conclusions instead of a decider of questions.

How it works

A deterministic router staffs the task, and records both who it called in and what would have called in each specialist it did not, so the first question anyone asks about a team — who is missing — has an answer. Each specialist drafts against a written persona and a house brief that every one of them receives. A critic then challenges that draft under a fixed protocol: state the strongest version of the argument first, run a pre-mortem, then file findings against named dimensions. The critic never edits. The author answers every point in writing and reissues the work, so authorship and accountability stay together. An Engagement Lead reads the finished positions, builds one plan, and hands up the questions no advisor can settle.

Everything the engagement accumulates feeds forward. Discovery notes, transcripts and client material dropped into the engagement folder are read before any advisor drafts, and each run records exactly which documents it read, truncated or dropped, so nothing is silently missed.

Where the human sits

At the gate, and the gate is not decorative. Approval attests to the artifact bytes, not to a summary of them: every artifact is hashed when produced and re-verified before any decision is recorded, so editing the manifest or swapping a file is refused. When the team escalates a question it says only a person can answer, the gate will not release the run at all until it is overridden — and the override costs a name and a written reason, recorded permanently as forced.

The parts that make it usable are the quiet ones. A reservation can be recorded without rejecting the whole run, because a gate with only approve and reject teaches people to use neither. A decision can be reopened, superseding the earlier one without erasing it, because a decision nobody can revisit is a decision people avoid making. The history reads forward: approved, reopened, rejected, with who and why at every step.

Demo

A recorded terminal walkthrough of a real run going through the gate ships in the repository. It shows the first approval being refused because the team escalated three questions, the override going through with a written reason, a reservation recorded without rejecting, and the decision being reopened when the reservation turned out to matter. Real provider output from three separate runs is committed alongside it, so the system can be judged without a key or a spend.

Outcomes

The team catches things a person would have to catch. On a live run designing an AI enablement programme, one specialist cited an OSHA page in support of an override-and-accountability design; its critic challenged the citation as not supporting the claim, and the author replaced it with the NIST AI Risk Management Framework, which does. That is a fabricated citation caught by a peer rather than by me, and it is the failure most likely to embarrass a consultant in front of a client.

On the same run, the Engagement Lead escalated three questions upward, including whether the go-live date had been fixed by a commitment nobody had disclosed, and the observation that no safety, security or operations seat had produced anything even though the override protocol depended on all three. The gate then refused to let the run be approved. That is the whole thesis working in one interaction: the team knew what it did not know, said so, and the machine declined to let a human wave it through.

What broke in production

Four things, all instructive, none hidden.

The repository shipped publicly with a live API key committed in a .env file and a 306 MB virtualenv accounting for 14,526 of its 14,673 tracked files. The keys were revoked and the history scrubbed; .git went from 100 MB to 168 KB.

Adversarial review then found three ways to forge an approval in code I had already reviewed three times. Artifact bytes were never verified, so editing one field in the manifest turned a flagged run into a clean approval with no trace. Model output was printed raw to the terminal, so an artifact containing escape sequences could repaint the reviewer’s screen immediately before they typed approve. And the .env loader would set any variable at all, including one that redirects model calls, key attached, to another host. All three are now refused and regression-tested.

The first run against a real model produced excellent content that stopped mid-sentence, because a token budget I had set for cost control was too small — and the missing sections were reported as a content failure rather than a truncation. An operational fault wearing a quality fault’s clothes is worse than an outright crash, so truncation is now named explicitly and reported first.

And the one I am least comfortable with: in the first two live critique rounds, the authors accepted every single challenge, nine of nine and then thirteen of thirteen. That is either consistently good critique or the same agreeableness the loop exists to fight, pointed in the other direction. I could not tell from two runs, so rather than claim it as a success metric I built the instrument that flags it — total agreement now surfaces as a warning to investigate, not a score to celebrate.