Tony Bleything

COURSE PRODUCTION

It fact-checked the course I wrote by hand.

The human gate: every finding is my call.

Product

Any content operation where publishing a wrong claim is expensive to undo.

In legal and regulated marketing: Client-facing guidance checked line by line against primary sources, with every flagged claim a lawyer's call to fix, cut or keep.

A pattern, not a past engagement: it describes how this system would be used, not where it has been.

no unattended path to a buyer-facing URLResearches the topic90 days of evidenceDesigns and writesspine, then every wordCompiles and rendersvalidated, three widthsThree reviewerspedagogy · sources · voicethey write findings; none can edit the coursefindingsTHE GATEI rule on every findingmy callCourse publishedto a URL I chose, by hand
The machine checks every claim. It never decides what to do about one.
Eight stages from topic to published site.
Eight stages from topic to published site.
The gate: every finding is my call, fix, cut, reword or record.
The gate: every finding is my call, fix, cut, reword or record.
Published by hand, from a terminal.
Published by hand, from a terminal.

The stages

  1. Researches the topic against the last 90 days of evidence
  2. Works out what learners should be able to do, and the backbone of the lessons
  3. Writes every word
  4. Builds it into a course file and checks the format is actually valid
  5. Opens every screen in a real browser at three different widths, to see it as a learner would
  6. Reviews, publishes, then drafts the announcements

Where the gate sits

I set the standard and decide on every finding. The publish step checks that a real person is actually sitting there and refuses to run without one, including the time I told it to go ahead anyway.

What moves between stages

Files move through one folder at a time, one stage per folder. Three reviewers work at once with deliberately separate remits — teaching quality, sources, and voice — and none of them is allowed to change the course itself. They write up what they found, the writer revises, and the final check reads the actual file rather than any reviewer's account of it.

What broke, and how it surfaced

A reviewer asked to check accuracy checks meaning, so quotations that had been subtly altered sailed through reviews that declared them word for word. The fix was to make it compare every quoted passage character by character. A format checker then got it wrong in both directions at once, rejecting seven perfectly good objectives and waving through two problems it should have caught. And a safety check that gets overridden every time it fires — three legitimate rebuilds refused in a single day — is a design problem, not a discipline problem.

Outcomes, with sources

90claims checked against primary sources in one audit — the scale a person will not sustain by handSource The two accuracy reviews dated 8 Sep 2026, kept in the courseforge project
13flagged as wrong, and every one of them my call: fix, cut, reword, or record and leaveSource The same two reviews. My ruling on each flag is written in the course's decisions log
0courses published to a buyer without me at the keyboard; the script checks, and refuses, including the time I told it to go aheadSource How it's built: the publish script's final check turned down every publish an AI agent started, and every live release was run by hand

I built a system where AI workers research, design, write, assemble and quality-check a course, and I decide on what they produce. Then I pointed its accuracy checker at a course I had written by hand and published myself. It found thirteen platform claims the vendor's own documentation contradicts, and four quotations that had been altered inside their quote marks.

Read the full story · about 2 minutes

The use case

As an instructional designer, I want the mechanical half of course production — research, structure, drafting, compiling, visual QA, fact-checking — executed by agents, so that my time goes to the judgment calls: is this the right topic, does this lesson actually change what someone does on Monday, and is this claim true.

The problem

Course production has a quality problem that reads like a writing problem. The failure that damages a course is rarely bad prose. It is a confidently stated instruction that is wrong — a menu path that moved, a feature that belongs to a different product, a statistic quoted with three words removed. Fluent writing hides these, because a claim whose meaning is intact reads as correct. Neither compiling, nor rendering, nor reading the course aloud will catch one. Only opening the source will, and no human sustains that across two hundred blocks.

How it works

Eight stages, each run by an agent or a script, passing files through one folder a stage at a time: research the topic against the last 90 days of evidence, design the objectives and lesson spine, write every word, compile to a validated course format, render every screen at three widths in a real browser, review, publish, draft distribution. Three reviewers run in parallel at the QA stage with deliberately non-overlapping authority — one judges pedagogy, one opens every source, one judges voice — and none of them can edit the course. They write findings; the writer revises; the gate reads the artifact rather than any agent’s claim about it.

Where the human sits

I own the topic, the standard, and every ruling. The system is built so it cannot quietly move that line: no agent grades its own work, the gate verifies the file rather than a status field, and unattended publishing is restricted by name to an internal namespace, so anything buyer-facing refuses to deploy without a person at a terminal. When the fix loop hits its round limit with defects still open, it records them and publishes anyway rather than looping, because a live URL with a known list of faults is more useful to me than no URL and a clean report.

Demo

Two live courses, both produced by the pipeline: courseforge-brief-is-the-job.vercel.app and courseforge-playbook-by-pipeline.vercel.app. The second was built as a controlled benchmark, on the same ground as a course I had already written by hand, with every agent blinded to the existing one. The comparison is written up lesson by lesson in the repo, evidence first and without a verdict.

Outcomes

The benchmark’s real finding was not that the pipeline writes well. It was where it fails: the same pipeline that produced three factual faults on a conceptual topic produced seven on a topic whose truth lives inside a product, because it can read documentation but cannot use software. That is a precise, useful limit, and it tells you what to keep a human for.

The more valuable result ran the other way. I turned the accuracy reviewer on the course I had written myself and shipped. It checked 90 platform claims and found 13 the vendor’s current documentation contradicts — including one that stopped every new learner on the first real task of Day 1, and a whole lesson built around a tool that exists in a different product. Those errors were not careless. They were true when written, and the product moved. That is the half-life of documentation inside teaching material, and this reviewer is the only thing I have found that measures it.

What broke in production

Four things worth naming. A fix loop’s characteristic failure is not carelessness but scope: three consecutive rounds of careful instruction failed to carry one conditional path correctly through five lessons, and deleting the path fixed it in a single round. Reviewers checking accuracy check meaning, so quote alteration survives them — three altered quotations passed reviews that called them verbatim, and the fourth surfaced only when I changed the instruction from “verify these claims” to “check every quoted span character by character.” A validation gate that matches a syntactic pattern fails in both directions: mine once rejected seven perfectly good objectives over formatting, and later reported two false positives as real defects. And a safety guard that everyone routes around has stopped being a guard — mine refused three legitimate recompiles in one day and was overridden every time, which is a design problem rather than a discipline problem.