The short version
I built a system where AI workers research, design, write, assemble and quality-check a course, and I decide on what they produce. Then I pointed its accuracy checker at a course I had written by hand and published myself. It found thirteen platform claims the vendor's own documentation contradicts, and four quotations that had been altered inside their quote marks.
Read the full story · about 2 minutes
The use case
As an instructional designer, I want the mechanical half of course production — research, structure, drafting, compiling, visual QA, fact-checking — executed by agents, so that my time goes to the judgment calls: is this the right topic, does this lesson actually change what someone does on Monday, and is this claim true.
The problem
Course production has a quality problem that reads like a writing problem. The failure that damages a course is rarely bad prose. It is a confidently stated instruction that is wrong — a menu path that moved, a feature that belongs to a different product, a statistic quoted with three words removed. Fluent writing hides these, because a claim whose meaning is intact reads as correct. Neither compiling, nor rendering, nor reading the course aloud will catch one. Only opening the source will, and no human sustains that across two hundred blocks.
How it works
Eight stages, each run by an agent or a script, passing files through one folder a stage at a time: research the topic against the last 90 days of evidence, design the objectives and lesson spine, write every word, compile to a validated course format, render every screen at three widths in a real browser, review, publish, draft distribution. Three reviewers run in parallel at the QA stage with deliberately non-overlapping authority — one judges pedagogy, one opens every source, one judges voice — and none of them can edit the course. They write findings; the writer revises; the gate reads the artifact rather than any agent’s claim about it.
Where the human sits
I own the topic, the standard, and every ruling. The system is built so it cannot quietly move that line: no agent grades its own work, the gate verifies the file rather than a status field, and unattended publishing is restricted by name to an internal namespace, so anything buyer-facing refuses to deploy without a person at a terminal. When the fix loop hits its round limit with defects still open, it records them and publishes anyway rather than looping, because a live URL with a known list of faults is more useful to me than no URL and a clean report.
Demo
Two live courses, both produced by the pipeline: courseforge-brief-is-the-job.vercel.app and courseforge-playbook-by-pipeline.vercel.app. The second was built as a controlled benchmark, on the same ground as a course I had already written by hand, with every agent blinded to the existing one. The comparison is written up lesson by lesson in the repo, evidence first and without a verdict.
Outcomes
The benchmark’s real finding was not that the pipeline writes well. It was where it fails: the same pipeline that produced three factual faults on a conceptual topic produced seven on a topic whose truth lives inside a product, because it can read documentation but cannot use software. That is a precise, useful limit, and it tells you what to keep a human for.
The more valuable result ran the other way. I turned the accuracy reviewer on the course I had written myself and shipped. It checked 90 platform claims and found 13 the vendor’s current documentation contradicts — including one that stopped every new learner on the first real task of Day 1, and a whole lesson built around a tool that exists in a different product. Those errors were not careless. They were true when written, and the product moved. That is the half-life of documentation inside teaching material, and this reviewer is the only thing I have found that measures it.
What broke in production
Four things worth naming. A fix loop’s characteristic failure is not carelessness but scope: three consecutive rounds of careful instruction failed to carry one conditional path correctly through five lessons, and deleting the path fixed it in a single round. Reviewers checking accuracy check meaning, so quote alteration survives them — three altered quotations passed reviews that called them verbatim, and the fourth surfaced only when I changed the instruction from “verify these claims” to “check every quoted span character by character.” A validation gate that matches a syntactic pattern fails in both directions: mine once rejected seven perfectly good objectives over formatting, and later reported two false positives as real defects. And a safety guard that everyone routes around has stopped being a guard — mine refused three legitimate recompiles in one day and was overridden every time, which is a design problem rather than a discipline problem.


