The short version
One person and a system of ten AI workers produced a 20-course, roughly 380-lesson education library for a client in six weeks of active development. Three points inside the process required a person to approve, and none of them could move forward on its own: every course a human passed, a human had read.
Read the full story · about 2 minutes
The use case
As a solo consultant holding a twenty-course production contract for a wellness-industry education company, I want an agent pipeline to carry the production load, so that six weeks of my judgment goes into instructional design and review instead of typing.
The problem
Twenty courses and roughly 380 lessons is a team-sized workload. The naive way to point AI at it, one long conversation that drafts course after course, degrades fast: context bloats, voice drifts, and by course twelve the model is quietly contradicting decisions made in course two. The hard part of course production at this scale is not generating words. It is holding a design standard steady across 380 lessons, and knowing which outputs a human has actually looked at.
How it works
The engine is a ten-agent pipeline run in strict sequence, each agent with one job: strategy, objectives, architecture, instructional design, structure, storyboarding, and quality assurance, through to packaged lessons. Two rules give it its reliability. Every agent conversation starts fresh, with no accumulated context to drift in; the only state that moves between agents is a structured handoff packet that a human reads and pastes forward. And no content ever lives in the chat itself; every output is a file on disk, versioned and reviewable. When a review found a problem, a targeted revision command reran the one agent responsible and everything downstream of it, rather than regenerating the world.
Where the human sits
Three review gates, and they cannot auto-advance. The pipeline stops at each one until a human has read the work and passed it, and the handoff packet design means I saw the state of the build at every transfer, not just at the end. This is the conducted model in its most literal form: the agents execute, the human owns the score. The pipeline never decided that its own work was good enough.
Demo
This is client work, so the demo is a self-authored, sanitized architecture diagram of the pipeline and its gates, not a recording. No client artifact appears anywhere on this site.
Outcomes
Twenty courses shipped, meaning published to the client, in six weeks of active development by one person and the pipeline. Roughly 380 lessons across two tracks. The consistency the pipeline was built for showed up where it counts: 396 lesson headers conformed to the design standard with zero defects, and over 900 editable diagram sources were produced alongside the published courses so the client can maintain what they own. All figures are from the engagement’s own records.
What broke in production
The honest part: shipped is not finished. The library published with a known punch list. Some learner downloads referenced in lesson copy did not yet exist at publication, several courses lacked a final packaged source, and a settings pass that should have run before publication ran after it instead. The pipeline’s gates caught what they were designed to catch, structure and instructional quality, and did not catch what nobody had assigned them: finishing details outside the generation flow. The close-out named every gap rather than papering over it, and a post-publication QA pass has been working the list down since.
The lesson carried forward is about gate design, not gate quantity. A human gate inspects what it is pointed at. If the checklist does not include the downloads, three careful reviewers will still approve a course with a missing download. The gates were the right architecture; the punch list taught me what else belonged on their checklists.


