Iterative Lesson Refinement
Overview
A QA loop that turns a draft lesson into one that actually teaches and that the learner can
defend live, by exposing it to isolated personas instead of grading your own homework.
Core principle: a lesson's author cannot see its holes, and a persona you role-play inline just
echoes what you already know; isolation is the whole point.
When to Use
- A teaching artifact must survive a skeptical follow-up ("you said cosine, why not Euclidean?"),
not just read well.
- The learner is time-starved: the artifact must be as SHORT as possible while still letting them
do the thing and survive the second question. Brevity is a graded criterion, not an afterthought.
- When NOT to use: content nobody must defend (internal notes, logs); or a first draft that
hasn't been written yet (write the tight draft first, then run the loop).
The Loop (per topic)
- Build/revise a tight draft. Ground every claim to the learner's real work honestly
(HAVE / PARTIAL / GAP). If a companion HTML-craft skill exists, use it for the page itself.
- Dispatch 4 ISOLATED subagents, each reading ONLY the artifact file: no answer key, no
authoring context. Locked roles:
- NOOB: technical generalist who has never met the topic; no outside lookups. Reports
confusion, drag, and density/overwhelm flags; takes the cold quiz. Your density detector.
- BASICS: surface familiarity; mandate = depth gaps, half-explained terms, "asserted
without justifying." Takes the cold quiz.
- EXPERT: practitioner + interviewer; mandate = (a) ACCURACY (quote → problem → fix),
(b) DEFENSIBILITY (the exact follow-up that exposes each weak claim), (c) TIGHTNESS
(a specific cut list). Does NOT take the quiz.
- READER-TWIN (mandatory since 2026-07-06; the slot whose absence let a stall ship):
derive it per artifact from
references/reader-profile.md (the durable profile of the
actual reader: home-stack expertise, zero-knowledge topics, reading traits; update the
profile as the reader demonstrably learns topics). ⚠️ The profile rots in BOTH
directions: it can lag behind learning (stale-low) AND it can overclaim (a 2026-08-14 run
found the profile asserting "production experience" for a skill the reader's own source-of-
truth record bounds to personal projects). When a twin flags a conflict between its profile
and the artifact, that is the twin working: adjudicate against the PRIMARY record (the
reader's ground-truth documents, not either derived file), and fix the PROFILE in the same
pass, or every future twin inherits the error. The real profile is personal and
gitignored; copy references/reader-profile.template.md to create one for your reader. Shape: deep expert in A (their home
stack: the language, IDE, or domain the page translates FROM or leans on), ZERO knowledge
of B (the topic being taught). Mandate: reading in order,
flag (a) any term used before defined (hero copy, diagram labels, and code comments count
as uses), (b) any literal number without a stated origin, (c) any construct in language A
they cannot place that is not labeled invented/pseudo, (d) metaphor switches without a
bridge, (e) read-twice sentences. This slot exists because the other three CANNOT catch a
fake construct in the reader's home language or a home-turf falsehood: a 9-page field sweep
found the generic trio had passed pages 10/10 that the reader-twin then found ~75 defects in.
- Grade the cold quiz against a rubric (answers in the learner's own words, /10); log scores.
Synthesize the reports → verify before applying: an auditor's factual counter-claim ("the
page's number is wrong") is testimony, not truth; check it against a primary source before
editing (field case: an auditor confidently corrected a right number to a wrong one; the doc
fetch caught it). Then revise, and fact-check any NEW fact the revision itself introduces
(fixes are un-audited authoring; a fix pass once shipped a checkably false origin story for
the very number it was fixing). Record the generalizable teaching lesson (not topic facts)
back into this skill.
- Repeat ~2 passes, escalating quiz difficulty each pass: pass 1 tests the spine (can they
describe it?), pass 2 tests the defense layer (the second question). Early-exit when a fresh
NOOB scores ≥8/10 AND the expert flags no accuracy/tightness issues AND a fresh READER-TWIN
reports zero stalls (no term-before-defined, no unexplained number, no unplaceable construct).
- Finalize: bump the artifact's version footer with what changed; clean up test scraps.
The delivery gate (non-negotiable)
Run the WHOLE gate ONCE, UP FRONT, before the learner ever sees the page. Never serially.
The most expensive failure mode is not a bad page; it is reactive gating: ship, let the human
hit a hole, add one QA instrument, ship again, let them hit the next hole. Each round burns their
tokens, their time, and their trust, and it makes the human the trigger for your next QA step,
which is exactly backwards. Before first delivery of any teaching page, in ONE parallel batch:
(1) build it right, most teaching topics are standard, not novel research, so invest the craft up
front; (2) run the full gate at once, the READER-TWIN plus, for a defend-cold/interview page, the
nemesis + interviewer + accuracy panel; (3) fix what survives verification, and re-read the whole
page once (patches can sum to worse); (4) THEN deliver. The human should never meet a pre-gate
artifact, and should never have to name a missing review instrument. If you find yourself adding a
reviewer because the human asked why you didn't, the gate ran too late.
The learner is PRODUCTION, not staging. The loop exists so the struggle happens before they
arrive: they read to LEARN, never to debug. Therefore an artifact may not be handed to the
learner, announced as ready, or registered in any index until the loop, including a clean
READER-TWIN pass, has converged WITHOUT them.
Interview-prep / "defend-cold" pages need an ADVERSARIAL pass too, not just the reader-twin.
The reader-twin is sympathetic; it measures "can I follow this," which catches confusion but
NOT confident-but-wrong claims or unarmed follow-ups. A page whose stated job is surviving a
hostile interviewer must be reviewed by a hostile interviewer. So for any defend-cold page, add
to the gate a nemesis-review panel: the nemesis (armed with the audience's expertise) plus an
interviewer-follow-up skeptic ("for each answer, what is my next question, and does the page arm
it?") plus, for technical pages, a domain-accuracy skeptic. Field case 2026-07-07: a Docker page
that passed the reader-twin 9/10 still had three defend-cold answers that each detonated on the
first follow-up (a healthcheck fix that did not fix the race, an EXPOSE claim that ignored
-P, and a grounding line that overclaimed a skill the reader had not authored, the exact
trap that had already cost a real final-round interview). Only the adversarial panel caught them. This is why the gate is two instruments,
not one: reader-twin proves it teaches, nemesis proves it does not get the reader killed in the room. If the learner ever hits a defect while studying,
treat it as a build failure with a root cause in this loop (usually a persona slot that didn't
match them): fix the page, fold the lesson back into this skill, and tighten the reader profile.
Do not thank the learner for the QA or frame their stall as a contribution; the moment they are
"helping refine," the artifact has already failed at its one job.
Targeted validation mode (cheap variant, added 2026-07-05)
When REVAMPING an artifact whose spine already converged in earlier passes, do not re-run the
full 3-persona loop. Dispatch 2 agents in parallel: a fresh NOOB (cold quiz over old + new
content) and an EXPERT whose fire is pointed at the NEW sections. Give the EXPERT the
artifact's audience expertise explicitly (their stack, their era), not just the topic domain:
audience-home-turf claims are the highest-risk content and a domain-only expert sails past them.
Field result: this 2-agent pass caught a false claim about the reader's own database engine that
three earlier full passes had missed. When the artifact is a TRANSLATION ("B via the A you
know"), make the second agent a full READER-TWIN (A-expert, zero B) rather than a domain expert
with audience seasoning; only the twin catches invented constructs in language A.
Quiz Design (the real gate)
- Author a per-topic quiz; learners answer in their own words (no copy-paste).
- A perfect score on a recall quiz proves nothing: the page teaches its own claims well.
The defense-layer quiz (second-order follow-ups) is what catches recite-vs-own.
- If a learner cannot answer, require them to name WHICH word or idea the page failed to give
them; that names the fix.
The Generalizable Teaching Lessons (the payload)
- Quiz the defense layer, not the page's own claims. Escalate difficulty each pass.
- Define the load-bearing mechanism word on first use. Beginners stall exactly on the words
presented AS the machinery (the #1 observed failure mode across every run), not optional jargon.
- Never name an interview question you don't arm. Flagging "be ready to defend X" without
the material is worse than silence.
- A "defend cold" Q&A must carry NEW depth (second-order follow-ups), never restate the body.
- One metaphor per concept, reused deliberately; three competing metaphors read as padding.
- State key numbers once, with their justification, then reference them.
- When you ground a claim to the learner's real work, pre-arm the obvious follow-up.
"I ran X on my codebase" → "how exactly?" must be answered ON the page.
- Define every borrowed term even inside an advanced section; name-dropping lets a learner
recite without owning, the exact thing a follow-up exposes.
- Mark the difficulty step-up ("second-pass layer") so the beginner isn't silently drowned.
- Lead with the why before the how: plain-English claim → analogy → diagram → defend-it Q&A.
- A defense answer must be self-contained. Never let it lean on a term the page never
defined; that recreates the exact trap the Q&A exists to close.
- By the final pass, weight to TIGHTNESS. Cut any defend-cold card that restates the spine.
Stop signal: fresh NOOB ~10/10 cold AND expert says "ship after ≤2 one-line fixes."
- Check every claim about the reader's OWN home stack against its CURRENT state.
Translation sections ("new thing B via the A you already know") date fastest exactly where
the reader is most expert, and a stale claim there detonates on their home turf. Arm the
EXPERT persona with the reader's home-stack expertise so it hunts these.
- A worked example must not disprove its own lesson, and simulated outputs must be
watermarked IN the artifact. Demonstrating semantic search with a question that near-quotes
its source proves keyword search would have worked too: paraphrase until zero surface tokens
overlap. Realistic invented scores/terminal output without a "simulated" label become
"observed data" one lazy memory later; a caption saying "illustrative" is not enough, mark
the artifacts themselves.
- Define-on-first-use is a WHOLE-PAGE ordering constraint, not a per-section virtue. Hero
copy, diagram labels, and code comments count as uses; a later deep-dive section does not
excuse an unglossed first appearance. Walk the page in reading order before shipping. (The
dominant defect class in a 9-page field sweep, on 8 of the 9 pages.)
- Every literal number carries its origin at first use. What fixes it, whether the reader
can change it, and an analogy class ("a hash width"); mark example values AS examples ("k = 5
is this page's example, not a law") and name the corpus behind scale claims. Invented
precision ("90% of confusion") reads as fake and poisons trust in the real numbers.
- Label invented/pseudo-code as pseudo IN the artifact: before the panel, in the panel
label, and in a code comment, and name the real construct it stands for. An invented function
in the reader's home language makes the expert reader conclude they are ignorant or the page
is lying; both destroy the page.
- Home-turf falsehoods are the fastest trust killers. A wrong claim about the reader's OWN
tool ("your IDE refuses to start without X") outranks any topic error, because it is the one
claim the reader can check instantly. The READER-TWIN slot exists to hunt exactly these.
- Fact-check the ANSWERS harder than the prose, and hunt them as their own pass. A wrong
explanation inside a defend-cold answer is strictly worse than the same error in the body,
because the format instructs the reader to rehearse it and say it out loud under pressure.
Prose gets skimmed; an answer gets memorized. Reviewers reading for clarity glide past a
confidently-wrong answer exactly as fast as a correct one, so accuracy review of the Q&A
cards cannot ride along with the readability pass — give it its own sweep, and treat every
answer as a load-bearing factual claim. (Field case 2026-08-10: a hostile reviewer found a
mechanically wrong protocol-framing explanation that had been written into a closed-book
answer; the honest replacement was also the better answer, the recurring pattern here.)
- Resolve any contradiction the artifact manufactures, on the artifact. Two claims that are
each true and look mutually exclusive several sections apart are a trap the page built and
never disarmed — and when both sit in defend-cold cards, the reader has been instructed to
say both halves aloud. Sweep a finished artifact for pairs of confident claims a hostile
reader would put side by side, then reconcile them in place or cut one. Anticipating the
collision is cheap; meeting it live is not. (Field case 2026-08-10: a rule forbidding a null
identifier and a case that legitimately sends one; the resolution was one sentence — the two
rules govern different message directions.)
Answer-visibility note: past ~4 questions in one block, stacked question-above-answer pairs
stop being a test — a reader-twin scored such a section 1/10 and reported skimming it, "the
answer sitting right under the question... I'm not testing myself, I'm reading a FAQ." When the
artifact is HTML, educational-html-prep carries the collapsed-answer component (.qz) and the
"say it out loud first, then open the row" instruction line.
- Verify a CORRECTION at least as hard as a claim. A pass producing retractions struck a
TRUE sentence and inserted a FALSE one behind a dated banner. The scrub that caught two other
invented claims in the same document sailed past that one, because a paragraph announcing its
own rigour reads as already-checked. A verification artifact carries borrowed trust, which is
exactly what makes it the best hiding place for an error.
- The most confidently written section is the least verified one. Both blocker defects in one
gated build sat in the escalation section - the what-if-they-push-harder material - because it
is prose rather than code and therefore feels unfalsifiable. It is the opposite: it is the most
checkable thing on the page and nobody checks it. Charter one reviewer to check every claim in
the advanced/what-if section against the code the same page prints.
Subagent Prompt Skeleton
You are role-playing [NOOB/BASICS/EXPERT] for a lesson-quality test. Stay in character.
PERSONA: [knowledge level + mandate; for EXPERT also the artifact's AUDIENCE expertise].
HARD RULES: read ONLY ; no web/outside knowledge beyond your level; be specific,
not polite. TASK: [confusion/depth/accuracy log] + [cold quiz answered in your own words].
RETURN in fixed labeled sections (A)…(E).
Files Convention
Keep a run self-contained in a LessonLab/ folder: PROTOCOL.md (rules of the run),
quiz_<topic>.md (questions + rubric), quiz_log.md (scores + per-iteration findings + the
generalizable lessons). The log is durable state across the many turns a loop takes.
Common Mistakes
- Role-playing the personas inline. They echo the author's knowledge; only isolated
subagents produce real confusion data.
- Treating a 10/10 recall quiz as done. It proves the page states its own content well,
nothing more; escalate to the defense quiz.
- Skipping the loop because the draft "reads clean." Every field run found real defects in
drafts that read clean, including factual errors.
- Letting the loop inflate the page. Track length trending DOWN while scores hold or rise.
- Treating convergence as immunity. A 10/10 across passes proves the CURRENT persona set is
satisfied, nothing more; if no slot matches the actual reader's profile, the blind spot ships
with a perfect score. The persona set must span the reader, then convergence means something.
- Applying an auditor's factual "correction" unverified. Personas hallucinate with full
confidence; a counter-claim about a number, API, or doc must be checked against the primary
source before the edit (else the audit inserts the error it exists to catch).
- Trusting your own fix. New facts written during revision got zero persona scrutiny;
fact-check them like first-draft claims.
- Serial gating (the token-and-trust killer). Adding a QA instrument only after the human
catches its absence. Every review the page needs must run in the SAME up-front batch, before
first delivery. If the human ever asks "did you run it past X," X should already have run; the
question means the gate was incomplete, not that the human requested an optional extra.
- Under-investing the first artifact because the ask sounds basic. "Make a beginner page on
a standard tool" is not a research problem; build it to the finished standard the first time.
The gate confirms quality, it does not rescue a lazy first draft, and rescuing one costs more
than doing it right did.
- Ending a fix pass at "all findings applied + greps clean." Patches are local; convergence is
global. After applying findings, a fresh READER-TWIN must re-read the WHOLE artifact (field case:
8 bolted-in glosses + a 49-substitution punctuation purge left a page mechanically compliant on
every rule and materially worse to read; the reader bounced off it the next day). Fixes must be
WOVEN into the prose, not bolted on as parentheticals; mechanical character substitution requires
sentence-level rewriting.
- Not propagating a fix to its neighbors (the dominant regression mechanism). When a fix changes
a load-bearing WORD or CLAIM, it is not done until every place that referenced the old wording is
updated too: the hero/lede, sibling Q&A cards, SVG diagram labels AND their aria-labels, figure
captions, and the footer. A prose fix that leaves a diagram or a neighboring card still asserting the
old thing creates a fresh SELF-CONTRADICTION that reads worse than the original defect. Mechanical
guard: after any patch, grep the page for the OLD word/claim you just replaced, and check every hit.
(Field case 2026-07-07: applying an adversarial panel's fixes dropped THREE teaching pages from ~9 to
6.5 on the post-patch reader-twin, every time because a prose edit did not reach its diagram or its
sibling cards, e.g. hero still said "no compile step" after the body said "compiled to bytecode"; the
SVG still said "shippable" and "21" after the prose said "usable" and "modified Fibonacci". All three
re-verified to 9/10 once every neighbor was propagated. This is the mechanism under the "N fixes sum
to worse" rule: it is usually an un-updated neighbor, not mere density.)
Provenance (the test record)
Method developed and field-proven 2026-06-24 → 2026-07-06 across three topics plus a 9-page
library sweep. The sweep (2026-07-06) is the READER-TWIN's origin story: the page's real reader
stalled cold on a section that three converged generic passes (including two 10/10 fresh noobs)
had blessed; reader-matched twins then found ~75 defects across all 9 pages, two false claims
about the reader's own IDE, one wrong fix introduced BY a fix, and one auditor counter-claim
that was itself wrong (caught by fetching the primary docs). Baseline (RED):
a lesson built without the loop scored 10/10 on a recall quiz while carrying defensibility holes,
undefined load-bearing terms, and one factual error, confirming self-review can't see them.
With the loop: topic 1 converged in 3 passes, topic 2 (skill lessons applied up front) in 1 pass
with zero factual errors in its v1, and the 2026-07-05 targeted 2-agent mode caught a false
audience-home-turf claim plus an honesty trap (a first-person past-tense story card for work not
yet done) that three generic passes had missed. Quiz scores climbed or held at every pass while
page length held roughly flat.
1---2name: iterative-lesson-refinement3description: Use when a teaching or study artifact (HTML page, study pack, explainer, interview-prep doc) must hold up under live questioning, when a lesson "reads fine" but the learner keeps failing follow-up questions, when asked to "battle-test / tighten / refine" a lesson, or when a learner needs to genuinely own a topic fast instead of reciting it.4---56# Iterative Lesson Refinement78## Overview910A QA loop that turns a draft lesson into one that *actually teaches* and that the learner can11*defend live*, by exposing it to **isolated personas** instead of grading your own homework.12Core principle: a lesson's author cannot see its holes, and a persona you role-play inline just13echoes what you already know; isolation is the whole point.1415## When to Use1617- A teaching artifact must survive a skeptical follow-up ("you said cosine, why not Euclidean?"),18 not just read well.19- The learner is time-starved: the artifact must be as SHORT as possible while still letting them20 do the thing and survive the second question. Brevity is a graded criterion, not an afterthought.21- **When NOT to use:** content nobody must defend (internal notes, logs); or a first draft that22 hasn't been written yet (write the tight draft first, then run the loop).2324## The Loop (per topic)25261. **Build/revise** a tight draft. Ground every claim to the learner's real work honestly27 (HAVE / PARTIAL / GAP). If a companion HTML-craft skill exists, use it for the page itself.282. **Dispatch 4 ISOLATED subagents**, each reading ONLY the artifact file: no answer key, no29 authoring context. Locked roles:30 - **NOOB**: technical generalist who has never met the topic; no outside lookups. Reports31 confusion, drag, and density/overwhelm flags; takes the cold quiz. Your density detector.32 - **BASICS**: surface familiarity; mandate = depth gaps, half-explained terms, "asserted33 without justifying." Takes the cold quiz.34 - **EXPERT**: practitioner + interviewer; mandate = (a) ACCURACY (quote → problem → fix),35 (b) DEFENSIBILITY (the exact follow-up that exposes each weak claim), (c) TIGHTNESS36 (a specific cut list). Does NOT take the quiz.37 - **READER-TWIN** (mandatory since 2026-07-06; the slot whose absence let a stall ship):38 derive it per artifact from `references/reader-profile.md` (the durable profile of the39 actual reader: home-stack expertise, zero-knowledge topics, reading traits; update the40 profile as the reader demonstrably learns topics). ⚠️ **The profile rots in BOTH41 directions**: it can lag behind learning (stale-low) AND it can overclaim (a 2026-08-14 run42 found the profile asserting "production experience" for a skill the reader's own source-of-43 truth record bounds to personal projects). When a twin flags a conflict between its profile44 and the artifact, that is the twin working: adjudicate against the PRIMARY record (the45 reader's ground-truth documents, not either derived file), and fix the PROFILE in the same46 pass, or every future twin inherits the error. The real profile is personal and47 gitignored; copy `references/reader-profile.template.md` to create one for your reader. Shape: deep expert in A (their home48 stack: the language, IDE, or domain the page translates FROM or leans on), ZERO knowledge49 of B (the topic being taught). Mandate: reading in order,50 flag (a) any term used before defined (hero copy, diagram labels, and code comments count51 as uses), (b) any literal number without a stated origin, (c) any construct in language A52 they cannot place that is not labeled invented/pseudo, (d) metaphor switches without a53 bridge, (e) read-twice sentences. This slot exists because the other three CANNOT catch a54 fake construct in the reader's home language or a home-turf falsehood: a 9-page field sweep55 found the generic trio had passed pages 10/10 that the reader-twin then found ~75 defects in.563. **Grade the cold quiz** against a rubric (answers in the learner's own words, /10); log scores.57 Synthesize the reports → **verify before applying**: an auditor's factual counter-claim ("the58 page's number is wrong") is testimony, not truth; check it against a primary source before59 editing (field case: an auditor confidently corrected a right number to a wrong one; the doc60 fetch caught it). Then revise, and **fact-check any NEW fact the revision itself introduces**61 (fixes are un-audited authoring; a fix pass once shipped a checkably false origin story for62 the very number it was fixing). Record the *generalizable* teaching lesson (not topic facts)63 back into this skill.644. **Repeat ~2 passes**, escalating quiz difficulty each pass: pass 1 tests the spine (can they65 describe it?), pass 2 tests the defense layer (the second question). Early-exit when a fresh66 NOOB scores ≥8/10 AND the expert flags no accuracy/tightness issues AND a fresh READER-TWIN67 reports zero stalls (no term-before-defined, no unexplained number, no unplaceable construct).685. **Finalize:** bump the artifact's version footer with what changed; clean up test scraps.6970## The delivery gate (non-negotiable)7172**Run the WHOLE gate ONCE, UP FRONT, before the learner ever sees the page. Never serially.**73The most expensive failure mode is not a bad page; it is *reactive gating*: ship, let the human74hit a hole, add one QA instrument, ship again, let them hit the next hole. Each round burns their75tokens, their time, and their trust, and it makes the human the trigger for your next QA step,76which is exactly backwards. Before first delivery of any teaching page, in ONE parallel batch:77(1) build it right, most teaching topics are standard, not novel research, so invest the craft up78front; (2) run the full gate at once, the READER-TWIN plus, for a defend-cold/interview page, the79nemesis + interviewer + accuracy panel; (3) fix what survives verification, and re-read the whole80page once (patches can sum to worse); (4) THEN deliver. The human should never meet a pre-gate81artifact, and should never have to name a missing review instrument. If you find yourself adding a82reviewer *because the human asked why you didn't*, the gate ran too late.8384**The learner is PRODUCTION, not staging.** The loop exists so the struggle happens before they85arrive: they read to LEARN, never to debug. Therefore an artifact may not be handed to the86learner, announced as ready, or registered in any index until the loop, including a clean87READER-TWIN pass, has converged WITHOUT them.8889**Interview-prep / "defend-cold" pages need an ADVERSARIAL pass too, not just the reader-twin.**90The reader-twin is sympathetic; it measures "can I follow this," which catches confusion but91NOT confident-but-wrong claims or unarmed follow-ups. A page whose stated job is surviving a92hostile interviewer must be reviewed by a hostile interviewer. So for any defend-cold page, add93to the gate a `nemesis-review` panel: the nemesis (armed with the audience's expertise) plus an94interviewer-follow-up skeptic ("for each answer, what is my next question, and does the page arm95it?") plus, for technical pages, a domain-accuracy skeptic. Field case 2026-07-07: a Docker page96that passed the reader-twin 9/10 still had three defend-cold answers that each detonated on the97first follow-up (a healthcheck fix that did not fix the race, an EXPOSE claim that ignored98`-P`, and a grounding line that overclaimed a skill the reader had not authored, the exact99trap that had already cost a real final-round interview). Only the adversarial panel caught them. This is why the gate is two instruments,100not one: reader-twin proves it teaches, nemesis proves it does not get the reader killed in the room. If the learner ever hits a defect while studying,101treat it as a build failure with a root cause in this loop (usually a persona slot that didn't102match them): fix the page, fold the lesson back into this skill, and tighten the reader profile.103Do not thank the learner for the QA or frame their stall as a contribution; the moment they are104"helping refine," the artifact has already failed at its one job.105106### Targeted validation mode (cheap variant, added 2026-07-05)107108When REVAMPING an artifact whose spine already converged in earlier passes, do not re-run the109full 3-persona loop. Dispatch **2 agents in parallel**: a fresh NOOB (cold quiz over old + new110content) and an EXPERT whose fire is pointed at the NEW sections. **Give the EXPERT the111artifact's audience expertise explicitly** (their stack, their era), not just the topic domain:112audience-home-turf claims are the highest-risk content and a domain-only expert sails past them.113Field result: this 2-agent pass caught a false claim about the reader's own database engine that114three earlier full passes had missed. When the artifact is a TRANSLATION ("B via the A you115know"), make the second agent a full READER-TWIN (A-expert, zero B) rather than a domain expert116with audience seasoning; only the twin catches invented constructs in language A.117118## Quiz Design (the real gate)119120- Author a per-topic quiz; learners answer **in their own words** (no copy-paste).121- A perfect score on a recall quiz proves nothing: the page teaches *its own claims* well.122 The defense-layer quiz (second-order follow-ups) is what catches recite-vs-own.123- If a learner cannot answer, require them to name WHICH word or idea the page failed to give124 them; that names the fix.125126## The Generalizable Teaching Lessons (the payload)1271281. **Quiz the defense layer, not the page's own claims.** Escalate difficulty each pass.1292. **Define the load-bearing mechanism word on first use.** Beginners stall exactly on the words130 presented AS the machinery (the #1 observed failure mode across every run), not optional jargon.1313. **Never name an interview question you don't arm.** Flagging "be ready to defend X" without132 the material is worse than silence.1334. **A "defend cold" Q&A must carry NEW depth** (second-order follow-ups), never restate the body.1345. **One metaphor per concept**, reused deliberately; three competing metaphors read as padding.1356. **State key numbers once, with their justification**, then reference them.1367. **When you ground a claim to the learner's real work, pre-arm the obvious follow-up.**137 "I ran X on my codebase" → "how exactly?" must be answered ON the page.1388. **Define every borrowed term even inside an advanced section**; name-dropping lets a learner139 recite without owning, the exact thing a follow-up exposes.1409. **Mark the difficulty step-up** ("second-pass layer") so the beginner isn't silently drowned.14110. **Lead with the why before the how:** plain-English claim → analogy → diagram → defend-it Q&A.14211. **A defense answer must be self-contained.** Never let it lean on a term the page never143 defined; that recreates the exact trap the Q&A exists to close.14412. **By the final pass, weight to TIGHTNESS.** Cut any defend-cold card that restates the spine.145 Stop signal: fresh NOOB ~10/10 cold AND expert says "ship after ≤2 one-line fixes."14613. **Check every claim about the reader's OWN home stack against its CURRENT state.**147 Translation sections ("new thing B via the A you already know") date fastest exactly where148 the reader is most expert, and a stale claim there detonates on their home turf. Arm the149 EXPERT persona with the reader's home-stack expertise so it hunts these.15014. **A worked example must not disprove its own lesson, and simulated outputs must be151 watermarked IN the artifact.** Demonstrating semantic search with a question that near-quotes152 its source proves keyword search would have worked too: paraphrase until zero surface tokens153 overlap. Realistic invented scores/terminal output without a "simulated" label become154 "observed data" one lazy memory later; a caption saying "illustrative" is not enough, mark155 the artifacts themselves.15615. **Define-on-first-use is a WHOLE-PAGE ordering constraint, not a per-section virtue.** Hero157 copy, diagram labels, and code comments count as uses; a later deep-dive section does not158 excuse an unglossed first appearance. Walk the page in reading order before shipping. (The159 dominant defect class in a 9-page field sweep, on 8 of the 9 pages.)16016. **Every literal number carries its origin at first use.** What fixes it, whether the reader161 can change it, and an analogy class ("a hash width"); mark example values AS examples ("k = 5162 is this page's example, not a law") and name the corpus behind scale claims. Invented163 precision ("90% of confusion") reads as fake and poisons trust in the real numbers.16417. **Label invented/pseudo-code as pseudo IN the artifact**: before the panel, in the panel165 label, and in a code comment, and name the real construct it stands for. An invented function166 in the reader's home language makes the expert reader conclude they are ignorant or the page167 is lying; both destroy the page.16818. **Home-turf falsehoods are the fastest trust killers.** A wrong claim about the reader's OWN169 tool ("your IDE refuses to start without X") outranks any topic error, because it is the one170 claim the reader can check instantly. The READER-TWIN slot exists to hunt exactly these.17119. **Fact-check the ANSWERS harder than the prose, and hunt them as their own pass.** A wrong172 explanation inside a defend-cold answer is strictly worse than the same error in the body,173 because the format instructs the reader to rehearse it and say it out loud under pressure.174 Prose gets skimmed; an answer gets memorized. Reviewers reading for clarity glide past a175 confidently-wrong answer exactly as fast as a correct one, so accuracy review of the Q&A176 cards cannot ride along with the readability pass — give it its own sweep, and treat every177 answer as a load-bearing factual claim. (Field case 2026-08-10: a hostile reviewer found a178 mechanically wrong protocol-framing explanation that had been written into a closed-book179 answer; the honest replacement was also the better answer, the recurring pattern here.)18020. **Resolve any contradiction the artifact manufactures, on the artifact.** Two claims that are181 each true and look mutually exclusive several sections apart are a trap the page built and182 never disarmed — and when both sit in defend-cold cards, the reader has been instructed to183 say both halves aloud. Sweep a finished artifact for pairs of confident claims a hostile184 reader would put side by side, then reconcile them in place or cut one. Anticipating the185 collision is cheap; meeting it live is not. (Field case 2026-08-10: a rule forbidding a null186 identifier and a case that legitimately sends one; the resolution was one sentence — the two187 rules govern different message directions.)188189**Answer-visibility note:** past ~4 questions in one block, stacked question-above-answer pairs190stop being a test — a reader-twin scored such a section 1/10 and reported skimming it, *"the191answer sitting right under the question... I'm not testing myself, I'm reading a FAQ."* When the192artifact is HTML, `educational-html-prep` carries the collapsed-answer component (`.qz`) and the193"say it out loud first, then open the row" instruction line.19419521. **Verify a CORRECTION at least as hard as a claim.** A pass producing retractions struck a196 TRUE sentence and inserted a FALSE one behind a dated banner. The scrub that caught two other197 invented claims in the same document sailed past that one, because **a paragraph announcing its198 own rigour reads as already-checked.** A verification artifact carries borrowed trust, which is199 exactly what makes it the best hiding place for an error.20022. **The most confidently written section is the least verified one.** Both blocker defects in one201 gated build sat in the *escalation* section - the what-if-they-push-harder material - because it202 is prose rather than code and therefore feels unfalsifiable. It is the opposite: it is the most203 checkable thing on the page and nobody checks it. **Charter one reviewer to check every claim in204 the advanced/what-if section against the code the same page prints.**205206## Subagent Prompt Skeleton207208> You are role-playing [NOOB/BASICS/EXPERT] for a lesson-quality test. Stay in character.209> PERSONA: [knowledge level + mandate; for EXPERT also the artifact's AUDIENCE expertise].210> HARD RULES: read ONLY <file path>; no web/outside knowledge beyond your level; be specific,211> not polite. TASK: [confusion/depth/accuracy log] + [cold quiz answered in your own words].212> RETURN in fixed labeled sections (A)…(E).213214## Files Convention215216Keep a run self-contained in a `LessonLab/` folder: `PROTOCOL.md` (rules of the run),217`quiz_<topic>.md` (questions + rubric), `quiz_log.md` (scores + per-iteration findings + the218generalizable lessons). The log is durable state across the many turns a loop takes.219220## Common Mistakes221222- **Role-playing the personas inline.** They echo the author's knowledge; only isolated223 subagents produce real confusion data.224- **Treating a 10/10 recall quiz as done.** It proves the page states its own content well,225 nothing more; escalate to the defense quiz.226- **Skipping the loop because the draft "reads clean."** Every field run found real defects in227 drafts that read clean, including factual errors.228- **Letting the loop inflate the page.** Track length trending DOWN while scores hold or rise.229- **Treating convergence as immunity.** A 10/10 across passes proves the CURRENT persona set is230 satisfied, nothing more; if no slot matches the actual reader's profile, the blind spot ships231 with a perfect score. The persona set must span the reader, then convergence means something.232- **Applying an auditor's factual "correction" unverified.** Personas hallucinate with full233 confidence; a counter-claim about a number, API, or doc must be checked against the primary234 source before the edit (else the audit inserts the error it exists to catch).235- **Trusting your own fix.** New facts written during revision got zero persona scrutiny;236 fact-check them like first-draft claims.237- **Serial gating (the token-and-trust killer).** Adding a QA instrument only after the human238 catches its absence. Every review the page needs must run in the SAME up-front batch, before239 first delivery. If the human ever asks "did you run it past X," X should already have run; the240 question means the gate was incomplete, not that the human requested an optional extra.241- **Under-investing the first artifact because the ask sounds basic.** "Make a beginner page on242 a standard tool" is not a research problem; build it to the finished standard the first time.243 The gate confirms quality, it does not rescue a lazy first draft, and rescuing one costs more244 than doing it right did.245- **Ending a fix pass at "all findings applied + greps clean."** Patches are local; convergence is246 global. After applying findings, a fresh READER-TWIN must re-read the WHOLE artifact (field case:247 8 bolted-in glosses + a 49-substitution punctuation purge left a page mechanically compliant on248 every rule and materially worse to read; the reader bounced off it the next day). Fixes must be249 WOVEN into the prose, not bolted on as parentheticals; mechanical character substitution requires250 sentence-level rewriting.251- **Not propagating a fix to its neighbors (the dominant regression mechanism).** When a fix changes252 a load-bearing WORD or CLAIM, it is not done until every place that referenced the old wording is253 updated too: the hero/lede, sibling Q&A cards, SVG diagram labels AND their aria-labels, figure254 captions, and the footer. A prose fix that leaves a diagram or a neighboring card still asserting the255 old thing creates a fresh SELF-CONTRADICTION that reads worse than the original defect. Mechanical256 guard: after any patch, grep the page for the OLD word/claim you just replaced, and check every hit.257 (Field case 2026-07-07: applying an adversarial panel's fixes dropped THREE teaching pages from ~9 to258 6.5 on the post-patch reader-twin, every time because a prose edit did not reach its diagram or its259 sibling cards, e.g. hero still said "no compile step" after the body said "compiled to bytecode"; the260 SVG still said "shippable" and "21" after the prose said "usable" and "modified Fibonacci". All three261 re-verified to 9/10 once every neighbor was propagated. This is the mechanism under the "N fixes sum262 to worse" rule: it is usually an un-updated neighbor, not mere density.)263264## Provenance (the test record)265266Method developed and field-proven 2026-06-24 → 2026-07-06 across three topics plus a 9-page267library sweep. The sweep (2026-07-06) is the READER-TWIN's origin story: the page's real reader268stalled cold on a section that three converged generic passes (including two 10/10 fresh noobs)269had blessed; reader-matched twins then found ~75 defects across all 9 pages, two false claims270about the reader's own IDE, one wrong fix introduced BY a fix, and one auditor counter-claim271that was itself wrong (caught by fetching the primary docs). Baseline (RED):272a lesson built without the loop scored 10/10 on a recall quiz while carrying defensibility holes,273undefined load-bearing terms, and one factual error, confirming self-review can't see them.274With the loop: topic 1 converged in 3 passes, topic 2 (skill lessons applied up front) in 1 pass275with zero factual errors in its v1, and the 2026-07-05 targeted 2-agent mode caught a false276audience-home-turf claim plus an honesty trap (a first-person past-tense story card for work not277yet done) that three generic passes had missed. Quiz scores climbed or held at every pass while278page length held roughly flat.