AI writing audit
What this returns, and what it will not claim
This skill answers one question: does this text read as machine written, and where.
It does not answer whether a machine wrote it. That is a different question and the evidence says nobody answers it reliably. The source guide records a 2025 study finding human accuracy at chance level, a second study at 57 percent on AI text, and a preprint putting heavy model users at about 90 percent, which still means one false accusation in ten. Detector tools are described in the same source as having non-trivial error rates and as breaking under paraphrase, markup changes, or an unfamiliar model.
So the verdict is about reading, not authorship. Never write "this was AI generated". Write "these seven patterns read as machine written, here is the repair for each".
The two layers
Surface layer. Word choice, sentence shape, punctuation, formatting. Catalogued in
references/surface-tells.md, built from Wikipedia's "Signs of AI writing" field guide.
Cheap to detect and cheap to fix, which is also its weakness: these signatures get removed
by each new model generation and by a single editing pass.
Discourse layer. Theme, causality, time, agency, stance. Catalogued in
references/discourse-tells.md, built from StoryScope (COLM 2026), which measured 61,608
stories and reached 93.2 macro-F1 on human versus machine using narrative structure alone,
with zero style signals. Its conclusion states the reason this layer exists: surface
signatures are transient and post-editable, while narrative features require structural
rewrites to change. The measured rate, weight, and repair for each discourse check are in
"Research-backed tells" below, and that section also carries the two procedure rules the
same paper measured.
Measured together in that paper: a style-only model scored 85.8 and a 30 feature narrative model scored 84.8. Within a point of each other. Neither layer is sufficient, so run both, and never report a pass from one layer alone.
The gate. references/false-positives.md holds the patterns that look like tells and
are not, the constructions that lean human, and the confidence each format deserves. Every
finding passes through it before it reaches the output.
The sources disagree with each other and with the tools built on them.
references/conflicts.md lists every disagreement and the ruling this skill applies. Four
of those rulings are load bearing and are repeated here, because getting them wrong changes
the output rather than refining it.
Two modes, and they are not the same question
Deliverable mode, the default. The text belongs to the person asking. The goal is a better text. A single hit is worth fixing when the fix costs nothing, no cluster required, and no claim about authorship is made or needed.
Assessment mode. The text came from somewhere else and the question is how it reads. A
single hit supports nothing. Spread is the bar: findings touching four or more of the seven
measured themes in "Research-backed tells" below, before the reading verdict means much. The
cautions in references/false-positives.md about false accusations apply in full, and so
does the measured floor: on 5,000 word fiction, with a purpose-built classifier, about one
human text in nine still came back as machine written.
Say which mode ran. The same text can pass one and fail the other without either being wrong.
The rule that outranks every other rule here
Never add anything to a text that was not true of it.
No invented first person. No anecdote that did not happen. No opinion the writer did not hold. No source, statistic, name, or number that was not already there or supplied by the writer.
Some cleanup tools, when given no writing sample, fall back on manufacturing a voice: add uncertainty, add a personal aside, add an admission. That produces text that reads more human and states things that are false, which in marketing copy is a false claim about a real business.
When a draft is clean, flat, and voiceless, report it as a finding. Ask for the missing specific. Do not fill the hole yourself. Voice can be matched from a sample, and it cannot be invented from nothing.
Optional input: a voice sample
If the writer supplies two or three samples of their own past writing, use them. Sample evidence outranks every generic list in this skill.
Measure the samples before auditing: sentence length spread, dash and punctuation habits, paragraph length, opening moves, recurring phrases, whether they use contractions, fragments, emoji, first person. Then build a protect list from the result.
The reason this matters: a real writer's tics look like machine filler to a generic pass. Spoken glue, a repeated stock phrase, an unusual dash habit, an "okay?" every few lines. A generic sweep strips exactly the things that make the text theirs. Anything on the protect list is weight 0 for that writer, and stripping it is a finding against the audit rather than against the text.
Without samples, run the generic lists and say in the verdict that no sample was available.
Run it
Pass 0. Profile the text
Name the format. Open references/formats.md and load its profile. If the format is not
listed and the audit is a one-off, answer the four questions in section 9 of that file and
write the derived profile into the output so the reader can argue with it. If the format is
one the writer works in and will bring back, run "Adapting to a new format" below instead,
which produces a profile with a silence map and registers it.
One piece can hold two formats. A short video ships as a spoken script plus a written caption, and those get different passes: the script gets the spoken profile with the word list demoted, the caption gets the full written pass. Audit them separately and say so.
Check the language. The discourse checks and the formatting checks work in any language. The vocabulary lists were measured on English tokens and do not survive translation. For text in another language, run both other layers, report the vocabulary layer as not run, and build a local list only from observed output in that language rather than by translating this one.
Ask for provenance if it is unknown and cheap to get. If the text predates 30 November 2022, machine authorship is ruled out by date and the audit becomes a writing review only.
Count the words. The count sets which layer leads:
| Words | Layer that leads | Discourse checks that run |
|---|---|---|
| Under 60 | Surface | Checks 1, 5 and 6 only: self-explaining, moral polarity, named reference |
| 60 to 300 | Both, equal | The format profile's list. Check 10 needs two facts in sequence and often stays silent here, which is a silence and not a pass |
| 300 to 800 | Both, discourse rising | All ten, full carry-over table |
| Over 800 | Discourse | All ten, full carry-over table, surface as supporting evidence |
Below 60 words the seven theme bar cannot be met, because three checks cannot touch four themes. So assessment mode has no verdict to give on a short text, and says that instead of lowering the bar. Deliverable mode still runs and still repairs.
Pass 1. Surface sweep
Go through references/surface-tells.md and mark hits with exact quotes and positions.
Mark, do not fix yet. Fixing during the sweep hides the density, and density is the finding.
Weight each hit:
- Weight 2, hard. Documented, specific, and rare in unedited human writing: the flagged vocabulary items, negative parallelism, vague attribution, participial pseudo-analysis, promotional register with no fact under it, self-summarizing conclusions, formatting artifacts from a chat interface.
- Weight 1, weak. Real but common in human writing: dash density, the rule of three, a symmetrical pair, boldface as texture, one transition word.
- Weight 0, never counted. The ineffective indicator list in the gate file. Perfect grammar, formal prose, bland prose, mixed register, a transition word on its own. Do not raise these as findings, in any format, at any length.
Pass 2. Discourse sweep
Before the checks, build the structural template. The paper measured that inducing features from raw prose and inducing them from a structured template of the same text return different feature sets, only 6 of the top 20 overlapping, and that the prose route returns style features while the template route returns structural ones. Fill the template in "Research-backed tells" below, mark every absent field null, then run the checks against the template. Quote the prose only to evidence a finding.
Run each check as its own pass. Applying the paper's features in one call covered 68.4 percent of them against 95.4 percent one dimension at a time, and the loss concentrated in revelation and temporal structure, which is where checks 8, 9 and 10 live. A single sweep under-detects the human-leaning signals specifically.
Run the carry-over table in references/discourse-tells.md at the depth Pass 0 set. Ten
checks in full, weights from "Research-backed tells" below:
- Self-explaining. Does the text state its own lesson, moral, or significance rather than leave it to the reader.
- Single track. Does everything pull one way, with no aside, exception, or second thread.
- Tidy causality. Does each sentence follow cleanly from the last, with no gap, reversal, or admitted cost.
- Internal resolution. Does it end by asking the reader to decide, realize, or commit, rather than to do something external and small.
- Moral polarity. Is every actor clearly right or clearly wrong, with nothing ambivalent.
- Named reference. Are there real names, numbers, dates, places, tools. Or only diffuse echoes of them.
- Emotion handling. Are feelings delivered as body sensations by default. The paper found embodied emotional expression is the machine-leaning value and an explicit label is the human-leaning one, which reverses the usual "show, do not tell" instruction. Judge the default, not one instance.
- Order. Is the text told in the flattest possible order, when the material had another one available.
- Reader address. Does the text know it is being read. Machine text writes as though no one is watching. The check has two halves and they behave differently, so run them apart. The second person half is silent wherever second person is the native register: a message, caption, ad, script, email or landing page, where its presence and its absence both carry nothing. Score it only in a format whose default register is third person. Profile 12, product description, is the first such profile and case 16 is the first run where this half fired. Question 1 of "Adapting to a new format" is what decides it, so any derived third person format turns it live. The medium naming half does fire here: "this is a cold message", "last one from me", "ignore this if it is not live". It fires in the human direction, so its absence is a note and never a fault, and its presence belongs in WHAT IS WORKING.
- Recontextualization. Does any later line change what an earlier line meant. Machine text discloses in the order it was assembled, so nothing arrives that forces a re-reading. Absence is common in human writing too, so this one is evidence and never a verdict.
Every discourse finding carries the theme it belongs to, because Pass 3 scores spread and spread is not reproducible without this mapping. The themes are numbered as they appear in "Research-backed tells" below.
| Check | Theme |
|---|---|
| 1. Self-explaining | 1, thematic over-determination |
| 2. Single track | 1 when the complaint is thematic unity, 3 when it is the missing second thread |
| 3. Tidy causality | 3, structural streamlining |
| 4. Internal resolution | 3, structural streamlining |
| 5. Moral polarity | 7, narrative diversity |
| 6. Named reference | 4, intertextual richness |
| 7. Emotion handling | 2 when feeling arrives as a body sensation, 7 when no feeling is named at all |
| 8. Order | 6, temporal complexity |
| 9. Reader address | 5, reader engagement |
| 10. Recontextualization | 6, temporal complexity |
Three consequences worth reading off the table. Ten checks cover seven themes, so themes 1, 3, 6 and 7 can each be reached by more than one check and a second hit inside a theme adds density without adding spread. The three checks eligible below 60 words reach themes 1, 4 and 7 only, which is the arithmetic behind the rule in Pass 0 that assessment mode has no verdict at that length. And theme 5 is reachable only through check 9, whose scored half is silent in every second person format, so in those formats the bar is four themes drawn from six rather than seven. In a third person format, derived or listed, theme 5 opens and the bar is drawn from all seven. Case 16 is the recorded instance.
Do not run the fiction-only features on non-fiction. references/discourse-tells.md lists
which ones stay in fiction and why porting them produces nonsense.
Pass 3. Gate
For every finding, three questions:
- Is it on the ineffective list. If yes, delete the finding.
- Does the format profile call it native. If yes, delete the finding.
- Is it a construction the source records as leaning human. Plain verbs, "there is a", "in order to", "the fact that", "very", "perhaps", a superlative. If yes, delete the finding, and if the draft is thin on these, note it as a repair opportunity rather than a fault. Sanding these off makes the text more machine-like, not less.
- For a discourse finding, is the human rate for that feature already high. Three of the highest ranked machine-leaning features sit close to the human value: thematic unity 4.41 against 4.74, causal continuity 3.92 against 4.20, moral weighting 3.26 against 3.68, all on a 1 to 5 scale. Keep the finding, label it weak, and never let one of these carry a verdict on its own. Compare with emotion carried by the body, 38 against 81, which is a real separation.
Then score. Density is weighted findings per 100 words:
| Density | Verdict |
|---|---|
| Under 1.0 | Reads human |
| 1.0 to 3.0 | Mixed |
| Over 3.0 | Reads machine |
For text under 60 words the denominator is unstable, so score by count: 0 to 1 weighted findings reads human, 2 to 3 mixed, 4 and above reads machine.
In assessment mode, density is not enough and spread decides. The paper groups its 30 core features into seven themes, three machine-leaning and four human-leaning. Count how many of those seven the findings touch. Hits in four or more, with at least one weight 2 hit in each, before a reading verdict means anything about a text you did not write. A high density inside one theme is a writing problem, not a reading verdict. The seven themes are the paper's. The number four is this skill's.
In deliverable mode the bar does not apply and spread is reported without gating anything. The text belongs to the person asking, no origin claim is being made, and a single weight 2 hit is worth repairing whether or not six other themes are clean. Print the spread line anyway, because it tells the writer whether they have one habit or several.
These cutoffs are a judgment call layered on top of descriptive sources, not a measured threshold. The source page says of itself that it is descriptive, not prescriptive, a list of observations rather than rules. Say so when the score is close to a boundary.
Pass 4. Repair
The rule that governs every repair, from the source guide: the patterns are potential signs of a problem, not the problem itself, and treating the signs as the thing to fix "could just make detection harder".
So, in order:
- Cut first. Most findings are sentences doing no work. Deleting is the repair. Do not replace a hollow sentence with a better hollow sentence.
- If it stays, put something under it. A flagged promotional line needs a fact, a number, a name, or a cut. A synonym pass leaves the text empty and clean, which is the failure mode this skill exists to prevent.
- Repair the discourse findings before the surface findings. Surface repairs on a single-track text produce a polished single-track text. The reverse order wastes work, because restructuring rewrites the sentences anyway.
- Change one thing per finding. Keep the writer's voice, the argument, the offer, and the facts. This skill has no mandate to change what the text says.
- Do not add a tell while removing one. The common accident is replacing a banned word with a rhetorical flourish, replacing a dash with a colon everywhere, or replacing a summary line with a rhetorical question.
- Leave the roughness. If the draft has a plain verb, an abrupt sentence, a small digression, or an unbalanced rhythm, that is the human signal. Protect it.
- Never insert rhythm the draft did not have. Varying sentence length is a repair for a metronome cadence that is already present. It is not a style to apply on top. Do not add a fragment, an aside, or a one-line paragraph that this writer's voice did not already contain. A text that reads as a machine trying not to read as a machine has traded one pattern for a worse one.
- Add nothing that was not true. See the rule above the passes. This is where it gets broken, and it is the one repair failure this skill treats as disqualifying.
Pass 5. Report
MODE: deliverable | assessment
FORMAT: <name> (<derived / from profile list>)
LENGTH: <n> words LANGUAGE: <name> (vocabulary layer: run / not run)
VOICE SAMPLE: yes, <n> samples | none supplied
CONFIDENCE: <from the confidence table in false-positives.md>
MODEL: <named, if the draft's author is known> | unknown
VERDICT: reads human | mixed | reads machine
(density <x.x> per 100 words, spread <n> of 7 themes)
STRUCTURAL TEMPLATE, filled before the discourse checks
agents: <who acts, or null>
events: <what happens, or null>
causality: <the chain, or null>
revelation: <what is withheld and when it lands, or null>
temporal order: <linear, nonlinear, mixed>
setting: <where, or null>
FINDINGS
1. [surface, w2] "<exact quote>"
Tell: <name from the catalogue>
Why: <one line>
Repair: "<replacement, or CUT>"
2. [discourse, w2, theme <n>] <check number and name>
Evidence: <quote or structural description>
Repair: <what to restructure>
3. [discourse, w1, theme <n>, WEAK] <check number and name>
Weak because the human mean on this feature is already high. Cannot carry a verdict.
...
GATED (looked like findings, are not)
- "<quote>" : <which gate rule cleared it>
SILENCE MAP (derived profiles only, all ten rows, see "Adapting to a new format")
<check> live | silent : <reason, and which of the six questions or the length rule>
THEME SPREAD: <which of the seven, and whether each carries a weight 2 hit>
WHAT IS WORKING
- <the human-leaning constructions present, so the next edit does not remove them>
- <every discourse check that fired in the human direction, named by number>
REWRITE
<the repaired text in full, if a rewrite was asked for>
Always print the GATED section, even when empty. It is what stops the audit turning into a machine that finds seven problems in every text regardless of the text.
Print WHAT IS WORKING before the rewrite. An audit that only subtracts trains the next draft toward the safe middle, and the safe middle is where the measured machine cluster sits.
Adapting to a new format
Someone says "I use this for support replies" or "adapt this for our CRM notes" or names a niche, a channel, or an automation this file has never heard of. This section is how you answer, without inventing anything.
What adapts: the length band, which layers run, which catalogue entries are hard flags here, which patterns are native here, the pass bar, and the silence map.
What does not: the ten checks, the surface catalogue and its weights, the four gate questions, the density bands, the seven themes and the bar of four.
A format does not get its own rules. It gets its own answer to which of the fixed rules can fire in it, and why the rest cannot. That is the difference between adapting a tool and loosening one.
Start from the nearest listed profile
Open references/formats.md. If one of the listed profiles is close, take its hard flags
and its native list as the starting draft and change only what the six questions below
change.
A support reply is an objection reply that the reader asked for. A CRM note is a long page
with no reader. Borrowing is cheaper and more accountable than deriving from nothing.
Length and channel need no questions. Length sets eligible checks from the Pass 0 table. Heard means the spoken override in profile 6 applies in full. Seen means the deck-level audit in profile 5. Written once for many readers means the scale pass in profile 2 runs.
Six questions
Each one changes exactly one check. Answer from the user's declaration, not from a guess.
| Question | Yes | No |
|---|---|---|
| Is the default grammatical person third? | Check 9 is live, and theme 5 becomes reachable as a finding | Check 9's scored half is silent, second person is native and carries nothing |
| Did the reader ask for this text? | An explicit takeaway is native. Check 1 fires only on a takeaway nobody asked for | Check 1 is a hard flag |
| Is the structure itself the product? | Check 3 is silent, and headings, bullets and inline-header lists are native | Both stay live |
| Could the writer have known something specific about this reader or subject? | Check 6 runs at full weight | Check 6 runs at weight 1 and measures the format's ceiling rather than the writer's effort. Say so in the report |
| Is any feeling in scope? | Check 7 is live | Check 7 is silent |
| Is the order fixed by convention? | Check 8 is silent | Check 8 is live |
If an answer is not in the declaration, ask that one question. Do not pick a value to keep moving. A guessed answer produces a profile that looks derived and is not.
If two answers point opposite ways on the same check, the silencing one wins and the conflict goes in the silence map rather than being resolved quietly.
The silence map, required
Print all ten rows. A silent check is a reported result, not an omission, which is the same rule the audit already applies to a check that finds nothing.
SILENCE MAP
1 self-explaining live | silent <why, and which question or the length rule>
2 single track ...
3 tidy causality ...
4 internal resolution ...
5 moral polarity ...
6 named reference ...
7 emotion handling ...
8 order ...
9 reader address ...
10 recontextualization ...
A row reading silent with no reason is inadmissible: it means the agent decided rather than derived. Below 60 words, seven rows read silent by length, and that is the reason to write.
Three rules that stop invention
- Every hard flag names its source: an entry in
references/surface-tells.md, a check number, or a scale-pass bullet from profile 2. A hard flag citing nothing is struck. - The ten checks may only be marked live or silent. Not added to, removed, renamed, merged or reordered. The four gate questions run unchanged. Weights, bands, themes and the bar are untouched, with the single exception written into question four above.
- A native entry either appears in the format-native table in
references/false-positives.mdor follows from one of the six answers, and the profile says which.
Register it
Print the derived profile and its silence map ahead of the findings, so the reader can argue with the profile rather than only with the findings.
If the format will be audited again, add it to references/formats.md as the next numbered
profile and add a worked example to TESTS.md, in the same change. A profile with no
recorded run is a claim the test log does not support.
Adding a profile is not a contract change, so recorded cases stay valid. If a derived profile ever forces a new output field, that is a contract change, and every recorded case has to be rechecked for that field in the same change.
Research-backed tells (arXiv 2604.03136)
Source: Russell, Rajendhran, Pham, Iyyer, Wieting. "StoryScope: Investigating idiosyncrasies in AI fiction." COLM 2026, arXiv 2604.03136v6. Page numbers below are the printed page of that PDF.
This section does two things the rest of the skill does not. It puts a measured number and a weight on each discourse check, taken from the paper's own ranking rather than from taste. And it fixes two procedure defects that the paper measured directly.
Nothing here claims a text was machine written. Every rate below is a rate, and the human column is never zero.
How weight is set here
The paper ranks its 30 core features by a core score, mean SHAP multiplied by a stability score and by one plus the absolute human-AI gap (Appendix I, p.23). That ranking is the paper's. The mapping to weights is this skill's:
- Weight 2. Top 10 of the AI-characterizing list (Table 14, p.24) or top 6 of the human-characterizing list (Table 15, p.25).
- Weight 1. Any other member of the 30.
- Weight 0. Anything on the fiction-only list in
discourse-tells.md, whatever its rank.
A human-leaning feature is never a fault. Its absence is a repair opportunity, and stripping it is a finding against the audit.
Read the human column, not only the gap
Three of the highest ranked AI-elevated features sit high for humans too: Thematic Unity 4.41 human against 4.74 AI, Causal Chain Continuity 3.92 against 4.20, Moral and Philosophical Weighting 3.26 against 3.68, all on a 1 to 5 scale (Table 16, p.26). A hit on one of these is weak evidence, because human writing does close to the same thing. Compare that with emotion carried by the body, 38 percent human against 81 percent AI, which is the widest separation in the whole table. Say which kind of hit you have.
Two procedure changes the paper forces
Run the discourse checks one dimension at a time. The paper compared applying its features in a single call against applying them one narrative dimension per call. Coverage rose from 68.4 percent of features to 95.4 percent, and the single-call dropout was concentrated in revelation and temporal structure (Appendix C, p.18). Those are the two dimensions carrying most of the human-leaning checks, so a single-sweep audit under-detects exactly the signals that argue a text is human. Run the ten checks as ten passes.
Abstract the text before running the discourse pass. The paper compared inducing features from raw prose against inducing them from a structured template of the same story. Of the top 20 discriminative features, only 6 overlapped: the raw-prose variant produced style-heavy features (humor, vocabulary register, allusion types, dominant imagery), the template variant produced structure-heavy ones (emotional arcs, relationship trajectories, event density, flashback usage). Appendix B, p.17. Auditing the prose directly returns surface findings wearing a structural label.
The paper prints the template it used, so the summary step does not have to be improvised. Figure 8, p.28 to p.30, fields condensed:
| Group | Fields |
|---|---|
| agents | major characters with role, attributes, emotion trajectory, motivation trajectory, trope. Supporting characters with a one-line role |
| social network | enduring bonds, as A-B: relationship type and quality |
| events | ordered beats with who, where, what, when. Causal links as event1 -> event2: explanation. Narrative schema |
| plot | themes, summary, moral, central obstacle, central conflict, archetype, plot arc |
| setting | locations with scope, time period, atmosphere |
| revelation | what is withheld for suspense, what causal antecedents are withheld, what was revealed and when |
| temporal order | linear, nonlinear or mixed. Duration, flashbacks, time jumps, scene durations |
| perspective | point of view, focalization by section, who speaks |
| style | allusions, figurative language, imagery, sentence complexity, evaluative language |
Three instructions travel with it and matter as much as the fields. Stay inside what the
text conveys and do not interpret past it. Write null for anything not present. Mark each
field as story-level or section-level.
The null rule is the one to keep for non-fiction. A draft that returns null for revelation,
for causal links, for supporting agents and for time jumps has told you what the discourse
pass is about to find, before the pass runs. Fill the template first, then audit the
template, then quote the prose only to evidence a finding.
AI-elevated: thematic over-determination
Table 16 group 1, p.26. Section 4.1, p.7.
| Test, answerable yes or no | Human | AI | Weight | Repair |
|---|---|---|---|---|
| Does a sentence state the text's own point, lesson, or significance | 3.28 | 3.94 | 2 | Cut it. If the point does not survive the cut, it was not in the text |
| Does the writer step outside the material to comment on what it means | 52% | 77% | 2 | Cut, or replace with the fact that would let a reader conclude it |
| Does everything pull one way, with no aside and no exception | 4.41 | 4.74 | 2 | Restore the detail that was cut for being off-point |
| Is a moral or philosophical question foregrounded | 3.26 | 3.68 | 1 | Demote it under the concrete case |
| Do quoted words exist mainly to argue a position | 34% | 59% | 1 | Give the speaker something to do besides hold a view |
| Do references stay vague allusions rather than named things | 50% | 72% | 1 | Name the source, or cut the gesture |
Scale rows are means on 1 to 5. Percentage rows are the share of stories carrying that value.
AI-elevated: sensory and embodied performativity
Table 16 group 2, p.26. Section 4.1, p.7.
| Test, answerable yes or no | Human | AI | Weight | Repair |
|---|---|---|---|---|
| Is feeling delivered by default as a body sensation or bodily metaphor | 38% | 81% | 2 | Name the feeling and move on |
| Is the physical setting doing the work of the inner state | 3.58 | 4.07 | 2, fiction only | Not applied outside fiction |
| Is sensory description dense throughout | 3.66 | 3.93 | 2, fiction only | Not applied outside fiction |
| Is smell used as imagery | 57% | 82% | 1, fiction only | Not applied outside fiction |
The first row is the widest gap in the paper, 42 points, and it is the measured reversal of
"show, do not tell" already ruled on in conflicts.md item 9. This section raises the
confidence on that ruling rather than changing it. Verbatim, p.7:
Where a human author might write that a character "felt afraid," AI renders fear as a tightening chest, cold sweat, and dimming lamplight.
Judge the default across the text, not one instance.
AI-elevated: structural streamlining
Table 16 group 3, p.26. Section 4.1, p.7.
| Test, answerable yes or no | Human | AI | Weight | Repair |
|---|---|---|---|---|
| Does every sentence follow cleanly from the last, no gap, no reversal | 3.92 | 4.20 | 2 | Admit the cost, the exception, or the step that failed |
| Does the close turn on the reader's own choice or will | 46% | 69% | 2 | Ask for an external act, small and checkable |
| Does the text run one thread with no second thread at all | 57% | 79% | 1 | Restore the second thread if one existed |
| Does it end in understanding or acceptance rather than an action | 27% | 47% | 1 | End on the thing to do, not the thing to realize |
| Is the central figure introduced by outside description | 30% | 52% | 1, fiction only | Not applied outside fiction |
Human-elevated: intertextual richness
Table 16 group 4, p.26. Section 4.1, p.7. Absence is the finding, never presence.
| Test | Human | AI | Weight | If absent |
|---|---|---|---|---|
| Does the text name a specific text, author, tool, brand, place, or number | 47% | 24% | 2 | Ask the writer for the specific. Do not supply one |
| Does it mix named references with looser ones rather than only gesturing | 37% | 16% | 2 | Same |
The paper's wording, p.7: AI "generally sticks to vague allusions and avoids naming real brands, places, or works."
Human-elevated: reader engagement
Table 16 group 5, p.26. Section 4.1, p.7.
| Test | Human | AI | Weight | If absent |
|---|---|---|---|---|
| Does the text name its own medium or situation | 67% | 39% | 2, positive only | Add nothing. Record as a flat register |
| Does it address the reader as a party who is present | 28% | 7% | inert here, see below | Not scored |
This is the group that does not survive the transfer intact, and check 9 in Pass 2 is split for that reason. The paper measured both rows on fiction, where third person is the default and turning to the reader is rare enough to mean something. In a message, caption, ad, script, email or product page, second person is the native register, so its presence and its absence both carry nothing. Eight of the nine originally listed profiles have second person as their default, which is why this row went unscored for the first fourteen runs. Profile 12, product description, is third person, and case 16 is the run where the row finally fired. Question 1 of "Adapting to a new format" is the switch.
The first row does transfer, because naming the medium is a choice in any format: "this is a cold message", "last one from me", "ignore this if it is not live". It runs in the human direction only. Its presence goes in WHAT IS WORKING and its absence is a note, never a fault. In a second person format theme 5 is therefore reachable only as a positive and the practical spread ceiling is six. In a third person format both halves run and the ceiling is seven.
The paper's line, p.7: "AI writes as though no one is watching."
Human-elevated: temporal complexity
Table 16 group 6, p.26. Section 4.1, p.7.
| Test | Human | AI | Weight | If absent |
|---|---|---|---|---|
| Does a later line force a re-reading of an earlier one | 3.28 | 2.95 | 2 | Note the missing turn. Do not invent one |
| Does the text jump across time | 2.40 | 2.12 | 1 | Note flat order |
| Is any disclosure staged out of order | 1.96 | 1.68 | 1 | Same |
| Does it use a flashback or a flash-forward | 2.58 | 2.31 | 1 | Same |
The first row is new to this skill. It was in the paper's human core list (Table 15 row 4, p.25) and had no matching check in Pass 2. Non-fiction form: a later fact that changes what an earlier sentence meant. All four means sit in the low middle of the scale for both sources, so absence here is common in human writing too.
Human-elevated: narrative diversity
Table 16 group 7, p.26. Section 4.1, p.8.
| Test | Human | AI | Weight | If absent |
|---|---|---|---|---|
| Is the central figure allowed to be partly wrong | 59% | 38% | 1 | Ask for the cost or the case where it failed |
| Is any feeling named plainly rather than performed | 29% | 8% | 1 | Pair with the embodied row above, same feature |
| Where a second thread exists, does it run parallel to the main one | 42% | 21% | 1 | Note it |
| Does the text move across more than one setting | 1.34 | 1.08 | 1, fiction only | Not applied outside fiction |
The third row corrects a reading that is easy to get backwards. The AI tell is having no second thread at all, 79 percent against 57 percent. Among texts that do have one, running it parallel to the main line is the human-leaning value, 42 percent against 21 percent. So "everything connects" is not by itself the tell. "There is only one thing" is.
Scoring by cluster
The paper groups its 30 core features into seven themes, three AI-elevated and four human-elevated (Table 16, p.26). Those seven are the categories to count in assessment mode.
This replaces an invented number. conflicts.md item 6 sets the assessment bar at five or
more distinct categories and states plainly that the number is this skill's, not a source's.
The seven themes above are a measured grouping, so the bar can be restated against them:
hits spread across four or more of the seven themes, with at least one weight 2 hit in each,
before a reading verdict means anything about an unfamiliar text. The number four is still
this skill's judgment. The seven themes are not.
What this paper refuses to support
One check is never enough, inside the source domain or outside it. Trained on a single NarraBench dimension, the best model reached 80.2 percent binary macro-F1, and removing any single dimension cost at most 1.2 points (Appendix E, Table 8, p.21). The signal is redundant and spread out. A verdict resting on one check is not supported.
A structural audit still misreads about one human text in nine. In the narrative-only six-way model, genuinely human stories were classified as human 88.5 percent of the time (Figure 3, p.9), and per-class human F1 was 0.89 without style and 0.93 with it (Table 12, p.22). That is a purpose-built classifier on 5,000 word fiction, which is the best case. Read it as a floor on the error rate of any reading verdict, and carry it into the confidence line.
Oddity raises the odds and cannot clear a single text. Human stories are rarer in narrative space, mean rarity percentile 0.71 against 0.49, Cohen's d 0.83, AUC 0.73. But at the prompt level the human version was the rarest of the six only 57.8 percent of the time, and in raw counts AI stories fill more of the rare tail than human ones: top 1 percent, 42 human against 41 AI; top 5 percent, 180 against 234; top 10 percent, 340 against 487 (Appendix H, Table 13, p.22 to p.23). The paper's own summary of Figure 5 is that "all distributions overlap substantially." So the closing question of an audit, is there anything here another draft would not have had, stays useful as a prompt and does not become a test.
Length and topic are not the signal, which is a result in this skill's favour. A classifier on word count alone reached 55.9 macro-F1, and the narrative model held 93.2 before and after length matching, and 91.6, 94.3 and 93.7 across short, medium and long bands (Appendix G, Table 11, p.22). Across six topics, Kruskal-Wallis H = 4.69, p = 0.46. Do not discount a finding because the text is short, and do not add one because it is long.
Surface repair does not move this layer. Span-level rewriting of seven artifact categories over 278 stories left detection at 93.9 macro-F1 against 95.5 unedited, a drop of 1.6 (Section 4.2, p.8). Structural findings are repaired by structural rewrites only, which is the reason Pass 4 orders discourse repairs first.
Model-conditional tells, use only when the model is known
The paper's fingerprint features are per-source and are not usable when the author is
unknown, which is the ruling discourse-tells.md already carries. They become usable in the
one case where the model is known, which is a draft this assistant just produced.
Claude is the most distinctive of the five models, 77.1 percent per-class F1 on narrative features alone against 55.0 to 73.0 for the others (Table 12, p.22), and holds 26 fingerprint features against 3 for the least distinctive model (Appendix I, p.25). Its top-ranked fingerprints, by uniqueness ratio (Table 17, p.27), and the paper's summary at p.9:
- Event escalation is flatter than any other source, uniqueness 22.4, the highest in the table. Test: does intensity stay level from open to close.
- Event-type variety is low, uniqueness 10.7. Test: do the beats repeat one kind of move.
- Endings reach forward, epilogue or flash-forward. Test: does the last beat step outside the main span to report what happened later.
- Dream or vision as a break in time is avoided.
- Verbatim, p.9: Claude "takes a reverent/continui
…(truncated)