Evidence Grading Skill
"We have evidence" covers everything from a randomized experiment to the CEO's seatmate on a flight — and decisions made on ungraded evidence inherit the confusion. The grading discipline: inventory what actually supports the claim, place each item on the business-evidence hierarchy (what it is, not how confident it feels), test the graded weight against the decision's stakes (a reversible pilot needs less than a one-way rebrand — decision-journal logic), and when evidence falls short, name the cheapest upgrade — because the answer to weak evidence is usually a better test, not a braver bet.
What This Skill Produces
- The evidence inventory — everything supporting (and contradicting) the claim, each item graded on the hierarchy
- The weight assessment — what the graded pile actually supports, including the mixed-signals honest read
- The sufficiency verdict — enough for this decision's stakes / not yet — with the stakes analysis shown
- The upgrade path — the cheapest next evidence that would move the verdict ("a 2-week holdout test settles this for $0")
Required Inputs
Ask for these if not provided:
- The claim and the decision riding on it — "users want X" feeding a backlog item vs. feeding a repositioning are different sufficiency bars; the decision's reversibility and cost set the bar
- The evidence, itemized — every piece: the data pull, the survey, the five customer quotes, the competitor's move, the expert's opinion — including the inconvenient items (an inventory that omits contradicting evidence is advocacy)
- The evidence's provenance — n, selection, dates, who collected it and with what incentive (source-triangulation supplies the externals; internal evidence has incentives too)
Framework: The Grading Rules
- The hierarchy, applied without sentiment: experiments/A-B tests (causal) > usage/behavioral data (what people do, correlational) > surveys (what they say, at scale — survey-design-basics quality adjusts the grade) > interviews (rich, small-n — interview-synthesis counts matter) > anecdotes (existence proofs only) > expert opinion (informed priors) > internal conviction (a hypothesis, not evidence). Each item gets its rung and its quality-within-rung — a leading survey grades below honest interviews.
- Direction and independence both count: the inventory includes contradicting evidence at its own grade, and echo-checks the supporting pile (three anecdotes traceable to one loud customer are one anecdote). The weight is the net, honestly netted.
- Anecdotes prove existence, never prevalence: "a customer asked for X" establishes that the need exists somewhere — it cannot establish how common; the classic grading error is prevalence conclusions from existence evidence, and it's the error this skill most often catches.
- Sufficiency is relative to stakes: the verdict tests weight against the decision — reversible-and-cheap decisions legitimately run on interview-grade evidence (the pilot is the experiment); irreversible-and-expensive ones demand behavioral or experimental grade. "Weak evidence" isn't a verdict; "weak for this bet" is.
- The upgrade path is the constructive ending: when insufficient, name the cheapest test that would change the verdict — the holdout, the fake-door, the survey that sizes the interview theme, the pilot-with-metrics. Ranked by cost-to-confidence ratio; the skill's product is often not "no" but "this $0 two-week test first."
Output Format
Evidence Grade: "[the claim]" — feeding [the decision]
The Inventory
| Evidence |
Rung |
Quality notes (n, selection, date, independence) |
Direction |
The Weight
[What the graded net actually supports, in one honest paragraph — existence vs. prevalence vs. causation explicitly]
The Sufficiency Verdict
[The decision's stakes (reversibility × cost) · enough / not yet · the reasoning]
The Upgrade Path
[The cheapest evidence that moves the verdict · cost and time · the second option]
Quality Checks
Anti-Patterns
1---2name: evidence-grading3description: Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments > usage data > surveys > interviews > anecdotes > opinion), the fit-for-decision test, and the mixed-evidence verdicts that real questions produce. Use when asked how strong is our evidence for this, grade what we know before the decision, is this enough to bet on, or we have three anecdotes and a survey — now what. Produces the evidence inventory with grades, the sufficiency verdict against the decision's stakes, and the cheapest-upgrade path.4---5
6# Evidence Grading Skill
7
8"We have evidence" covers everything from a randomized experiment to the CEO's seatmate on a flight — and decisions made on ungraded evidence inherit the confusion. The grading discipline: inventory what actually supports the claim, place each item on the business-evidence hierarchy (what it *is*, not how confident it feels), test the graded weight against the decision's stakes (a reversible pilot needs less than a one-way rebrand — [decision-journal](../decision-journal/SKILL.md) logic), and when evidence falls short, name the *cheapest upgrade* — because the answer to weak evidence is usually a better test, not a braver bet.
9
10## What This Skill Produces
11
12- **The evidence inventory** — everything supporting (and contradicting) the claim, each item graded on the hierarchy
13- **The weight assessment** — what the graded pile actually supports, including the mixed-signals honest read
14- **The sufficiency verdict** — enough for this decision's stakes / not yet — with the stakes analysis shown
15- **The upgrade path** — the cheapest next evidence that would move the verdict ("a 2-week holdout test settles this for $0")
16
17## Required Inputs
18
19Ask for these if not provided:
20- **The claim and the decision riding on it** — "users want X" feeding a backlog item vs. feeding a repositioning are different sufficiency bars; the decision's reversibility and cost set the bar
21- **The evidence, itemized** — every piece: the data pull, the survey, the five customer quotes, the competitor's move, the expert's opinion — including the inconvenient items (an inventory that omits contradicting evidence is advocacy)
22- **The evidence's provenance** — n, selection, dates, who collected it and with what incentive ([source-triangulation](../source-triangulation/SKILL.md) supplies the externals; internal evidence has incentives too)
23
24## Framework: The Grading Rules
25
261. **The hierarchy, applied without sentiment:** experiments/A-B tests (causal) > usage/behavioral data (what people *do*, correlational) > surveys (what they *say*, at scale — [survey-design-basics](../survey-design-basics/SKILL.md) quality adjusts the grade) > interviews (rich, small-n — [interview-synthesis](../interview-synthesis/SKILL.md) counts matter) > anecdotes (existence proofs only) > expert opinion (informed priors) > internal conviction (a hypothesis, not evidence). Each item gets its rung *and its quality-within-rung* — a leading survey grades below honest interviews.
272. **Direction and independence both count:** the inventory includes contradicting evidence at its own grade, and echo-checks the supporting pile (three anecdotes traceable to one loud customer are one anecdote). The weight is the *net*, honestly netted.
283. **Anecdotes prove existence, never prevalence:** "a customer asked for X" establishes that the need exists somewhere — it cannot establish how common; the classic grading error is prevalence conclusions from existence evidence, and it's the error this skill most often catches.
294. **Sufficiency is relative to stakes:** the verdict tests weight against the decision — reversible-and-cheap decisions legitimately run on interview-grade evidence (the pilot *is* the experiment); irreversible-and-expensive ones demand behavioral or experimental grade. "Weak evidence" isn't a verdict; "weak for *this* bet" is.
305. **The upgrade path is the constructive ending:** when insufficient, name the cheapest test that would change the verdict — the holdout, the fake-door, the survey that sizes the interview theme, the pilot-with-metrics. Ranked by cost-to-confidence ratio; the skill's product is often not "no" but "this $0 two-week test first."
31
32## Output Format
33
34# Evidence Grade: "[the claim]" — feeding [the decision]
35
36## The Inventory
37| Evidence | Rung | Quality notes (n, selection, date, independence) | Direction |
38|---|---|---|---|
39
40## The Weight
41[What the graded net actually supports, in one honest paragraph — existence vs. prevalence vs. causation explicitly]
42
43## The Sufficiency Verdict
44[The decision's stakes (reversibility × cost) · enough / not yet · the reasoning]
45
46## The Upgrade Path
47[The cheapest evidence that moves the verdict · cost and time · the second option]
48
49## Quality Checks
50
51- [ ] Every item has a rung and within-rung quality notes
52- [ ] Contradicting evidence is in the inventory at its own grade
53- [ ] Echoes were collapsed before weighing
54- [ ] The verdict is stakes-relative, not absolute
55- [ ] The insufficient branch ends in a priced upgrade, not just a no
56
57## Anti-Patterns
58
59- [ ] Do not grade by vividness — the memorable anecdote outshines the boring dataset in every meeting; the hierarchy exists to resist exactly that
60- [ ] Do not conclude prevalence from existence — the most common grading felony
61- [ ] Do not omit the inconvenient items — an advocacy inventory grades the author, not the claim
62- [ ] Do not demand experimental grade for reversible bets — over-evidencing cheap decisions is its own waste
63- [ ] Do not end at "insufficient" — the upgrade path is the difference between rigor and obstruction