Self-Rate
Score your own work on a calibrated scale, then change the work to match the score — lower the claims you can't back, or do more work to earn a higher one. Honest self-assessment beats confident hand-waving. Pairs with [[but-for-real]] (which verifies) — this one grades.
Why calibrated
A number means nothing without anchors. Define what the scale's points actually mean before you score, or every answer drifts to "8/10, looks good".
Default 1–10 anchors:
- 1–3 — likely wrong, unverified, or missing the point; would embarrass me if shipped
- 4–6 — plausible but has known gaps, untested claims, or weak spots I can name
- 7–8 — solid; verified the core, minor caveats remain
- 9–10 — verified end-to-end, edge cases considered, I'd bet on it
Procedure
- Pick dimensions that fit the work — typically correctness/evidence, completeness, clarity, risk. For a single factual claim, one "confidence in this claim" score is enough.
- Score each dimension against the anchors, with a one-line justification naming the specific reason — "6: didn't run the migration on a populated table", not "6: seems okay".
- Act on the score — this is the point, not the number:
- Low score → add the caveat explicitly, or do the missing work to raise it
- Overclaim → downgrade the wording to match what you actually verified
- High score you can't justify → it's not high; re-score honestly
- Return the scores + the revised work.
Common mistakes
- Inflating to be agreeable — the user wants the real number, not reassurance.
- Scoring without anchors, so everything lands at 7–8.
- Justifications that restate the score ("8 because it's good") instead of naming evidence or the missing piece.
- Producing a score and then not changing anything — the rating exists to drive a revision.