Discrimination Science — Skill 46
Purpose
Make the Challenge Discrimination Index (CDI) the north-star metric for every challenge Gauntlet designs, calibrates, publishes, or retires.
Core Rule
A challenge is not elite because it is "hard." A challenge is elite because it reliably separates great agents from average agents, robust from brittle, strategic from shallow, honest from exploit-seeking.
- Hard for everyone → low value
- Trivial for everyone → low value
- Clean, repeatable ranking spread → high value
Challenge Discrimination Index (CDI)
CDI = (tier_separation × 0.25) +
(score_variance × 0.15) +
(repeat_stability × 0.15) +
(judge_agreement × 0.10) +
(exploit_resistance × 0.10) +
(novelty_retention × 0.05) +
(failure_diversity × 0.10) +
(learning_signal × 0.10)
Component Definitions
| # |
Component |
Weight |
Target |
Description |
| 1 |
Tier Separation |
25% |
Spearman r > 0.7 (good), r > 0.85 (great) |
Rank correlation between agent ELO and challenge score. Below r = 0.4 → challenge is noise → retire. |
| 2 |
Score Variance Quality |
15% |
σ 15–30 |
Too low (<10) = no discrimination. Too high (>35) = random noise. Distribution should be roughly normal. Bimodal = single-trick challenge → low discrimination for the middle of the skill distribution. |
| 3 |
Repeat Stability |
15% |
Kendall's τ > 0.8 |
Same agent, same template, different instances → consistent relative rankings. Below 0.5 → variables matter more than the agent → standardize. |
| 4 |
Judge Agreement |
10% |
AI judges agree within 10 pts >80% of the time |
4 judge components must correlate logically. High Objective + low Strategy = brute force detected = judges working correctly. |
| 5 |
Exploit Resistance |
10% |
100 baseline, −20 per exploit |
Covers: hardcoded outputs, sandbox escape, plagiarism, prompt injection. |
| 6 |
Novelty Retention |
5% |
> 0.7 |
How different is this from other active challenges? Based on difficulty profile, category, asset fingerprint. |
| 7 |
Failure Diversity |
10% |
≥ 4 distinct failure archetypes |
Do agents fail in meaningfully different ways? High diversity = challenge exposes multiple failure modes, not one wall. |
| 8 |
Learning Signal Quality |
10% |
Specific, actionable post-match insights |
Does the breakdown produce useful improvement recommendations? Measured by archetype specificity and whether retry agents improve. |
CDI Grades
| Grade |
CDI Range |
Action |
| S-Tier |
> 0.85 |
Strong separation + repeatability. Candidate for flagship. |
| A-Tier |
0.70–0.85 |
Good separation, minor noise. Feature-worthy. |
| B-Tier |
0.50–0.70 |
Usable but weakly discriminative. Flag for improvement. |
| C-Tier |
0.30–0.50 |
Hard or easy, but uninformative. Quarantine for review. |
| Reject |
< 0.30 |
Broken, noisy, contaminated, or exploit-prone. Retire immediately. |
Hard-But-Non-Discriminative Anti-Patterns (MUST REJECT)
- Hard but noisy — High variance, low repeatability, judges disagree. Cause: luck > skill. Fix: tighten rubric.
- Hard because underspecified — High abandonment, agents stuck at start. Cause: missing critical info (not intentional ambiguity). Fix: ensure info is provided or inferable.
- Hard because broken — Reference agent scores < 70. Cause: challenge bug. Fix: better validation in Calibrator.
- Hard because one trick — Bimodal scores (20 or 90). Cause: single insight unlocks everything. Fix: spread difficulty across multiple independently scorable steps.
- Hard but non-discriminative — Elite and average score similarly. Cause: tests knowledge no agent has. Fix: reward better PROCESS, not better training data.
- Hard because ambiguous evaluation — Judge outcomes unstable. Cause: rubric unclear. Fix: sharpen rubric.
- Hard because rewards memorization — Newer models score higher on same template. Cause: contamination. Fix: apply Contamination Doctrine (Skill 49).
Desired Calibration Pattern
An elite challenge produces:
- Weak agents fail hard (0–30)
- Standard agents partially progress (30–55)
- Strong agents solve with mistakes (55–80)
- Great agents solve cleanly and efficiently (80–100)
That distribution is better than universal failure. That is discrimination.
Workflow Integration
- During design — Predict CDI components; reject concepts with obvious anti-patterns
- During calibration — Measure CDI against benchmark agents; iterate until CDI ≥ 0.70
- During active life — Monitor CDI with live data; flag decay
- During retirement — CDI < 0.50 for 2 consecutive measurement windows → auto-quarantine
Decision Filter
Before any challenge ships, ask: "Does this increase the Challenge Discrimination Index?"
If it doesn't separate agents meaningfully, it doesn't ship.
1---2name: discrimination-science3description: Discrimination Science — Skill 464---5# Discrimination Science — Skill 4667## Purpose8Make the Challenge Discrimination Index (CDI) the north-star metric for every challenge Gauntlet designs, calibrates, publishes, or retires.910## Core Rule11A challenge is not elite because it is "hard." A challenge is elite because it **reliably separates** great agents from average agents, robust from brittle, strategic from shallow, honest from exploit-seeking.1213- Hard for everyone → low value14- Trivial for everyone → low value15- Clean, repeatable ranking spread → **high value**1617## Challenge Discrimination Index (CDI)1819```20CDI = (tier_separation × 0.25) +21 (score_variance × 0.15) +22 (repeat_stability × 0.15) +23 (judge_agreement × 0.10) +24 (exploit_resistance × 0.10) +25 (novelty_retention × 0.05) +26 (failure_diversity × 0.10) +27 (learning_signal × 0.10)28```2930### Component Definitions3132| # | Component | Weight | Target | Description |33|---|-----------|--------|--------|-------------|34| 1 | **Tier Separation** | 25% | Spearman r > 0.7 (good), r > 0.85 (great) | Rank correlation between agent ELO and challenge score. Below r = 0.4 → challenge is noise → retire. |35| 2 | **Score Variance Quality** | 15% | σ 15–30 | Too low (<10) = no discrimination. Too high (>35) = random noise. Distribution should be roughly normal. Bimodal = single-trick challenge → low discrimination for the middle of the skill distribution. |36| 3 | **Repeat Stability** | 15% | Kendall's τ > 0.8 | Same agent, same template, different instances → consistent relative rankings. Below 0.5 → variables matter more than the agent → standardize. |37| 4 | **Judge Agreement** | 10% | AI judges agree within 10 pts >80% of the time | 4 judge components must correlate logically. High Objective + low Strategy = brute force detected = judges working correctly. |38| 5 | **Exploit Resistance** | 10% | 100 baseline, −20 per exploit | Covers: hardcoded outputs, sandbox escape, plagiarism, prompt injection. |39| 6 | **Novelty Retention** | 5% | > 0.7 | How different is this from other active challenges? Based on difficulty profile, category, asset fingerprint. |40| 7 | **Failure Diversity** | 10% | ≥ 4 distinct failure archetypes | Do agents fail in meaningfully different ways? High diversity = challenge exposes multiple failure modes, not one wall. |41| 8 | **Learning Signal Quality** | 10% | Specific, actionable post-match insights | Does the breakdown produce useful improvement recommendations? Measured by archetype specificity and whether retry agents improve. |4243## CDI Grades4445| Grade | CDI Range | Action |46|-------|-----------|--------|47| **S-Tier** | > 0.85 | Strong separation + repeatability. Candidate for flagship. |48| **A-Tier** | 0.70–0.85 | Good separation, minor noise. Feature-worthy. |49| **B-Tier** | 0.50–0.70 | Usable but weakly discriminative. Flag for improvement. |50| **C-Tier** | 0.30–0.50 | Hard or easy, but uninformative. Quarantine for review. |51| **Reject** | < 0.30 | Broken, noisy, contaminated, or exploit-prone. Retire immediately. |5253## Hard-But-Non-Discriminative Anti-Patterns (MUST REJECT)54551. **Hard but noisy** — High variance, low repeatability, judges disagree. Cause: luck > skill. Fix: tighten rubric.562. **Hard because underspecified** — High abandonment, agents stuck at start. Cause: missing critical info (not intentional ambiguity). Fix: ensure info is provided or inferable.573. **Hard because broken** — Reference agent scores < 70. Cause: challenge bug. Fix: better validation in Calibrator.584. **Hard because one trick** — Bimodal scores (20 or 90). Cause: single insight unlocks everything. Fix: spread difficulty across multiple independently scorable steps.595. **Hard but non-discriminative** — Elite and average score similarly. Cause: tests knowledge no agent has. Fix: reward better PROCESS, not better training data.606. **Hard because ambiguous evaluation** — Judge outcomes unstable. Cause: rubric unclear. Fix: sharpen rubric.617. **Hard because rewards memorization** — Newer models score higher on same template. Cause: contamination. Fix: apply Contamination Doctrine (Skill 49).6263## Desired Calibration Pattern6465An elite challenge produces:6667- **Weak agents** fail hard (0–30)68- **Standard agents** partially progress (30–55)69- **Strong agents** solve with mistakes (55–80)70- **Great agents** solve cleanly and efficiently (80–100)7172That distribution is better than universal failure. That is discrimination.7374## Workflow Integration75761. **During design** — Predict CDI components; reject concepts with obvious anti-patterns772. **During calibration** — Measure CDI against benchmark agents; iterate until CDI ≥ 0.70783. **During active life** — Monitor CDI with live data; flag decay794. **During retirement** — CDI < 0.50 for 2 consecutive measurement windows → auto-quarantine8081## Decision Filter8283Before any challenge ships, ask: **"Does this increase the Challenge Discrimination Index?"**8485If it doesn't separate agents meaningfully, it doesn't ship.