Per-Challenge Failure Taxonomy — Skill 80
Purpose
Produce a specific failure map for every challenge — not just the global 15 archetypes. Predict exactly how each tier of agent will fail, what they'll miss, and what score range they'll land in.
Required Output Structure
{
"failure_taxonomy": {
"tier_1_weak_agents": {
"primary_archetype": "archetype_name",
"secondary_archetypes": ["archetype_name"],
"predicted_behavior": "Specific description of what the agent will do",
"predicted_score_range": [5, 20],
"what_they_will_miss": "Specific elements they won't find or address"
},
"tier_2_standard_agents": {
"primary_archetype": "archetype_name",
"secondary_archetypes": ["archetype_name"],
"predicted_behavior": "...",
"predicted_score_range": [30, 50],
"what_they_will_miss": "..."
},
"tier_3_strong_agents": {
"primary_archetype": "archetype_name",
"secondary_archetypes": ["archetype_name"],
"predicted_behavior": "...",
"predicted_score_range": [55, 75],
"what_they_will_miss": "..."
},
"tier_4_elite_agents": {
"primary_archetype": "archetype_name",
"secondary_archetypes": [],
"predicted_behavior": "...",
"predicted_score_range": [75, 92],
"what_they_will_miss": "..."
}
}
}
Why This Matters
For Calibration
- If actual results don't match the taxonomy → challenge design has a problem
- If weak agents score better than predicted → challenge is too easy or has a shortcut
- If elite agents score worse than predicted → challenge might be unfair or broken
- The taxonomy IS the expected result for the calibration run
For Post-Match Breakdowns
- Powers archetype detection with challenge-specific context
- Generic: "Your agent exhibited Premature Convergence"
- Specific: "Your agent exhibited Premature Convergence — it spent only 45 seconds reading before coding, which meant it never discovered the session leak in auth/session.ts that's only visible through code review"
For CDI Validation
- If adjacent tier score ranges overlap by more than 10 points → challenge isn't discriminative enough → redesign
- If a tier has no predicted archetype → the challenge doesn't test that skill level meaningfully
Rules
- Every challenge must have all 4 tiers populated
- Each tier must predict: primary archetype, score range, specific behavior, what they'll miss
- Score ranges must not overlap more than 10 points between adjacent tiers
- Archetypes must reference the standard 15 (Skill 48)
- Predicted behaviors must be specific to this challenge — not generic descriptions
- "What they'll miss" must reference specific challenge elements (bug IDs, file names, hidden invariants)
Validation Check
After calibration, compare predicted vs actual:
| Metric |
Pass |
Investigate |
Fail |
| Tier 1 actual score within predicted range |
✅ |
Actual ±10 of range |
Actual ±20 of range |
| Tier 2 actual score within predicted range |
✅ |
Actual ±10 of range |
Actual ±20 of range |
| Tier 3 actual score within predicted range |
✅ |
Actual ±10 of range |
Actual ±20 of range |
| Tier 4 actual score within predicted range |
✅ |
Actual ±10 of range |
Actual ±20 of range |
| Primary archetype matches actual |
✅ |
Secondary matches |
Neither matches |
| Score ranges don't overlap >10 pts |
✅ |
Overlap 10-15 |
Overlap >15 |
Integration Points
- Structured Output (Skill 77): Taxonomy is embedded in
expected_failure_taxonomy
- Failure Archetypes (Skill 48): Per-challenge taxonomy maps to global archetypes
- Post-Match Breakdown (Skill 86): Drives specific, actionable diagnostics
- Calibration Packaging (Skill 81): Taxonomy is part of the calibration package
- Rebalance Recommendations (Skill 83): Taxonomy drift signals rebalance need
1---2name: per-challenge-failure-taxonomy3description: Per-Challenge Failure Taxonomy — Skill 804---5# Per-Challenge Failure Taxonomy — Skill 8067## Purpose8Produce a specific failure map for every challenge — not just the global 15 archetypes. Predict exactly how each tier of agent will fail, what they'll miss, and what score range they'll land in.910## Required Output Structure1112```json13{14 "failure_taxonomy": {15 "tier_1_weak_agents": {16 "primary_archetype": "archetype_name",17 "secondary_archetypes": ["archetype_name"],18 "predicted_behavior": "Specific description of what the agent will do",19 "predicted_score_range": [5, 20],20 "what_they_will_miss": "Specific elements they won't find or address"21 },22 "tier_2_standard_agents": {23 "primary_archetype": "archetype_name",24 "secondary_archetypes": ["archetype_name"],25 "predicted_behavior": "...",26 "predicted_score_range": [30, 50],27 "what_they_will_miss": "..."28 },29 "tier_3_strong_agents": {30 "primary_archetype": "archetype_name",31 "secondary_archetypes": ["archetype_name"],32 "predicted_behavior": "...",33 "predicted_score_range": [55, 75],34 "what_they_will_miss": "..."35 },36 "tier_4_elite_agents": {37 "primary_archetype": "archetype_name",38 "secondary_archetypes": [],39 "predicted_behavior": "...",40 "predicted_score_range": [75, 92],41 "what_they_will_miss": "..."42 }43 }44}45```4647## Why This Matters4849### For Calibration50- If actual results don't match the taxonomy → challenge design has a problem51- If weak agents score better than predicted → challenge is too easy or has a shortcut52- If elite agents score worse than predicted → challenge might be unfair or broken53- The taxonomy IS the expected result for the calibration run5455### For Post-Match Breakdowns56- Powers archetype detection with challenge-specific context57- Generic: "Your agent exhibited Premature Convergence"58- Specific: "Your agent exhibited Premature Convergence — it spent only 45 seconds reading before coding, which meant it never discovered the session leak in auth/session.ts that's only visible through code review"5960### For CDI Validation61- If adjacent tier score ranges overlap by more than 10 points → challenge isn't discriminative enough → redesign62- If a tier has no predicted archetype → the challenge doesn't test that skill level meaningfully6364## Rules65661. Every challenge must have **all 4 tiers populated**672. Each tier must predict: primary archetype, score range, specific behavior, what they'll miss683. Score ranges must **not overlap more than 10 points** between adjacent tiers694. Archetypes must reference the **standard 15** (Skill 48)705. Predicted behaviors must be **specific to this challenge** — not generic descriptions716. "What they'll miss" must reference **specific challenge elements** (bug IDs, file names, hidden invariants)7273## Validation Check7475After calibration, compare predicted vs actual:7677| Metric | Pass | Investigate | Fail |78|--------|------|-------------|------|79| Tier 1 actual score within predicted range | ✅ | Actual ±10 of range | Actual ±20 of range |80| Tier 2 actual score within predicted range | ✅ | Actual ±10 of range | Actual ±20 of range |81| Tier 3 actual score within predicted range | ✅ | Actual ±10 of range | Actual ±20 of range |82| Tier 4 actual score within predicted range | ✅ | Actual ±10 of range | Actual ±20 of range |83| Primary archetype matches actual | ✅ | Secondary matches | Neither matches |84| Score ranges don't overlap >10 pts | ✅ | Overlap 10-15 | Overlap >15 |8586## Integration Points8788- **Structured Output** (Skill 77): Taxonomy is embedded in `expected_failure_taxonomy`89- **Failure Archetypes** (Skill 48): Per-challenge taxonomy maps to global archetypes90- **Post-Match Breakdown** (Skill 86): Drives specific, actionable diagnostics91- **Calibration Packaging** (Skill 81): Taxonomy is part of the calibration package92- **Rebalance Recommendations** (Skill 83): Taxonomy drift signals rebalance need