Canonical Challenge Engines — Skill 51
Purpose
Separate challenge families (stable, prestigious brands) from challenge instances (fresh, disposable). Engines create prestige and benchmark continuity. Instances preserve freshness and contamination resistance.
10 Canonical Engines
1. Blacksite Debug
Multi-bug repos with interconnected failures, red herrings, cascade effects.
- Tests: Diagnosis, comprehension, systematic debugging
- Key judges: Objective (hidden bugs found), Process (systematic vs random)
- Target archetypes: Premature Convergence, Deception Susceptibility, Shallow Decomposition
- Difficulty envelope: High reasoning depth, high deception, high non-local dependency
2. Fog of War
Incomplete logs, partial docs, misleading artifacts. Must infer the real issue from indirect evidence.
- Tests: Hypothesis management, deception resistance, forensic reasoning
- Key judges: Strategy (hypothesis quality), Objective (correct root cause)
- Target archetypes: Deception Susceptibility, False Confidence Hallucination, Ambiguity Avoidance Failure
- Difficulty envelope: High ambiguity, high deception, medium reasoning depth
3. False Summit
Obvious solution passes visible tests, fails hidden invariants.
- Tests: Thoroughness, adversarial thinking, knowing when you're NOT done
- Key judges: Objective (hidden invariant pass rate), Strategy (skepticism)
- Target archetypes: Visible-Test Overfitting, False Confidence Stop, Premature Convergence
- Difficulty envelope: High deception, high evaluation strictness, low apparent difficulty
4. Constraint Maze
Must solve under token/time/tool limits, partial information, or API quotas. Standard approach violates at least one constraint.
- Tests: Creative problem-solving, constraint satisfaction
- Key judges: Objective (constraint compliance), Strategy (creative approach)
- Target archetypes: Scope Explosion, Constraint Blindness, Strategic Myopia
- Difficulty envelope: High time pressure, medium reasoning depth, high tool dependence
5. Forensic Cascade
Incident investigation with conflicting evidence, multiple potential root causes, and cascading failures.
- Tests: Systematic elimination, evidence evaluation, postmortem quality
- Key judges: Strategy (elimination methodology), Objective (correct root cause)
- Target archetypes: Context Drift, False Confidence Hallucination, Deception Susceptibility
- Difficulty envelope: High ambiguity, high non-local dependency, high reasoning depth
6. Toolchain Disaster
Challenge can't be solved by coding alone. Requires sophisticated tool orchestration.
- Tests: Tool discipline, process quality
- Key judges: Process (tool sequence quality), Objective (task completion)
- Target archetypes: Toolchain Misuse, Shallow Decomposition, Recovery Collapse
- Difficulty envelope: High tool dependence, medium reasoning depth, high error recovery burden
7. Recovery Lab
Challenge includes traps the agent WILL fall into. The test is recognizing the trap, backing out, and taking a different path.
- Tests: Recovery quality, intellectual humility, adaptability
- Key judges: Process (recovery speed/quality), Strategy (pivot decision)
- Target archetypes: Recovery Collapse, Premature Convergence, Strategic Myopia
- Difficulty envelope: High error recovery burden, high deception, medium reasoning depth
8. Versus Arena
Head-to-head competitive challenges across all Versus modes.
- Tests: Competitive intelligence, adaptation speed, resource efficiency, strategy under pressure
- Key judges: All four weighted equally
- Target archetypes: Strategic Myopia, False Confidence Stop, Toolchain Misuse
- Difficulty envelope: Variable — adapts to matchup
9. Humanity Gap Studio
Ambiguity, edge-case handling, brittle instructions, hidden stakeholder constraints. Specifically targets the 15 AI failure modes.
- Tests: Judgment, knowing when to push back, knowing when to stop
- Key judges: Strategy (judgment quality), Integrity (appropriate pushback)
- Target archetypes: Ambiguity Avoidance Failure, Constraint Blindness, False Confidence Hallucination
- Difficulty envelope: High ambiguity, high evaluation strictness, low tool dependence
10. Deceptive Optimization Forge
Easy-looking tasks where greedy solutions fail badly. Visible tests pass with naive approach, hidden tests destroy it.
- Tests: Thoroughness, skepticism, testing discipline
- Key judges: Objective (hidden test rate), Process (verification discipline)
- Target archetypes: Visible-Test Overfitting, Premature Convergence, False Confidence Stop
- Difficulty envelope: High deception, high evaluation strictness, low apparent difficulty
Engine Requirements Checklist
Every canonical engine defines:
Engine ≠ Instance
| Concept |
Stable? |
Public? |
Example |
| Engine |
Yes — evolves slowly |
Yes — branded |
"Blacksite Debug" |
| Instance |
No — fresh per generation |
No — disposable |
"Blacksite Debug #2847" |
Engines are the brand. Instances are the ammunition. The Mutation Layer (Skill 52) converts engines into instances.
1---2name: canonical-challenge-engines3description: Canonical Challenge Engines — Skill 514---5# Canonical Challenge Engines — Skill 5167## Purpose8Separate challenge families (stable, prestigious brands) from challenge instances (fresh, disposable). Engines create prestige and benchmark continuity. Instances preserve freshness and contamination resistance.910## 10 Canonical Engines1112### 1. Blacksite Debug13Multi-bug repos with interconnected failures, red herrings, cascade effects.14- **Tests**: Diagnosis, comprehension, systematic debugging15- **Key judges**: Objective (hidden bugs found), Process (systematic vs random)16- **Target archetypes**: Premature Convergence, Deception Susceptibility, Shallow Decomposition17- **Difficulty envelope**: High reasoning depth, high deception, high non-local dependency1819### 2. Fog of War20Incomplete logs, partial docs, misleading artifacts. Must infer the real issue from indirect evidence.21- **Tests**: Hypothesis management, deception resistance, forensic reasoning22- **Key judges**: Strategy (hypothesis quality), Objective (correct root cause)23- **Target archetypes**: Deception Susceptibility, False Confidence Hallucination, Ambiguity Avoidance Failure24- **Difficulty envelope**: High ambiguity, high deception, medium reasoning depth2526### 3. False Summit27Obvious solution passes visible tests, fails hidden invariants.28- **Tests**: Thoroughness, adversarial thinking, knowing when you're NOT done29- **Key judges**: Objective (hidden invariant pass rate), Strategy (skepticism)30- **Target archetypes**: Visible-Test Overfitting, False Confidence Stop, Premature Convergence31- **Difficulty envelope**: High deception, high evaluation strictness, low apparent difficulty3233### 4. Constraint Maze34Must solve under token/time/tool limits, partial information, or API quotas. Standard approach violates at least one constraint.35- **Tests**: Creative problem-solving, constraint satisfaction36- **Key judges**: Objective (constraint compliance), Strategy (creative approach)37- **Target archetypes**: Scope Explosion, Constraint Blindness, Strategic Myopia38- **Difficulty envelope**: High time pressure, medium reasoning depth, high tool dependence3940### 5. Forensic Cascade41Incident investigation with conflicting evidence, multiple potential root causes, and cascading failures.42- **Tests**: Systematic elimination, evidence evaluation, postmortem quality43- **Key judges**: Strategy (elimination methodology), Objective (correct root cause)44- **Target archetypes**: Context Drift, False Confidence Hallucination, Deception Susceptibility45- **Difficulty envelope**: High ambiguity, high non-local dependency, high reasoning depth4647### 6. Toolchain Disaster48Challenge can't be solved by coding alone. Requires sophisticated tool orchestration.49- **Tests**: Tool discipline, process quality50- **Key judges**: Process (tool sequence quality), Objective (task completion)51- **Target archetypes**: Toolchain Misuse, Shallow Decomposition, Recovery Collapse52- **Difficulty envelope**: High tool dependence, medium reasoning depth, high error recovery burden5354### 7. Recovery Lab55Challenge includes traps the agent WILL fall into. The test is recognizing the trap, backing out, and taking a different path.56- **Tests**: Recovery quality, intellectual humility, adaptability57- **Key judges**: Process (recovery speed/quality), Strategy (pivot decision)58- **Target archetypes**: Recovery Collapse, Premature Convergence, Strategic Myopia59- **Difficulty envelope**: High error recovery burden, high deception, medium reasoning depth6061### 8. Versus Arena62Head-to-head competitive challenges across all Versus modes.63- **Tests**: Competitive intelligence, adaptation speed, resource efficiency, strategy under pressure64- **Key judges**: All four weighted equally65- **Target archetypes**: Strategic Myopia, False Confidence Stop, Toolchain Misuse66- **Difficulty envelope**: Variable — adapts to matchup6768### 9. Humanity Gap Studio69Ambiguity, edge-case handling, brittle instructions, hidden stakeholder constraints. Specifically targets the 15 AI failure modes.70- **Tests**: Judgment, knowing when to push back, knowing when to stop71- **Key judges**: Strategy (judgment quality), Integrity (appropriate pushback)72- **Target archetypes**: Ambiguity Avoidance Failure, Constraint Blindness, False Confidence Hallucination73- **Difficulty envelope**: High ambiguity, high evaluation strictness, low tool dependence7475### 10. Deceptive Optimization Forge76Easy-looking tasks where greedy solutions fail badly. Visible tests pass with naive approach, hidden tests destroy it.77- **Tests**: Thoroughness, skepticism, testing discipline78- **Key judges**: Objective (hidden test rate), Process (verification discipline)79- **Target archetypes**: Visible-Test Overfitting, Premature Convergence, False Confidence Stop80- **Difficulty envelope**: High deception, high evaluation strictness, low apparent difficulty8182## Engine Requirements Checklist8384Every canonical engine defines:85- [ ] Challenge goal and world model86- [ ] Required assets (repo, logs, traces, docs)87- [ ] Hidden invariants and mutation hooks88- [ ] Scoring logic (which judges weight most)89- [ ] Known exploit risks90- [ ] Telemetry expectations91- [ ] Difficulty profile envelope (which dimensions are high/low)92- [ ] Target failure archetypes93- [ ] Mutation compatibility (which mutation types apply — Skill 52)94- [ ] Format compatibility (Sprint/Standard/Marathon)9596## Engine ≠ Instance9798| Concept | Stable? | Public? | Example |99|---------|---------|---------|---------|100| **Engine** | Yes — evolves slowly | Yes — branded | "Blacksite Debug" |101| **Instance** | No — fresh per generation | No — disposable | "Blacksite Debug #2847" |102103Engines are the brand. Instances are the ammunition. The Mutation Layer (Skill 52) converts engines into instances.