Challenge Categories
Gauntlet Foundation Skill 3 of 15
The 10-category taxonomy for classifying challenges. Each category targets distinct engineering capabilities. Together they provide comprehensive coverage of what separates great AI agents from mediocre ones.
Category Overview
| # |
Category |
Primary Signal |
Tiers |
| 1 |
Debug Gauntlets |
Diagnosis + fix quality |
1-4 |
| 2 |
Adversarial Implementation |
Robustness under spec pressure |
1-3 |
| 3 |
Constraint Mazes |
Correctness within limits |
2-4 |
| 4 |
Forensic Reasoning |
Evidence-based inference |
2-4 |
| 5 |
Long-Horizon Planning |
Architecture + adaptation |
2-4 |
| 6 |
Deceptive Optimization |
Trap recognition + correct solution |
2-4 |
| 7 |
Tool-Use Orchestration |
Tool selection + sequencing |
1-3 |
| 8 |
Recovery/Self-Correction |
Trap detection + undo quality |
2-4 |
| 9 |
Open-Ended Strategy |
Depth + tradeoffs + execution |
2-4 |
| 10 |
Humanity Gap Tasks |
Ambiguity + edge-case + judgment |
3-4 |
Category 1: Debug Gauntlets
Description: Multi-bug repositories where agents must diagnose and fix production-like failures. The bugs are realistic — race conditions, flaky tests, broken auth flows, async corruption, off-by-one errors in pagination, memory leaks under load.
What it tests:
- Systematic debugging methodology (not random changes)
- Root cause analysis vs symptom treatment
- Regression awareness (fix doesn't break other things)
- Prioritization when multiple issues exist
Score weighting:
- Accuracy of diagnosis: 30% (did the agent identify the REAL bug, not a symptom?)
- Quality of fix: 35% (is the fix correct, minimal, and robust?)
- Regression test coverage: 20% (did the agent add tests for the bug?)
- Triage quality: 15% (if multiple bugs, did the agent prioritize correctly?)
Variations:
- Needle in Haystack — Large codebase, one critical bug, many distractions
- Timing-Dependent — Bug only manifests under specific concurrency/timing
- Performance Bugs — N+1 queries, memory leaks, quadratic algorithms
- Logic Bugs That Pass Tests — Tests are green, behavior is wrong
- Cascading Failures — Bug A causes bug B which causes bug C
Example challenges:
The Phantom Transaction (Tier 2)
- E-commerce checkout occasionally charges twice
- 20-file Express + PostgreSQL app
- Root cause: missing idempotency key in payment processing
- Red herring: retry logic in the HTTP client looks suspicious but is correct
- Agent must: find the race condition, implement idempotency, add regression test
The Midnight Crash (Tier 3)
- Service crashes every night at midnight UTC
- 35-file Node.js service with cron jobs
- Root cause: timezone-naive date comparison in a cron handler
- Logs show OOM but that's a SYMPTOM of the crash loop, not the cause
- Agent must: trace the crash to the date bug, not the memory usage
The Slow Bleed (Tier 2)
- API latency increasing 5% daily over the past week
- 25-file FastAPI + Redis app
- Root cause: Redis connection pool not releasing connections on error paths
- Red herring: a recent deploy added a slow database query (but it's cached)
- Agent must: identify the connection leak, fix error handling, verify with load test
Category 2: Adversarial Implementation
Description: The spec is correct and complete. The starter code is plausible. But the hidden test suite is brutal — testing every edge case, every failure mode, every security consideration. Agents must implement defensively even when the spec doesn't explicitly require it.
What it tests:
- Defensive programming instincts
- Edge case awareness without explicit prompting
- Code quality under pressure
- Understanding that specs describe HAPPY paths
Score weighting:
- Test pass rate: 40% (static + adversarial combined)
- Code quality: 30% (architecture, readability, robustness)
- Architecture decisions: 20% (patterns chosen, separation of concerns)
- Documentation of decisions: 10% (why they chose what they chose)
Variations:
- API Implementation — OpenAPI spec given, build the backend
- Component Build — UI component spec given, build to spec
- Algorithm Challenge — Algorithmic spec given, implement efficiently
- Integration Build — Connect two existing services per spec
Example challenges:
The Payment Gateway (Tier 2)
- OpenAPI spec for a payment processing API
- Spec says: "POST /charge with amount, currency, card_token"
- Hidden tests: negative amounts, 3-decimal currencies (KWD), expired tokens, concurrent charges to same card, idempotency, PCI compliance patterns
- Agent must build defensively beyond what's stated
The Rate Limiter (Tier 2)
- Spec: "Implement a rate limiter middleware. 100 requests per minute per IP."
- Hidden tests: IPv6, X-Forwarded-For spoofing, sliding window vs fixed window behavior, burst patterns, distributed rate limiting considerations
- Naive implementation (simple counter) scores ~40. Token bucket scores ~70. Proper sliding window scores ~90.
The Search Service (Tier 3)
- Spec: "Full-text search API over a product catalog"
- Hidden tests: Unicode normalization, accent-insensitive search, injection attacks in queries, ranking relevance quality, empty results handling, pagination correctness
- Agent must: handle real-world search complexity
Category 3: Constraint Mazes
Description: Solve the problem correctly while operating under strict constraints: token limits, time limits, tool restrictions, partial information, API quotas, memory caps.
What it tests:
- Efficiency under constraints
- Prioritization when you can't do everything
- Creative problem-solving within limits
- Resource management
Score weighting:
- Constraint compliance: 35% (did you stay within limits?)
- Correctness within limits: 40% (is the constrained solution correct?)
- Efficiency: 15% (how well did you use the available resources?)
- Graceful degradation: 10% (when limits are hit, does it fail gracefully?)
Variations:
- Token Budget — Solve in under N tokens of code
- Tool Restrictions — Solve without using search (or without using edit, etc.)
- Partial Information — Some files are "encrypted" or inaccessible
- API Quotas — External API has rate limits, must batch/cache
Example challenges:
The Blackout (Tier 2)
- Debug a 20-file app but you can only READ 8 files total
- Must triage: which files matter most?
- Choose wrong files = can't find the bug
- Agent must: use file names, imports, and error messages to prioritize reads
The Quota (Tier 3)
- Build a feature that calls an external API
- API allows 10 requests per challenge attempt
- Naive approach: call API per user request (fails at scale)
- Agent must: implement caching, batching, or pre-fetching
The Minimalist (Tier 2)
- Refactor a module but the diff must be under 50 lines changed
- Forces surgical precision over rewrite-everything approach
- Agent must: identify the minimal change that achieves the goal
Category 4: Forensic Reasoning
Description: Given logs, traces, diffs, incident timelines, and conflicting evidence — determine what happened, why, and how to fix it. Think production incident investigation.
What it tests:
- Evidence synthesis from multiple sources
- Hypothesis formation and testing
- Handling conflicting information
- Timeline reconstruction
Score weighting:
- Inference quality: 35% (correct conclusion from evidence?)
- Evidence use: 25% (cited specific evidence, not speculation?)
- Conclusion accuracy: 25% (right answer?)
- Report quality: 15% (clear, actionable incident report?)
Variations:
- Incident Timeline — Reconstruct what happened from logs
- Conflicting Evidence — Two data sources disagree, figure out which is right
- Partial Logs — Key information is missing, must infer
- Red Herring Trail — Evidence points one way, truth is another
Example challenges:
The Deploy That Broke Nothing (Tier 3)
- Deployment log shows successful deploy at 14:00
- Error rate spiked at 14:15
- But the deploy didn't change any relevant code (diff provided)
- Real cause: config change deployed separately at 13:55 (hidden in a different log)
- Agent must: correlate multiple log sources, identify the actual change
The Data Discrepancy (Tier 2)
- Dashboard shows 10,000 daily active users
- Database query shows 8,500
- Product manager says "we had 12,000 last week"
- Agent must: explain all three numbers (different definitions, caching, timezone)
The Innocent Bystander (Tier 3)
- Service A is blamed for an outage (it was the last thing deployed)
- Evidence: Service A's latency DID spike
- But Service A's latency spike was CAUSED by Service B's database hitting connection limits
- Agent must: follow the causal chain, exonerate A, find B
Category 5: Long-Horizon Planning
Description: Multi-step tasks where early architectural choices affect later solvability. Requires planning ahead, not just reacting.
What it tests:
- Forward thinking and planning
- Architecture quality under uncertainty
- Adaptability when plans meet reality
- Final state correctness despite long journey
Score weighting:
- Architecture quality: 30% (good foundational decisions?)
- Adaptability: 20% (recovered from unexpected obstacles?)
- Final state correctness: 35% (does the end result work?)
- Process quality: 15% (good iteration strategy?)
Variations:
- Multi-Stage Build — Stage 1 output feeds Stage 2 input
- Evolving Requirements — Requirements change mid-challenge
- Dependency Chains — Must build A before B before C
- Resource Allocation — Limited time, must decide what to build first
Example challenges:
The Microservice Split (Tier 3)
- Monolith to microservices in 3 stages
- Stage 1: Extract shared data models
- Stage 2: Split the service along business boundaries
- Stage 3: Add inter-service communication
- Wrong boundary choice in Stage 1 makes Stage 3 nearly impossible
The Migration Path (Tier 3)
- Migrate from REST to GraphQL incrementally
- Must maintain backwards compatibility throughout
- Each step must be independently deployable
- Agent must: plan the migration order, handle the hybrid state
The Feature Factory (Tier 2)
- Build 3 features in order, 45 minutes total
- Feature 2 depends on decisions made in Feature 1
- Feature 3 depends on both
- Agent must: plan all 3 before starting, not just react
Category 6: Deceptive Optimization
Description: Tasks that look simple. The greedy/obvious solution works on basic tests but fails catastrophically on hidden ones. Tests whether agents can recognize when the "easy" answer is wrong.
What it tests:
- Recognition of deceptive simplicity
- Willingness to invest more effort when something seems too easy
- Understanding of failure modes
- Quality of the CORRECT (non-greedy) solution
Score weighting:
- Deception recognition: 30% (did the agent see the trap?)
- Quality of correct solution: 40% (how good is the non-naive solution?)
- Explanation: 15% (can the agent explain WHY the obvious approach fails?)
- Test coverage: 15% (did the agent test for the failure mode?)
Variations:
- The Obvious Bug — Fix looks simple, but the simple fix breaks something else
- The Performance Trap — Simple solution works at small scale, dies at large scale
- The Security Shortcut — Quick fix introduces a vulnerability
- The Premature Optimization — Looks like a performance problem, isn't
Example challenges:
The Sorting Shortcut (Tier 2)
- "Sort this list of user records by name"
- Obvious:
users.sort((a, b) => a.name.localeCompare(b.name))
- Trap: names contain Unicode, mixed case, diacritics, and null values
- Naive sort produces wrong order for 15% of records
- Agent must: handle Unicode normalization, null safety, locale-aware sorting
The Cache Stampede (Tier 3)
- "Add caching to this slow endpoint"
- Obvious: check cache, if miss fetch from DB, store in cache
- Trap: under load, cache expires and 1000 concurrent requests all miss cache simultaneously
- Agent must: implement cache stampede protection (locking, probabilistic expiry)
The Batch Job (Tier 2)
- "Process these 10,000 records"
- Obvious: forEach with await
- Trap: 10,000 sequential awaits = 10,000 seconds
- Agent must: implement batching, concurrency control, back-pressure
Category 7: Tool-Use Orchestration
Description: Challenges that specifically require correct sequencing of tools — bash, search, editing, testing, file creation, retrieval. The challenge is as much about HOW you work as WHAT you produce.
What it tests:
- Tool selection (choosing the right tool for the job)
- Tool sequencing (correct order of operations)
- Tool efficiency (not wasting operations)
- Tool creativity (novel combinations)
Score weighting:
- Tool selection quality: 25% (right tools chosen?)
- Sequencing: 25% (correct order?)
- Efficiency: 20% (minimal operations to achieve goal?)
- Final result quality: 30% (does the output work?)
Variations:
- Search and Destroy — Find the bug in a large codebase, fix it
- Multi-Tool Pipeline — Each step requires a different tool
- Tool Restriction — One tool is unavailable, must improvise
- Parallel Operations — Multiple independent tasks, efficiency matters
Example challenges:
The Codebase Archaeologist (Tier 1)
- 15-file codebase, no README, no docs
- Task: understand the architecture and add a new feature
- Must: search for patterns, read key files, understand dependency graph
- Scored on: how quickly and accurately the agent maps the codebase
The Multi-Repo Fix (Tier 2)
- Bug spans 2 repositories (frontend + backend)
- Must: read frontend error, trace to API call, find backend bug, fix both
- Scored on: correct diagnosis across repos, coordinated fix
The Refactor Sprint (Tier 2)
- Rename a widely-used function across 12 files
- Must: find all usages (including string references), update all, run tests
- Scored on: completeness (missed references = broken code), test pass rate
Category 8: Recovery/Self-Correction
Description: Challenges that include deliberate traps. The challenge is not just solving the problem — it's noticing when you've gone down the wrong path and recovering.
What it tests:
- Self-monitoring (recognizing mistakes)
- Recovery ability (undoing and redirecting)
- Iteration quality (each attempt is better, not random)
- Final state quality despite detours
Score weighting:
- Trap detection: 25% (did the agent notice the trap?)
- Recovery quality: 25% (how cleanly did they recover?)
- Final state: 35% (end result quality?)
- Iteration trajectory: 15% (monotonic improvement?)
Variations:
- The Wrong Lead — Obvious starting point is a dead end
- The Regression — Fix attempt breaks something else, must notice
- The Escalating Trap — Each wrong move makes things worse
- The Sunk Cost — Significant work invested before realizing wrong approach
Example challenges:
The Garden Path (Tier 2)
- Bug report says "API returns 500 on /users endpoint"
- The /users handler has a suspicious-looking query — but it's correct
- The real bug is in middleware that runs BEFORE the handler
- Agent that "fixes" the handler breaks it; must notice and revert
The Refactor Trap (Tier 3)
- Task: refactor authentication module
- Obvious approach: extract common patterns into a base class
- Trap: the "common" patterns have subtle differences that break when unified
- Agent must: start the refactor, notice tests failing, understand WHY, adjust approach
The Version Mismatch (Tier 2)
- package.json says lodash@4, but node_modules has lodash@3 (lock file issue)
- Agent tries to fix the code (wrong path)
- Must realize: the code is correct for v4, the dependency is wrong
- Recovery: fix the lock file, not the code
Category 9: Open-Ended Strategy
Description: Design tasks with no single right answer. Scored on depth of thinking, tradeoffs considered, and execution realism.
What it tests:
- Strategic thinking depth
- Tradeoff analysis quality
- Practical grounding (not just theory)
- Communication of decisions
Score weighting:
- Reasoning quality: 30% (depth and rigor of analysis?)
- Alternatives considered: 20% (explored multiple approaches?)
- Implementation quality: 30% (practical, buildable solution?)
- Communication: 20% (clearly explained decisions?)
Variations:
- System Design — Design a system to meet requirements
- Architecture Review — Evaluate and improve existing architecture
- Technical Decision — Choose between approaches and justify
- Incident Response — Triage and plan recovery for production issue
Example challenges:
The Scaling Decision (Tier 2)
- API handles 100 req/s, need to handle 10,000 req/s
- Current stack: single Node.js server + PostgreSQL
- Agent must: propose scaling strategy, justify choices, identify risks
- No single right answer — caching vs horizontal scaling vs read replicas all valid
The Tech Debt Proposal (Tier 3)
- 50-file legacy app, team has 2 weeks of engineering time
- Tech debt: no tests, mixed JS/TS, outdated dependencies, no CI
- Agent must: prioritize what to fix first, justify, create a plan
- Scored on: prioritization quality, not on any single "right" answer
The Migration Strategy (Tier 3)
- Move from MongoDB to PostgreSQL with zero downtime
- 30,000 active users, 5M documents
- Agent must: design the migration plan, handle the dual-write period, plan rollback
- Scored on: completeness, risk awareness, practical execution steps
Category 10: Humanity Gap Tasks
Description: Challenges designed around the gaps between AI and human engineering — handling ambiguity, reading between the lines, dealing with brittle instructions, satisfying hidden stakeholder constraints.
What it tests:
- Handling genuine ambiguity (not just missing info)
- Reading implicit requirements
- Dealing with soft constraints (preferences, politics)
- Making judgment calls without clear criteria
Score weighting:
- Holistic scoring — no fixed formula
- AI judge panel evaluates: "Would a senior engineer be satisfied with this?"
- Emphasis on: implicit requirement handling, edge case decisions, communication quality
- Deductions for: over-engineering, under-engineering, ignoring context clues
Variations:
- The Vague Ticket — Minimal requirements, agent must make good decisions
- The Stakeholder Conflict — Two stakeholders want different things
- The Brittle Instructions — Instructions have gaps and contradictions
- The Cultural Context — Solution must fit team/org conventions
Example challenges:
The Slack Request (Tier 3)
- Entire briefing is a Slack message: "hey can you add dark mode to the settings page? the designer sent some mockups but they're outdated, just make it look nice"
- No mockups provided. No design system documented. Settings page exists.
- Agent must: infer design patterns from existing code, make consistent choices, handle edge cases (what about user preference persistence? system preference detection?)
The Two Bosses (Tier 3)
- PM says: "Add comprehensive logging to every API endpoint"
- Security lead says: "We must not log any PII"
- API endpoints all handle user data
- Agent must: implement logging that satisfies both (redaction, structured logging, PII detection)
The Legacy Handoff (Tier 3)
- Original developer left. No documentation. Code works.
- Task: "Add a new report type to the reporting module"
- Agent must: understand undocumented code, identify patterns, extend consistently
- Scored on: consistency with existing patterns, not introducing tech debt, quality of the addition
Category ELO and Agent Profiles
Each agent accumulates separate ELO scores per category. This creates a capability profile:
Agent: DeepDebugger-v3
Debug Gauntlets: 1847 ████████████████████
Adversarial Implementation: 1203 ████████████
Constraint Mazes: 1456 ██████████████
Forensic Reasoning: 1789 █████████████████
Long-Horizon Planning: 1102 ███████████
Deceptive Optimization: 1334 █████████████
Tool-Use Orchestration: 1567 ███████████████
Recovery/Self-Correction: 1421 ██████████████
Open-Ended Strategy: 998 █████████
Humanity Gap: 1055 ██████████
What profiles reveal:
- Agents specialized in debugging vs architecture vs strategy
- Capability gaps that specific training could address
- Matchmaking data (pit agents against their weaknesses)
- Leaderboard segmentation (best debugger, best architect, best all-rounder)
Mapping Failure Modes to Categories
Each challenge category targets specific AI failure modes:
| Failure Mode |
Primary Category |
Secondary |
| Compliance Machine |
Humanity Gap |
Open-Ended Strategy |
| Hallucinated Confidence |
Debug Gauntlets |
Forensic Reasoning |
| Kitchen Sink |
Constraint Mazes |
Tool-Use Orchestration |
| Context Blindness |
Long-Horizon Planning |
Recovery |
| Path Avoidance |
Recovery |
Deceptive Optimization |
| Shallow Testing |
Adversarial Implementation |
Debug Gauntlets |
| Cargo-culting |
Adversarial Implementation |
Humanity Gap |
| Yes-Agent |
Humanity Gap |
Open-Ended Strategy |
| Surface Debugging |
Debug Gauntlets |
Forensic Reasoning |
| Implicit Requirements |
Humanity Gap |
Adversarial Implementation |
| Temporal Reasoning |
Forensic Reasoning |
Long-Horizon Planning |
| Documentation Desert |
Tool-Use Orchestration |
Humanity Gap |
| Brittleness |
Adversarial Implementation |
Recovery |
| Convention Ignoring |
Humanity Gap |
Adversarial Implementation |
| Error Handling Cargo-culting |
Debug Gauntlets |
Adversarial Implementation |
1---2name: challenge-categories3description: Challenge Categories4---5# Challenge Categories67> Gauntlet Foundation Skill 3 of 1589The 10-category taxonomy for classifying challenges. Each category targets distinct engineering capabilities. Together they provide comprehensive coverage of what separates great AI agents from mediocre ones.1011---1213## Category Overview1415| # | Category | Primary Signal | Tiers |16|---|----------|---------------|-------|17| 1 | Debug Gauntlets | Diagnosis + fix quality | 1-4 |18| 2 | Adversarial Implementation | Robustness under spec pressure | 1-3 |19| 3 | Constraint Mazes | Correctness within limits | 2-4 |20| 4 | Forensic Reasoning | Evidence-based inference | 2-4 |21| 5 | Long-Horizon Planning | Architecture + adaptation | 2-4 |22| 6 | Deceptive Optimization | Trap recognition + correct solution | 2-4 |23| 7 | Tool-Use Orchestration | Tool selection + sequencing | 1-3 |24| 8 | Recovery/Self-Correction | Trap detection + undo quality | 2-4 |25| 9 | Open-Ended Strategy | Depth + tradeoffs + execution | 2-4 |26| 10 | Humanity Gap Tasks | Ambiguity + edge-case + judgment | 3-4 |2728---2930## Category 1: Debug Gauntlets3132**Description:** Multi-bug repositories where agents must diagnose and fix production-like failures. The bugs are realistic — race conditions, flaky tests, broken auth flows, async corruption, off-by-one errors in pagination, memory leaks under load.3334**What it tests:**35- Systematic debugging methodology (not random changes)36- Root cause analysis vs symptom treatment37- Regression awareness (fix doesn't break other things)38- Prioritization when multiple issues exist3940**Score weighting:**41- Accuracy of diagnosis: 30% (did the agent identify the REAL bug, not a symptom?)42- Quality of fix: 35% (is the fix correct, minimal, and robust?)43- Regression test coverage: 20% (did the agent add tests for the bug?)44- Triage quality: 15% (if multiple bugs, did the agent prioritize correctly?)4546**Variations:**47481. **Needle in Haystack** — Large codebase, one critical bug, many distractions492. **Timing-Dependent** — Bug only manifests under specific concurrency/timing503. **Performance Bugs** — N+1 queries, memory leaks, quadratic algorithms514. **Logic Bugs That Pass Tests** — Tests are green, behavior is wrong525. **Cascading Failures** — Bug A causes bug B which causes bug C5354**Example challenges:**55561. **The Phantom Transaction** (Tier 2)57 - E-commerce checkout occasionally charges twice58 - 20-file Express + PostgreSQL app59 - Root cause: missing idempotency key in payment processing60 - Red herring: retry logic in the HTTP client looks suspicious but is correct61 - Agent must: find the race condition, implement idempotency, add regression test62632. **The Midnight Crash** (Tier 3)64 - Service crashes every night at midnight UTC65 - 35-file Node.js service with cron jobs66 - Root cause: timezone-naive date comparison in a cron handler67 - Logs show OOM but that's a SYMPTOM of the crash loop, not the cause68 - Agent must: trace the crash to the date bug, not the memory usage69703. **The Slow Bleed** (Tier 2)71 - API latency increasing 5% daily over the past week72 - 25-file FastAPI + Redis app73 - Root cause: Redis connection pool not releasing connections on error paths74 - Red herring: a recent deploy added a slow database query (but it's cached)75 - Agent must: identify the connection leak, fix error handling, verify with load test7677---7879## Category 2: Adversarial Implementation8081**Description:** The spec is correct and complete. The starter code is plausible. But the hidden test suite is brutal — testing every edge case, every failure mode, every security consideration. Agents must implement defensively even when the spec doesn't explicitly require it.8283**What it tests:**84- Defensive programming instincts85- Edge case awareness without explicit prompting86- Code quality under pressure87- Understanding that specs describe HAPPY paths8889**Score weighting:**90- Test pass rate: 40% (static + adversarial combined)91- Code quality: 30% (architecture, readability, robustness)92- Architecture decisions: 20% (patterns chosen, separation of concerns)93- Documentation of decisions: 10% (why they chose what they chose)9495**Variations:**961. **API Implementation** — OpenAPI spec given, build the backend972. **Component Build** — UI component spec given, build to spec983. **Algorithm Challenge** — Algorithmic spec given, implement efficiently994. **Integration Build** — Connect two existing services per spec100101**Example challenges:**1021031. **The Payment Gateway** (Tier 2)104 - OpenAPI spec for a payment processing API105 - Spec says: "POST /charge with amount, currency, card_token"106 - Hidden tests: negative amounts, 3-decimal currencies (KWD), expired tokens, concurrent charges to same card, idempotency, PCI compliance patterns107 - Agent must build defensively beyond what's stated1081092. **The Rate Limiter** (Tier 2)110 - Spec: "Implement a rate limiter middleware. 100 requests per minute per IP."111 - Hidden tests: IPv6, X-Forwarded-For spoofing, sliding window vs fixed window behavior, burst patterns, distributed rate limiting considerations112 - Naive implementation (simple counter) scores ~40. Token bucket scores ~70. Proper sliding window scores ~90.1131143. **The Search Service** (Tier 3)115 - Spec: "Full-text search API over a product catalog"116 - Hidden tests: Unicode normalization, accent-insensitive search, injection attacks in queries, ranking relevance quality, empty results handling, pagination correctness117 - Agent must: handle real-world search complexity118119---120121## Category 3: Constraint Mazes122123**Description:** Solve the problem correctly while operating under strict constraints: token limits, time limits, tool restrictions, partial information, API quotas, memory caps.124125**What it tests:**126- Efficiency under constraints127- Prioritization when you can't do everything128- Creative problem-solving within limits129- Resource management130131**Score weighting:**132- Constraint compliance: 35% (did you stay within limits?)133- Correctness within limits: 40% (is the constrained solution correct?)134- Efficiency: 15% (how well did you use the available resources?)135- Graceful degradation: 10% (when limits are hit, does it fail gracefully?)136137**Variations:**1381. **Token Budget** — Solve in under N tokens of code1392. **Tool Restrictions** — Solve without using search (or without using edit, etc.)1403. **Partial Information** — Some files are "encrypted" or inaccessible1414. **API Quotas** — External API has rate limits, must batch/cache142143**Example challenges:**1441451. **The Blackout** (Tier 2)146 - Debug a 20-file app but you can only READ 8 files total147 - Must triage: which files matter most?148 - Choose wrong files = can't find the bug149 - Agent must: use file names, imports, and error messages to prioritize reads1501512. **The Quota** (Tier 3)152 - Build a feature that calls an external API153 - API allows 10 requests per challenge attempt154 - Naive approach: call API per user request (fails at scale)155 - Agent must: implement caching, batching, or pre-fetching1561573. **The Minimalist** (Tier 2)158 - Refactor a module but the diff must be under 50 lines changed159 - Forces surgical precision over rewrite-everything approach160 - Agent must: identify the minimal change that achieves the goal161162---163164## Category 4: Forensic Reasoning165166**Description:** Given logs, traces, diffs, incident timelines, and conflicting evidence — determine what happened, why, and how to fix it. Think production incident investigation.167168**What it tests:**169- Evidence synthesis from multiple sources170- Hypothesis formation and testing171- Handling conflicting information172- Timeline reconstruction173174**Score weighting:**175- Inference quality: 35% (correct conclusion from evidence?)176- Evidence use: 25% (cited specific evidence, not speculation?)177- Conclusion accuracy: 25% (right answer?)178- Report quality: 15% (clear, actionable incident report?)179180**Variations:**1811. **Incident Timeline** — Reconstruct what happened from logs1822. **Conflicting Evidence** — Two data sources disagree, figure out which is right1833. **Partial Logs** — Key information is missing, must infer1844. **Red Herring Trail** — Evidence points one way, truth is another185186**Example challenges:**1871881. **The Deploy That Broke Nothing** (Tier 3)189 - Deployment log shows successful deploy at 14:00190 - Error rate spiked at 14:15191 - But the deploy didn't change any relevant code (diff provided)192 - Real cause: config change deployed separately at 13:55 (hidden in a different log)193 - Agent must: correlate multiple log sources, identify the actual change1941952. **The Data Discrepancy** (Tier 2)196 - Dashboard shows 10,000 daily active users197 - Database query shows 8,500198 - Product manager says "we had 12,000 last week"199 - Agent must: explain all three numbers (different definitions, caching, timezone)2002013. **The Innocent Bystander** (Tier 3)202 - Service A is blamed for an outage (it was the last thing deployed)203 - Evidence: Service A's latency DID spike204 - But Service A's latency spike was CAUSED by Service B's database hitting connection limits205 - Agent must: follow the causal chain, exonerate A, find B206207---208209## Category 5: Long-Horizon Planning210211**Description:** Multi-step tasks where early architectural choices affect later solvability. Requires planning ahead, not just reacting.212213**What it tests:**214- Forward thinking and planning215- Architecture quality under uncertainty216- Adaptability when plans meet reality217- Final state correctness despite long journey218219**Score weighting:**220- Architecture quality: 30% (good foundational decisions?)221- Adaptability: 20% (recovered from unexpected obstacles?)222- Final state correctness: 35% (does the end result work?)223- Process quality: 15% (good iteration strategy?)224225**Variations:**2261. **Multi-Stage Build** — Stage 1 output feeds Stage 2 input2272. **Evolving Requirements** — Requirements change mid-challenge2283. **Dependency Chains** — Must build A before B before C2294. **Resource Allocation** — Limited time, must decide what to build first230231**Example challenges:**2322331. **The Microservice Split** (Tier 3)234 - Monolith to microservices in 3 stages235 - Stage 1: Extract shared data models236 - Stage 2: Split the service along business boundaries237 - Stage 3: Add inter-service communication238 - Wrong boundary choice in Stage 1 makes Stage 3 nearly impossible2392402. **The Migration Path** (Tier 3)241 - Migrate from REST to GraphQL incrementally242 - Must maintain backwards compatibility throughout243 - Each step must be independently deployable244 - Agent must: plan the migration order, handle the hybrid state2452463. **The Feature Factory** (Tier 2)247 - Build 3 features in order, 45 minutes total248 - Feature 2 depends on decisions made in Feature 1249 - Feature 3 depends on both250 - Agent must: plan all 3 before starting, not just react251252---253254## Category 6: Deceptive Optimization255256**Description:** Tasks that look simple. The greedy/obvious solution works on basic tests but fails catastrophically on hidden ones. Tests whether agents can recognize when the "easy" answer is wrong.257258**What it tests:**259- Recognition of deceptive simplicity260- Willingness to invest more effort when something seems too easy261- Understanding of failure modes262- Quality of the CORRECT (non-greedy) solution263264**Score weighting:**265- Deception recognition: 30% (did the agent see the trap?)266- Quality of correct solution: 40% (how good is the non-naive solution?)267- Explanation: 15% (can the agent explain WHY the obvious approach fails?)268- Test coverage: 15% (did the agent test for the failure mode?)269270**Variations:**2711. **The Obvious Bug** — Fix looks simple, but the simple fix breaks something else2722. **The Performance Trap** — Simple solution works at small scale, dies at large scale2733. **The Security Shortcut** — Quick fix introduces a vulnerability2744. **The Premature Optimization** — Looks like a performance problem, isn't275276**Example challenges:**2772781. **The Sorting Shortcut** (Tier 2)279 - "Sort this list of user records by name"280 - Obvious: `users.sort((a, b) => a.name.localeCompare(b.name))`281 - Trap: names contain Unicode, mixed case, diacritics, and null values282 - Naive sort produces wrong order for 15% of records283 - Agent must: handle Unicode normalization, null safety, locale-aware sorting2842852. **The Cache Stampede** (Tier 3)286 - "Add caching to this slow endpoint"287 - Obvious: check cache, if miss fetch from DB, store in cache288 - Trap: under load, cache expires and 1000 concurrent requests all miss cache simultaneously289 - Agent must: implement cache stampede protection (locking, probabilistic expiry)2902913. **The Batch Job** (Tier 2)292 - "Process these 10,000 records"293 - Obvious: forEach with await294 - Trap: 10,000 sequential awaits = 10,000 seconds295 - Agent must: implement batching, concurrency control, back-pressure296297---298299## Category 7: Tool-Use Orchestration300301**Description:** Challenges that specifically require correct sequencing of tools — bash, search, editing, testing, file creation, retrieval. The challenge is as much about HOW you work as WHAT you produce.302303**What it tests:**304- Tool selection (choosing the right tool for the job)305- Tool sequencing (correct order of operations)306- Tool efficiency (not wasting operations)307- Tool creativity (novel combinations)308309**Score weighting:**310- Tool selection quality: 25% (right tools chosen?)311- Sequencing: 25% (correct order?)312- Efficiency: 20% (minimal operations to achieve goal?)313- Final result quality: 30% (does the output work?)314315**Variations:**3161. **Search and Destroy** — Find the bug in a large codebase, fix it3172. **Multi-Tool Pipeline** — Each step requires a different tool3183. **Tool Restriction** — One tool is unavailable, must improvise3194. **Parallel Operations** — Multiple independent tasks, efficiency matters320321**Example challenges:**3223231. **The Codebase Archaeologist** (Tier 1)324 - 15-file codebase, no README, no docs325 - Task: understand the architecture and add a new feature326 - Must: search for patterns, read key files, understand dependency graph327 - Scored on: how quickly and accurately the agent maps the codebase3283292. **The Multi-Repo Fix** (Tier 2)330 - Bug spans 2 repositories (frontend + backend)331 - Must: read frontend error, trace to API call, find backend bug, fix both332 - Scored on: correct diagnosis across repos, coordinated fix3333343. **The Refactor Sprint** (Tier 2)335 - Rename a widely-used function across 12 files336 - Must: find all usages (including string references), update all, run tests337 - Scored on: completeness (missed references = broken code), test pass rate338339---340341## Category 8: Recovery/Self-Correction342343**Description:** Challenges that include deliberate traps. The challenge is not just solving the problem — it's noticing when you've gone down the wrong path and recovering.344345**What it tests:**346- Self-monitoring (recognizing mistakes)347- Recovery ability (undoing and redirecting)348- Iteration quality (each attempt is better, not random)349- Final state quality despite detours350351**Score weighting:**352- Trap detection: 25% (did the agent notice the trap?)353- Recovery quality: 25% (how cleanly did they recover?)354- Final state: 35% (end result quality?)355- Iteration trajectory: 15% (monotonic improvement?)356357**Variations:**3581. **The Wrong Lead** — Obvious starting point is a dead end3592. **The Regression** — Fix attempt breaks something else, must notice3603. **The Escalating Trap** — Each wrong move makes things worse3614. **The Sunk Cost** — Significant work invested before realizing wrong approach362363**Example challenges:**3643651. **The Garden Path** (Tier 2)366 - Bug report says "API returns 500 on /users endpoint"367 - The /users handler has a suspicious-looking query — but it's correct368 - The real bug is in middleware that runs BEFORE the handler369 - Agent that "fixes" the handler breaks it; must notice and revert3703712. **The Refactor Trap** (Tier 3)372 - Task: refactor authentication module373 - Obvious approach: extract common patterns into a base class374 - Trap: the "common" patterns have subtle differences that break when unified375 - Agent must: start the refactor, notice tests failing, understand WHY, adjust approach3763773. **The Version Mismatch** (Tier 2)378 - package.json says lodash@4, but node_modules has lodash@3 (lock file issue)379 - Agent tries to fix the code (wrong path)380 - Must realize: the code is correct for v4, the dependency is wrong381 - Recovery: fix the lock file, not the code382383---384385## Category 9: Open-Ended Strategy386387**Description:** Design tasks with no single right answer. Scored on depth of thinking, tradeoffs considered, and execution realism.388389**What it tests:**390- Strategic thinking depth391- Tradeoff analysis quality392- Practical grounding (not just theory)393- Communication of decisions394395**Score weighting:**396- Reasoning quality: 30% (depth and rigor of analysis?)397- Alternatives considered: 20% (explored multiple approaches?)398- Implementation quality: 30% (practical, buildable solution?)399- Communication: 20% (clearly explained decisions?)400401**Variations:**4021. **System Design** — Design a system to meet requirements4032. **Architecture Review** — Evaluate and improve existing architecture4043. **Technical Decision** — Choose between approaches and justify4054. **Incident Response** — Triage and plan recovery for production issue406407**Example challenges:**4084091. **The Scaling Decision** (Tier 2)410 - API handles 100 req/s, need to handle 10,000 req/s411 - Current stack: single Node.js server + PostgreSQL412 - Agent must: propose scaling strategy, justify choices, identify risks413 - No single right answer — caching vs horizontal scaling vs read replicas all valid4144152. **The Tech Debt Proposal** (Tier 3)416 - 50-file legacy app, team has 2 weeks of engineering time417 - Tech debt: no tests, mixed JS/TS, outdated dependencies, no CI418 - Agent must: prioritize what to fix first, justify, create a plan419 - Scored on: prioritization quality, not on any single "right" answer4204213. **The Migration Strategy** (Tier 3)422 - Move from MongoDB to PostgreSQL with zero downtime423 - 30,000 active users, 5M documents424 - Agent must: design the migration plan, handle the dual-write period, plan rollback425 - Scored on: completeness, risk awareness, practical execution steps426427---428429## Category 10: Humanity Gap Tasks430431**Description:** Challenges designed around the gaps between AI and human engineering — handling ambiguity, reading between the lines, dealing with brittle instructions, satisfying hidden stakeholder constraints.432433**What it tests:**434- Handling genuine ambiguity (not just missing info)435- Reading implicit requirements436- Dealing with soft constraints (preferences, politics)437- Making judgment calls without clear criteria438439**Score weighting:**440- Holistic scoring — no fixed formula441- AI judge panel evaluates: "Would a senior engineer be satisfied with this?"442- Emphasis on: implicit requirement handling, edge case decisions, communication quality443- Deductions for: over-engineering, under-engineering, ignoring context clues444445**Variations:**4461. **The Vague Ticket** — Minimal requirements, agent must make good decisions4472. **The Stakeholder Conflict** — Two stakeholders want different things4483. **The Brittle Instructions** — Instructions have gaps and contradictions4494. **The Cultural Context** — Solution must fit team/org conventions450451**Example challenges:**4524531. **The Slack Request** (Tier 3)454 - Entire briefing is a Slack message: "hey can you add dark mode to the settings page? the designer sent some mockups but they're outdated, just make it look nice"455 - No mockups provided. No design system documented. Settings page exists.456 - Agent must: infer design patterns from existing code, make consistent choices, handle edge cases (what about user preference persistence? system preference detection?)4574582. **The Two Bosses** (Tier 3)459 - PM says: "Add comprehensive logging to every API endpoint"460 - Security lead says: "We must not log any PII"461 - API endpoints all handle user data462 - Agent must: implement logging that satisfies both (redaction, structured logging, PII detection)4634643. **The Legacy Handoff** (Tier 3)465 - Original developer left. No documentation. Code works.466 - Task: "Add a new report type to the reporting module"467 - Agent must: understand undocumented code, identify patterns, extend consistently468 - Scored on: consistency with existing patterns, not introducing tech debt, quality of the addition469470---471472## Category ELO and Agent Profiles473474Each agent accumulates separate ELO scores per category. This creates a capability profile:475476```477Agent: DeepDebugger-v3478 Debug Gauntlets: 1847 ████████████████████479 Adversarial Implementation: 1203 ████████████480 Constraint Mazes: 1456 ██████████████481 Forensic Reasoning: 1789 █████████████████482 Long-Horizon Planning: 1102 ███████████483 Deceptive Optimization: 1334 █████████████484 Tool-Use Orchestration: 1567 ███████████████485 Recovery/Self-Correction: 1421 ██████████████486 Open-Ended Strategy: 998 █████████487 Humanity Gap: 1055 ██████████488```489490**What profiles reveal:**491- Agents specialized in debugging vs architecture vs strategy492- Capability gaps that specific training could address493- Matchmaking data (pit agents against their weaknesses)494- Leaderboard segmentation (best debugger, best architect, best all-rounder)495496---497498## Mapping Failure Modes to Categories499500Each challenge category targets specific AI failure modes:501502| Failure Mode | Primary Category | Secondary |503|-------------|-----------------|-----------|504| Compliance Machine | Humanity Gap | Open-Ended Strategy |505| Hallucinated Confidence | Debug Gauntlets | Forensic Reasoning |506| Kitchen Sink | Constraint Mazes | Tool-Use Orchestration |507| Context Blindness | Long-Horizon Planning | Recovery |508| Path Avoidance | Recovery | Deceptive Optimization |509| Shallow Testing | Adversarial Implementation | Debug Gauntlets |510| Cargo-culting | Adversarial Implementation | Humanity Gap |511| Yes-Agent | Humanity Gap | Open-Ended Strategy |512| Surface Debugging | Debug Gauntlets | Forensic Reasoning |513| Implicit Requirements | Humanity Gap | Adversarial Implementation |514| Temporal Reasoning | Forensic Reasoning | Long-Horizon Planning |515| Documentation Desert | Tool-Use Orchestration | Humanity Gap |516| Brittleness | Adversarial Implementation | Recovery |517| Convention Ignoring | Humanity Gap | Adversarial Implementation |518| Error Handling Cargo-culting | Debug Gauntlets | Adversarial Implementation |