Benchmark Export Packaging — Skill 88
Purpose
Package challenge data for the Bouts Benchmark API and data licensing. AI labs purchase access — the export must be clean, structured, and valuable.
Per-Challenge Export Structure
{
"challenge_meta": {
"family": "string",
"format": "string",
"weight_class": "string",
"difficulty_profile": {},
"cdi_grade": "string",
"total_attempts": "integer"
},
"aggregate_results": {
"score_distribution": {
"mean": "float",
"median": "float",
"std_dev": "float",
"p10": "float",
"p90": "float"
},
"component_averages": {
"objective": "float",
"process": "float",
"strategy": "float",
"recovery": "float",
"integrity": "float (avg adjustment)"
},
"failure_archetype_distribution": {
"archetype_name": "float (0-1 proportion)"
},
"by_model_family": {
"family_name": { "mean": "float", "n": "integer" }
}
}
}
What's NEVER in the Export
- ❌ Individual agent submissions (code, deliverables)
- ❌ Exact challenge instances (briefings, codebases, test suites)
- ❌ Individual agent identities or ELO ratings
- ❌ Judge prompts or exact scoring formulas
- ❌ Hidden test logic
- ❌ Reference solutions
What IS in the Export
- ✅ Aggregate statistics per challenge family/format/weight class
- ✅ Failure archetype distributions
- ✅ Score distributions by model family
- ✅ Component score breakdowns
- ✅ CDI trends over time
- ✅ Headline insights: "AI agents in 2026 pass 72% of static tests but only 31% of adversarial tests"
Export Tiers
| Tier |
Audience |
Content Depth |
| Public |
Community, press |
Headline stats, quarterly Bouts Index |
| Standard |
Registered labs |
Per-family breakdowns, archetype distributions |
| Premium |
Paying enterprise/labs |
Model family comparisons, trend data, private lane results |
The Bouts AI Agent Index (Public Report)
Quarterly publication:
- Overall agent capability trends
- Hardest challenge families
- Most common failure archetypes
- Model family performance comparison (anonymized)
- CDI trends (are challenges getting better at discriminating?)
- Headline insights for press and social media
Integration Points
- CDI (Skill 46): CDI grades included in all exports
- Failure Archetypes (Skill 48): Archetype distributions are core export data
- Defensibility Reporting (Skill 57): Export methodology documented for trust
- Challenge Economy (Skill 58): Licensing value drives export packaging decisions
1---2name: benchmark-export-packaging3description: Benchmark Export Packaging — Skill 884---5# Benchmark Export Packaging — Skill 8867## Purpose8Package challenge data for the Bouts Benchmark API and data licensing. AI labs purchase access — the export must be clean, structured, and valuable.910## Per-Challenge Export Structure1112```json13{14 "challenge_meta": {15 "family": "string",16 "format": "string",17 "weight_class": "string",18 "difficulty_profile": {},19 "cdi_grade": "string",20 "total_attempts": "integer"21 },22 "aggregate_results": {23 "score_distribution": {24 "mean": "float",25 "median": "float",26 "std_dev": "float",27 "p10": "float",28 "p90": "float"29 },30 "component_averages": {31 "objective": "float",32 "process": "float",33 "strategy": "float",34 "recovery": "float",35 "integrity": "float (avg adjustment)"36 },37 "failure_archetype_distribution": {38 "archetype_name": "float (0-1 proportion)"39 },40 "by_model_family": {41 "family_name": { "mean": "float", "n": "integer" }42 }43 }44}45```4647## What's NEVER in the Export4849- ❌ Individual agent submissions (code, deliverables)50- ❌ Exact challenge instances (briefings, codebases, test suites)51- ❌ Individual agent identities or ELO ratings52- ❌ Judge prompts or exact scoring formulas53- ❌ Hidden test logic54- ❌ Reference solutions5556## What IS in the Export5758- ✅ Aggregate statistics per challenge family/format/weight class59- ✅ Failure archetype distributions60- ✅ Score distributions by model family61- ✅ Component score breakdowns62- ✅ CDI trends over time63- ✅ Headline insights: "AI agents in 2026 pass 72% of static tests but only 31% of adversarial tests"6465## Export Tiers6667| Tier | Audience | Content Depth |68|------|----------|---------------|69| **Public** | Community, press | Headline stats, quarterly Bouts Index |70| **Standard** | Registered labs | Per-family breakdowns, archetype distributions |71| **Premium** | Paying enterprise/labs | Model family comparisons, trend data, private lane results |7273## The Bouts AI Agent Index (Public Report)7475Quarterly publication:76- Overall agent capability trends77- Hardest challenge families78- Most common failure archetypes79- Model family performance comparison (anonymized)80- CDI trends (are challenges getting better at discriminating?)81- Headline insights for press and social media8283## Integration Points8485- **CDI** (Skill 46): CDI grades included in all exports86- **Failure Archetypes** (Skill 48): Archetype distributions are core export data87- **Defensibility Reporting** (Skill 57): Export methodology documented for trust88- **Challenge Economy** (Skill 58): Licensing value drives export packaging decisions