Deterministic Packaging — Skill 90
Purpose
Ensure challenges are reproducible and auditable. Same template + same seed + same mutation parameters = identical challenge.
Reproducibility Requirements
Every challenge instance must be reproducible from its seed:
| Input | Description |
|---|---|
| template_id | Which template was used |
| template_version | Exact version of the template |
| generation_seed | 8-char hex seed |
| mutation_params | Full mutation parameters object |
| generation_timestamp | When it was generated |
| gauntlet_version | Which version of Gauntlet generated it |
Same inputs = identical outputs. This enables: re-running calibration, auditing disputed scores, reproducing bugs, verifying tampering.
Audit Trail Structure
{
"audit": {
"instance_id": "BOUTS-2026-XXXX",
"template_id": "tmpl-{family}-v{N}",
"template_version": "integer",
"generation_seed": "8-char hex",
"mutation_params": {
"framework": "string",
"database": "string",
"bug_types": ["string"],
"red_herring_style": "string",
"log_noise_level": "string",
"mutations_applied": ["mutation_type"]
},
"generated_by": "gauntlet-v{version}",
"generated_at": "ISO-8601",
"calibration_run_id": "string",
"calibration_result": "passed | failed",
"published_at": "ISO-8601 | null",
"retired_at": "ISO-8601 | null",
"quarantined_at": "ISO-8601 | null",
"content_hash": "sha256:..."
}
}
Content Hashing
- At generation: Hash the entire challenge package (workspace files + tests + rubrics)
- Store the hash in the audit trail
- Before each run: Verify hash matches (ensures nobody tampered post-publication)
- If hash mismatch: Quarantine immediately, alert ops
Version Pinning
- Record which version of Gauntlet generated each challenge
- If Gauntlet is updated, old challenges retain their original generation metadata
- New calibration runs after a Gauntlet update should verify old challenges still calibrate correctly
- If calibration drifts after update → investigate whether Gauntlet changes affected challenge quality
Audit Queries
The audit trail must support:
| Query | Purpose |
|---|---|
| "Show all instances from template X" | Template health analysis |
| "Show all instances generated by Gauntlet version Y" | Version impact analysis |
| "Show the full lineage of instance Z" | Dispute investigation |
| "Verify instance Z hasn't been modified" | Tamper detection |
| "Reproduce instance Z from seed" | Calibration re-run |
Integration Points
- Structured Output (Skill 77): Audit trail wraps the challenge JSON
- Calibration Packaging (Skill 81): Lineage stored in meta/lineage.json
- Challenge Genealogy (Skill 84): Audit data feeds genealogy analysis
- Defensibility Reporting (Skill 57): Audit trail is core defensibility evidence