BullshitBench Technical Guide
This guide is for maintainers and contributors working on benchmark operations and data publishing.
Pipeline Overview
The end-to-end flow is:
collectgrade-panelpublish_latest_to_viewer.sh(updatesdata/latest)
Run all three steps with:
./scripts/run_end_to_end.sh
Serve locally after publish:
./scripts/run_end_to_end.sh --serve --port 8877
Then open http://localhost:8877/viewer/index.html.
Publish Existing Run Artifacts
Use this when you already have run outputs and only want to refresh data/latest:
./scripts/publish_latest_to_viewer.sh \
--responses-file <path/to/responses.jsonl> \
--collection-stats <path/to/collection_stats.json> \
--panel-summary <path/to/panel_summary.json> \
--aggregate-summary <path/to/aggregate_summary.json> \
--aggregate-rows <path/to/aggregate.jsonl>
The publish step strips local-machine path fields from public artifacts.
Launch-Date Metadata Pipeline
Build model launch-date inventory/buckets and export review/candidate/canonical launch datasets:
./scripts/model_launch_pipeline.py run
This writes:
data/model_metadata/tested_models_inventory.csvdata/model_metadata/model_buckets.csvdata/model_metadata/model_launch_sources.csv(template if missing)data/model_metadata/model_launch_collection.csvdata/model_metadata/model_launch_judged.csvdata/model_metadata/model_launch_attempts.csvdata/model_metadata/model_launch_dates_review.csvdata/model_metadata/model_launch_dates_candidates.csvdata/model_metadata/model_launch_dates.csv(canonical accepted rows)
Publishing also exports:
data/latest/model_launch_dates.csvdata/latest/leaderboard_with_launch.csv
Current Config Notes
- Main config:
config.json - Question set:
questions.json - Recent published snapshot includes
openai/gpt-5.2-codexandopenai/gpt-5.3-codexvariants across reasoning levels (low,high,xhigh) indata/latest/leaderboard.csv.
Repository Layout
scripts/openrouter_benchmark.py: core CLI (collect,grade,grade-panel,aggregate,report)scripts/run_end_to_end.sh: one-command pipeline runnerscripts/publish_latest_to_viewer.sh: publish run outputs intodata/latestscripts/cleanup_generated_outputs.sh: remove generated local artifactsscripts/model_launch_pipeline.py: launch-date collection/judging pipelineviewer/index.html: canonical interactive viewerdata/latest/*: canonical published dataset used by the viewerruns/*: local run history
Published Dataset Files
data/latest contains:
responses.jsonlcollection_stats.jsonpanel_summary.jsonaggregate_summary.jsonaggregate.jsonlleaderboard.csvleaderboard_with_launch.csvmodel_launch_dates.csvmanifest.json
Environment
Required:
OPENROUTER_API_KEY
Optional:
OPENROUTER_REFEREROPENROUTER_APP_NAME