Evaluation Protocol Comparison Tactic
Compare how different papers implement the same benchmark to expose hidden protocol variance that undermines cross-paper score comparability.
Stages
Stage 1: Paper Collection (Same Benchmark)
Collect 10-15 papers that report results on the target benchmark:
- Prioritize diversity: different labs, years, model families
- Include the original benchmark paper as reference protocol
- Include papers from different venues (top conferences, workshops, preprints)
- Search via dare-ss (ss_relevance_search) and dare-scholar (paper_searching)
Search queries: "[benchmark name] evaluation", "[benchmark name] results", "[benchmark name] state-of-the-art"
Stage 2: Protocol Element Extraction
For each paper, run protocol-element-extraction SOP to extract:
| Element Category |
Specific Parameters |
| Data |
Split version, subset selection, preprocessing, filtering |
| Prompting |
Template format, few-shot examples (count, selection), instruction wording |
| Generation |
Decoding strategy, temperature, top-p/top-k, max tokens, stop criteria |
| Evaluation |
Metric implementation, postprocessing, normalization, scoring script version |
| Infrastructure |
Framework, precision (fp16/bf16/fp32), batch size, hardware |
Stage 3: Difference Matrix Construction
Build a comparison matrix:
- Rows = protocol elements
- Columns = papers
- Cells = specific value used
- Highlight deviations from original protocol
Compute per-element variance:
- None: All papers use identical value
- Low: Minor variations (e.g., different random seeds)
- Medium: Substantive differences (e.g., different few-shot examples)
- High: Fundamental disagreements (e.g., different splits, different metrics)
- Extreme: Papers appear to evaluate different things under same name
Stage 4: Impact Assessment
For each high-variance element:
- Search for ablation studies showing impact of that element
- Estimate score range attributable to protocol choice vs model quality
- Identify which protocol choices systematically favor certain model families
- Flag "protocol p-hacking" — suspicious correlation between protocol choice and reported improvement
Output
protocol_comparison:
benchmark: string
papers_compared: int
reference_protocol: string # original benchmark paper
difference_matrix:
- element: string
category: data|prompting|generation|evaluation|infrastructure
variance_level: none|low|medium|high|extreme
values: list[{paper, value}]
impact_estimate: string
highest_variance_elements:
- element: string
score_impact: string
favors: string # which model family benefits
protocol_p_hacking_flags:
- paper: string
suspicious_choice: string
benefit: string
cross_paper_comparability: high|moderate|low|unreliable
standardization_recommendations:
- element: string
recommended_value: string
rationale: string
Yield Report
| Metric |
Minimum |
| Papers compared |
8 |
| Protocol elements extracted per paper |
10 |
| High-variance elements identified |
2 |
| Impact estimates produced |
3 |
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP |
When to use |
| protocol-element-extraction |
Extract evaluation protocol parameters from papers |
1---2name: evaluation-protocol-comparison3description: Compare implementation differences of same benchmark across papers4---56# Evaluation Protocol Comparison Tactic78Compare how different papers implement the same benchmark to expose hidden protocol variance that undermines cross-paper score comparability.910## Stages1112### Stage 1: Paper Collection (Same Benchmark)1314Collect 10-15 papers that report results on the target benchmark:15- Prioritize diversity: different labs, years, model families16- Include the original benchmark paper as reference protocol17- Include papers from different venues (top conferences, workshops, preprints)18- Search via dare-ss (ss_relevance_search) and dare-scholar (paper_searching)1920**Search queries**: "[benchmark name] evaluation", "[benchmark name] results", "[benchmark name] state-of-the-art"2122### Stage 2: Protocol Element Extraction2324For each paper, run protocol-element-extraction SOP to extract:2526| Element Category | Specific Parameters |27|-----------------|-------------------|28| **Data** | Split version, subset selection, preprocessing, filtering |29| **Prompting** | Template format, few-shot examples (count, selection), instruction wording |30| **Generation** | Decoding strategy, temperature, top-p/top-k, max tokens, stop criteria |31| **Evaluation** | Metric implementation, postprocessing, normalization, scoring script version |32| **Infrastructure** | Framework, precision (fp16/bf16/fp32), batch size, hardware |3334### Stage 3: Difference Matrix Construction3536Build a comparison matrix:37- Rows = protocol elements38- Columns = papers39- Cells = specific value used40- Highlight deviations from original protocol4142Compute per-element variance:43- **None**: All papers use identical value44- **Low**: Minor variations (e.g., different random seeds)45- **Medium**: Substantive differences (e.g., different few-shot examples)46- **High**: Fundamental disagreements (e.g., different splits, different metrics)47- **Extreme**: Papers appear to evaluate different things under same name4849### Stage 4: Impact Assessment5051For each high-variance element:52- Search for ablation studies showing impact of that element53- Estimate score range attributable to protocol choice vs model quality54- Identify which protocol choices systematically favor certain model families55- Flag "protocol p-hacking" — suspicious correlation between protocol choice and reported improvement5657## Output5859```yaml60protocol_comparison:61 benchmark: string62 papers_compared: int63 reference_protocol: string # original benchmark paper64 difference_matrix:65 - element: string66 category: data|prompting|generation|evaluation|infrastructure67 variance_level: none|low|medium|high|extreme68 values: list[{paper, value}]69 impact_estimate: string70 highest_variance_elements:71 - element: string72 score_impact: string73 favors: string # which model family benefits74 protocol_p_hacking_flags:75 - paper: string76 suspicious_choice: string77 benefit: string78 cross_paper_comparability: high|moderate|low|unreliable79 standardization_recommendations:80 - element: string81 recommended_value: string82 rationale: string83```8485## Yield Report8687| Metric | Minimum |88|--------|---------|89| Papers compared | 8 |90| Protocol elements extracted per paper | 10 |91| High-variance elements identified | 2 |92| Impact estimates produced | 3 |9394<!-- BEGIN available-tables (generated) -->9596## Available SOPs9798Optional, no fixed order; the final leaf is always a sop.99100| SOP | When to use |101| --- | --- |102| protocol-element-extraction | Extract evaluation protocol parameters from papers |103104<!-- END available-tables (generated) -->