Reproducibility Repair Eval

Evaluates the ability of LLMs and AI agents to automatically repair broken R-based social science code and restore computational reproducibility. It probes how well different workflows handle varying error complexities and contextual information. Use when the user wants to benchmark on Custom R-based Social Science Code Dataset, or asks about evaluating this task. Reports reproduction_success_rate.

qhjqhj00 2d97b6f 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/reproducibility-repair-eval commit 2d97b6fadc

Frequently asked questions

npx skillmds add qhjqhj00/reproducibility-repair-eval