Verisoftbench Eval

This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.

qhjqhj00 1f772f2 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/verisoftbench-eval commit 1f772f2c24

Frequently asked questions

npx skillmds add qhjqhj00/verisoftbench-eval