Veriequivbench Eval

Evaluates an LLM's ability to generate formally verifiable code that aligns with natural language problem descriptions and passes unit tests. It probes complex algorithmic reasoning and code-specification alignment without requiring manual ground-truth specifications. Use when the user wants to benchmark on VeriEquivBench, or asks about evaluating this task. Reports equivalence_score.

qhjqhj00 e9d9c85 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/veriequivbench-eval commit e9d9c858d5

Frequently asked questions

npx skillmds add qhjqhj00/veriequivbench-eval