Autoformalization Compile Eval

Probes a model's ability to translate informal natural language mathematical statements into syntactically and semantically valid formal code for theorem provers (Isabelle or Lean4). It measures how well the model captures formal syntax, type-checking rules, and prover-specific conventions without requiring proof generation. Use when the user wants to benchmark on miniF2F, ProofNet, or asks about evaluating this task. Reports Compilation rates (%).

qhjqhj00 470f30d 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/autoformalization-compile-eval commit 470f30db8e

Frequently asked questions

npx skillmds add qhjqhj00/autoformalization-compile-eval