guydav-restrictedpython-code-eval
Metric
guydav/restrictedpython_code_evalfrom the HuggingFaceevaluatelibrary.
When to invoke
User asks to compute guydav/restrictedpython_code_eval or wants HF evaluate's canonical version.
Recipe
import evaluate
metric = evaluate.load("guydav/restrictedpython_code_eval")
result = metric.compute(predictions=preds, references=refs)
print(result)
Don'ts
- Don't assume your in-house
guydav/restrictedpython_code_evalmatches HF — version conventions vary. - Many evaluate metrics have task-specific arguments (
average=,lang=,model_type=); read the metric card before reporting numbers.