Nl2sh Eval

This benchmark evaluates the ability of large language models to translate natural language instructions into executable Bash commands. It probes functional correctness by comparing model-generated commands against ground-truth commands using a functional equivalence heuristic that combines command execution with LLM-based output analysis. Use when the user wants to benchmark on NL2SH, InterCode-ALFA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 0021fbd 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/nl2sh-eval commit 0021fbdd0b

Frequently asked questions

npx skillmds add qhjqhj00/nl2sh-eval