Mint Eval

Evaluates large language models' ability to solve tasks using external tools across multiple interaction turns, and their capacity to leverage natural language feedback to improve performance. It also measures the rate of improvement per turn and identifies failure patterns like formatting issues or training data artifacts. Use when the user wants to benchmark on MINT, or asks about evaluating this task. Reports Success Rate (SR).

qhjqhj00 85c0c04 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mint-eval commit 85c0c040ba

Frequently asked questions

npx skillmds add qhjqhj00/mint-eval