Flageval Textual Eval

Evaluates large reasoning models on automatically verifiable textual problem-solving tasks, including academic coursework, word puzzles, cipher deciphering, and algorithmic coding. It probes the models' ability to follow instructions, perform logical deduction, and produce correctly formatted final answers under varying reasoning effort settings. Use when the user wants to benchmark on FlagEval Textual, or asks about evaluating this task. Reports accuracy.

qhjqhj00 937aceb 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/flageval-textual-eval commit 937aceb47b

Frequently asked questions

npx skillmds add qhjqhj00/flageval-textual-eval