Flask Eval

Evaluates LLMs on fine-grained alignment capabilities by decomposing instruction-following performance into 12 sub-skills across four domains (Logical Thinking, Background Knowledge, Problem Handling, User Alignment). It measures how well models adhere to specific quality criteria like factuality, logical robustness, and harmlessness on a per-instance basis. Use when the user wants to benchmark on FLASK, or asks about evaluating this task. Reports FLASK skill score.

qhjqhj00 e17c1ee 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/flask-eval commit e17c1ee1c7

Frequently asked questions

npx skillmds add qhjqhj00/flask-eval