Toolemu Eval

Evaluates the safety and helpfulness of language model agents interacting with tools in a simulated environment. It measures how well an automated emulator and evaluator align with human judgments, and quantifies agent failure rates under standard and adversarial conditions. Use when the user wants to benchmark on ToolEmu Agent Trajectories, or asks about evaluating this task. Reports Cohen's κ (Quadratic-weighted).

qhjqhj00 ef6cf99 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/toolemu-eval commit ef6cf99e38

Frequently asked questions

npx skillmds add qhjqhj00/toolemu-eval