C3 Bench Eval

Evaluates the robustness and multi-tasking capabilities of LLM-based agents by probing their ability to handle complex tool dependencies, propagate hidden information across tasks, and maintain stable decision policies under dynamic, multi-round interactions. Use when the user wants to benchmark on C^3-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 37f7869 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/c3-bench-eval commit 37f7869f8f

Frequently asked questions

npx skillmds add qhjqhj00/c3-bench-eval