Agenticcache Eval

Evaluates the ability of embodied multi-agent systems to execute long-horizon, coordinated tasks efficiently using cache-driven asynchronous planning. It probes how well agents can reuse cached plan transitions to reduce LLM inference latency and token costs while maintaining high task success rates across diverse 3D simulation environments. Use when the user wants to benchmark on TDW-MAT, TDW-COOK, TDW-GAME, BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate.

qhjqhj00 4702630 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agenticcache-eval commit 47026301f6

Frequently asked questions

npx skillmds add qhjqhj00/agenticcache-eval