Context Eval

Evaluate whether a context engineering harness actually improves agent outcomes. Use when the user wants to measure, benchmark, compare, or validate any context artifact — rules, instructions, guidelines, docs, retrieval pipelines, tool setups. Triggers on "does this context help?", "benchmark my context", "evaluate my prompt", "A/B test my context", "is this worth the tokens?", "eval my context", or when someone wants empirical evidence their context engineering produces better outcomes than baseline.

andurilcode 4c037af 12 files · 107.2 KB Updated

File contents

andurilcode/craftwork/tree/main/skills/context-eval commit 4c037af380

Frequently asked questions

npx skillmds@latest add andurilcode/context-eval