Composition evaluation
Use this skill when a composition policy needs reproducible comparison against a single-pass baseline.
- Start from
benchmarks/composition-policy-benchmark.v1.jsonand preserve fixed tasks, settings, budgets, seeds, metrics, and thresholds. - Run
aiwg composition benchmark <manifest.json>; use--raw-outand--summary-outto retain both evidence layers. - Compare success-conditioned cost and latency, not unconditioned cheap failures. Review speed-of-accuracy and strict-LCM-versus-adaptive deltas.
- Require an independent or human evaluation path and inspect self-judge bias.
- Review every failure-injection outcome and recovery receipt.
- Keep synthetic conformance distinct from provider evidence. Do not open the claim gate without repeated trusted runs, independent evaluation, confidence intervals, and task-family replication.
Composition graphs remain flow.aiwg.io/v1alpha1 FlowGraph; do not introduce
a fourth-level DNS API group. Do not request or persist private chain-of-thought.
@implements #2118