LLM Council Cost

Cuts the token spend and latency of a multi-model pipeline - route-first gating so the council fires only on hard queries, confidence-based escalation, shared-prefix caching, batching, tail-latency budgets. Use when calling several models per request costs too much or takes too long, or to cache across member calls. Not for whether to adopt a council, or for aggregation.

Paldom Updated

File contents

Paldom/llm-council-skills/tree/main/skills/llm-council-cost commit b16012997f

Frequently asked questions

npx skillmds@latest add paldom/llm-council-cost