LLM Eval Router

Shadow-test local Ollama models against a cloud baseline with a multi-judge ensemble. Automatically promotes models when statistically proven equivalent — reducing API costs with evidence, not hope.

reddinft Updated

File contents

reddinft/skill-llm-eval-router/tree/main/ commit 3cbc6ab202

Frequently asked questions

npx skillmds@latest add reddinft-skill-llm-eval-router/llm-eval-router