Omni AI Eval

Evaluate Omni AI accuracy using Omni's built-in eval system — define a prompt set, run a judged eval against a model (or branch), and read the accuracy-judge verdicts. Use this skill whenever someone wants to evaluate Omni AI, benchmark Blobby, run regression tests, compare AI output across branches or model-context changes, measure AI quality, run A/B tests on model changes, assess the impact of an ai_context or modeling change, or any variant of "run evals", "test Blobby", "benchmark query generation", "compare AI results", "regression test", "how accurate is the AI", or "measure the impact of my changes".

exploreomni c8f636f 3 files · 18.9 KB Updated

File contents

exploreomni/omni-agent-skills/tree/main/skills/omni-ai-eval commit c8f636f7b7

Frequently asked questions

npx skillmds@latest add exploreomni/omni-ai-eval