Eval

AI/LLM evaluation. Benchmark creation, regression testing, statistical significance, LLM-as-judge, promptfoo.

arbazkhan971 6831aae 6.3 KB Updated

File contents

arbazkhan971/godmode/tree/main/skills/eval commit 6831aaec8b

Frequently asked questions

npx skillmds@latest add arbazkhan971/eval