Compass Judger Reward Based Evaluation

Build generalist judge models for evaluating LLM outputs using verifiable rewards and policy gradient training. Create a 7B model competitive with much larger judges through reward-guided optimization and critical thinking decomposition. Use when you need reliable automated evaluation of model outputs across diverse tasks.

adu2021 6b0a9a8 20.8 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/compass-judger-reward-based-evaluation commit 6b0a9a8818

Frequently asked questions

npx skillmds@latest add adu2021/compass-judger-reward-based-evaluation