Together Evaluations

LLM-as-a-judge evaluation framework on Together AI. Classify, score, and compare model outputs, select judge models, use external-provider judges or targets, poll results and download reports. Reach for it whenever the user wants to benchmark outputs, grade responses, compare A/B variants, or operationalize automated evaluations.

zainhas 18d9340 5 files · 62.2 KB Updated

File contents

zainhas/togetherai-skills/tree/main/skills/together-evaluations commit 18d9340166

Frequently asked questions

npx skillmds add zainhas/together-evaluations