Together Evaluations

LLM-as-a-judge evaluation framework on Together AI. Classify, score, and compare model outputs, select judge models, use external-provider judges or targets, poll results and download reports. Reach for it whenever the user wants to benchmark outputs, grade responses, compare A/B variants, or operationalize automated evaluations.

togethercomputer fa3473d 5 files · 66.6 KB Updated

File contents

togethercomputer/skills/tree/main/skills/together-evaluations commit fa3473d8d3

Frequently asked questions

npx skillmds@latest add togethercomputer/together-evaluations