Results for “ratings”
2 skillsEval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
Agents
Evaluates execution transcripts and output files against a list of expectations, assigning pass/fail verdicts with cited evidence and critiquing the assertions themselves.
0 · bundle