Results for “hare-score”

15 skills
More results
jarbitechture
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
0
thedixitjain
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2
dylanckawalec
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
3
lucassantana-dev
rag-quality
Evaluate retrieval quality from the local RAG index
1 · bundle
sinhoneyy
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
11
alirezarezvani
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
qhjqhj00
eer
Compute the Equal Error Rate (EER) metric using torchmetrics for binary, multiclass, or multilabel classification tasks, with reference signatures and usage examples.
3
lingxling
arbor
Runs an autonomous optimization loop that iteratively improves an artifact against an objective and evaluator using Hypothesis Tree Refinement, with subagent executors in isolated git worktrees.
253 · bundle
doany-ai
happyhorse-1-0
Generate text-to-video with HappyHorse 1.0 on RunComfy. Documents HappyHorse 1.0's strengths (#1 on Artificial Analysis Video Arena, native 1080p with in-pass synchronized audio, multi-shot character consistency, 6-language prompt support), the duration / aspect-ratio / resolution schema, and when to route to Wan 2.7 / Seedance 2 / LTX 2 instead. Calls `runcomfy run happyhorse/happyhorse-1-0/text-to-video` through the local RunComfy CLI. Triggers on "happyhorse", "happy horse", "happyhorse 1.0", "happyhorse video", or any explicit ask to generate video with this model.
5
dvy1987
eval-judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
enuno
wolf-howl
Runs a nightly automated retrospective on autonomous trading strategy performance, computing win rates, fee drag, holding period buckets, direction bias, and producing data-driven improvement suggestions.
1 · bundle
qhjqhj00
spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
qhjqhj00
accuracy
Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
3
comeonoliver
songsee
Generates spectrograms and multi-panel audio feature visualizations from audio files via a command-line tool.
61