Results for “conviction-scoring”
12 skillsMore results
eval-judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
f1score
Compute the F1Score metric using torchmetrics when predictions and ground-truth labels are available.
3
accuracy
Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
3
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
11
cx-incentive-design
Use to design support incentives that improve behaviour without destroying the metric — pairing pay with guardrails, naming gaming modes, and choosing measures that survive Goodhart pressure. Trigger for "incentive plan", "agent bonus scheme", "SPIFF design", "pay for QA score", "what metric should we bonus", CSAT incentives, or reviewing whether a comp change is driving gaming.
1
recall
Computes the Recall metric using torchmetrics, including configuration for binary, multiclass, and multilabel tasks.
3
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2
c2c-eval
Benchmarks language model agents on the C2C multi-agent negotiation task, reporting win rate across starting positions.
3
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
evaluation
Build evaluation frameworks for agent systems with deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, and outcome measurement.
16.9k · bundle
eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
3