Results for “distribution-score”
10 skillsMore results
Rulebase QA Coverage Audit
Use to audit QA coverage and scorecard health in a Rulebase workspace via the Rulebase MCP server. Trigger for "audit our QA coverage", "which agents or channels aren't being evaluated", "are our QA scores meaningful", "is our scorecard working", QA blind spots, score distribution or ceiling effects, and checking whether QA scores relate to SLA or complaint outcomes.
1 · bundle
Anderson
Computes the Anderson-Darling test statistic and p-value using scipy.stats.anderson for evaluating predictions against ground truth.
3
Trak Attributing Model Behavior At Scale Arxiv 2303 14186v2
TRAK: Attributing Model Behavior at Scale
6
Polygraph
Assigns behavioral trust grades (A–F) to MCP servers by running probes for prompt injection, permission overreach, data leaks, and adversarial-input handling, and publishes reproducible onchain attestations.
1.2k · bundle
Alphagbm Options Score
Score and rank options contracts for any ticker using a multi-factor model covering liquidity, IV attractiveness, Greeks balance, and risk/reward. Returns scored option chains with the best contracts highlighted.
1.2k
Bleurt
Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
3
Eval Judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
3
Econ Audit
Audit economic analysis outputs (fiscal briefings, macro briefings, market research, longlists, and other quantitative economic documents) against methodology standards, academic literature, and common errors. Runs structured checks across core categories including counterfactual, additionality, discounting, double counting, distributional analysis, Aqua Book RIGOUR, and Flyvbjerg-style strategic misrepresentation detection. Returns a RAG scorecard with issues ranked by severity.
1k · bundle