Results for “credibility-scoring”
11 skillsMore results
Polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
Confidence Calibration Check
Capture confidence ratings before and after a learning attempt to identify overconfidence and underconfidence patterns. Use when a student wants to understand how well they actually know something versus how well they think they know it.
0
Aice
Tracks bidirectional confidence scores across five domains (TECH, OPS, JUDGMENT, COMMS, ORCH) for agents and users, with triggers, anti-patterns, and pool scoring per runtime.
32 · bundle
Accuracy
Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
3
Spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
Dior
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
3
8k4
Checks on-chain agent trustworthiness, discovers agents for tasks, profiles agents, looks up wallet/identity records, contacts or dispatches agents, and reads or writes hosted metadata via the 8K4 Protocol (ERC-8004).
2
Eval Judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
Eas
Validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity, using the AmbGIMT dataset and pairwise score comparisons.
3
Cider
Computes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
3