SkillMD SkillMD
Skills
All skills Official skills Leaderboard Saved
Categories
Coding & Dev Tools9777AI & ML5020DevOps & Infra2868Integrations & APIs2327Productivity2192Security1976 All categories →
Plugins Docs
menu-rounded
Skills Categories Plugins Docs My skills Saved
light-dark-mode
Profile My skills Saved Collections Edit profile Submit a skill
All Skills 25,836 ✦ Verified
Categories
AI & ML 5,020
Agent Building 585 Image & Video Generation 189 MCP Servers 94 Model Training & Fine-tuning 377 Prompt Engineering 130 RAG & Embeddings 120 Speech & Audio 107
Coding & Dev Tools 9,777 Data & Analytics 1,729 Design & Media 985 DevOps & Infra 2,868 Docs & Writing 1,311 Finance & Business 347 Integrations & APIs 2,327 Marketing & Growth 1,430 Product & Planning 1,037 Productivity 2,192 Research & Search 928 Security 1,976 Web & Frontend 1,641

Results for “credibility-scoring”

6 skills
qhjqhj00
Polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
qhjqhj00
Accuracy
Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
3
qhjqhj00
Spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
qhjqhj00
Dior
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
3
qhjqhj00
Eas
Validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity, using the AmbGIMT dataset and pairwise score comparisons.
3
qhjqhj00
Cider
Computes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
3
SKILLMD.com

The open registry of AI Agent Skills: safety-reviewed SKILL.md files for Claude, Cursor, Codex & 60+ agents.

$ npm i skillmds

Explore

All Skills Categories Agents Plugins New & Latest Leaderboard

Support

About Contact npm Terms Privacy

Learn

Docs Blog Stats FAQ Submit a Skill
© 2026 SkillMD.com Skills attributed to their authors under their original licenses.
SKILLMD