SkillMD SkillMD
Skills
All skills Official skills Leaderboard Saved
Categories
Coding & Dev Tools9777AI & ML5020DevOps & Infra2868Integrations & APIs2327Productivity2192Security1976 All categories →
Plugins Docs
menu-rounded
Skills Categories Plugins Docs My skills Saved
light-dark-mode
Profile My skills Saved Collections Edit profile Submit a skill
All Skills 25,836 ✦ Verified
Categories
AI & ML 5,020
Agent Building 585 Image & Video Generation 189 MCP Servers 94 Model Training & Fine-tuning 377 Prompt Engineering 130 RAG & Embeddings 120 Speech & Audio 107
Coding & Dev Tools 9,777 Data & Analytics 1,729 Design & Media 985 DevOps & Infra 2,868 Docs & Writing 1,311 Finance & Business 347 Integrations & APIs 2,327 Marketing & Growth 1,430 Product & Planning 1,037 Productivity 2,192 Research & Search 928 Security 1,976 Web & Frontend 1,641

Results for “spearman-rank”

6 skills
qhjqhj00
Posh
Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
3
More results
qhjqhj00
Ndcg 10
Evaluates how well internal model representations (hidden states) predict token-level information importance in summarization tasks, using NDCG@10 and Spearman's rank correlation.
3
qhjqhj00
Epsilon
Evaluates the correlation between a zero-cost NAS metric (epsilon) and actual training accuracy across different neural architecture search spaces, testing the metric's ability to rank architectures without training. It probes whether output dispersion from constant weight initializations can serve as a reliable.
3
qhjqhj00
Menli
Evaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions.
3
qhjqhj00
Spice
Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
3
qhjqhj00
Squad
Computes the SQuAD metric using torchmetrics, given predictions and ground truth. Use when evaluating question-answering outputs with exact match and F1 scores.
3
SKILLMD.com

The open registry of AI Agent Skills: safety-reviewed SKILL.md files for Claude, Cursor, Codex & 60+ agents.

$ npm i skillmds

Explore

All Skills Categories Agents Plugins New & Latest Leaderboard

Support

About Contact npm Terms Privacy

Learn

Docs Blog Stats FAQ Submit a Skill
© 2026 SkillMD.com Skills attributed to their authors under their original licenses.
SKILLMD