Results for “model-ranking”

35 skills
More results
qhjqhj00
Dior
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
3
huggingface
Huggingface Best
Queries Hugging Face benchmark leaderboards to find the best AI models for a task, filters by device constraints, and returns a ranked comparison table with scores.
10.8k
google
Agent Platform Model Registry
Manage machine learning models in the Agent Platform Model Registry: list, describe, upload, update, and delete models and their versions.
14.4k
projectious-work
Pk Model Refresh
Use the model-recommender skill, Workflow C (Roster Refresh), to research and update the model roster from live benchmarks.
0
snoodleboot-io
Model Evaluation
Every metric encodes an opinion about which mistake hurts.
2
projectious-work
Model Recommender
Recommend the right AI model for a task by scoring candidates across six dimensions (Reasoning, Engineering, Speed, Breadth, Reliability, Governance) and displaying a spider-chart profile.
0 · bundle
livelybug
Benchmark Models
Cross-model benchmark for gstack skills. (gstack)
0
jiachen-t-wang
Trak Attributing Model Behavior At Scale Arxiv 2303 14186v2
TRAK: Attributing Model Behavior at Scale
6
google
Agent Platform Tuning
Fine-tune open models or Gemini models using Agent Platform infrastructure, from environment setup through data preparation, job configuration, monitoring, and deployment.
14.4k · bundle
bankrbot
Aeon Huggingface Trending
Filters and ranks trending Hugging Face models, datasets, and spaces by novelty and significance, providing a 'why notable' explanation for each pick.
1.2k · bundle
jarbitechture
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
0
snoodleboot-io
Model Monitoring
The layers trade timeliness against definitiveness.
2
shenxingy
Model Research
Research latest Claude models and update selection guide — run when new models drop or periodically to stay current
8 · bundle
dvy1987
Model Selection
Plan which model tier handles which work BEFORE execution begins — a high-cognition model deeply understands the problem, lays the foundations, then emits a modular plan assigning each module the cheapest tier that can safely execute it, with escalation tripwires and one-way-door protection. Advisory only: it announces "next module → tier X / model Y" at each boundary and the HUMAN switches models — harnesses like Cursor cannot switch mid-run. Load when the user asks which model to use, wants a model plan, model tiers, model-tier routing, assign models to tasks or modules, says "cheap model got stuck", "which model for this task", "cost-efficient model choice", or when implementation-plan / problem-to-plan need a model: tier column. NOT dynamic-routing (plan-path selection after failure) — this skill assigns cognition tiers to work.
3 · bundle
alirezarezvani
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
michaelschecht
Model Selection
Recommend model families and validation strategy based on data, constraints, and objective. Use when: (1) choosing algorithms, (2) balancing bias/variance, (3) planning benchmark baselines. NOT for: final legal/compliance sign-off.
0
georgeqle
Key Moments
Rank a topic's user-flow branches by proof priority (value × risk × frequency) right after user-flow-map, ordering the branches, gating variation breadth, and promoting or pruning flows so state-model and ux-variations grow the tree in proof order — writes only existing flow-tree ordering fields, no schema change.
1 · bundle
snoodleboot-io
Feature Engineering
Cardinality and model family jointly determine the encoding.
2
machenjie
Profiling
`task-agent`/`review-agent`: use when CPU, memory, I/O, database, network, rendering, or cost needs measured bottleneck evidence; skip without a profiling need.
4 · bundle
claude-dev-suite
Reranking
Reranking retrieved documents with cross-encoders and LLM rerankers. Cohere Rerank v3, Voyage rerank-2, BGE reranker, ColBERT late interaction, Jina reranker. Cost and latency tradeoffs, top-K in / top-N out strategy. USE WHEN: user mentions "rerank", "reranker", "cross-encoder", "Cohere Rerank", "Voyage rerank", "BGE reranker", "ColBERT", "Jina reranker", "bi-encoder" DO NOT USE FOR: initial retrieval - use `advanced-retrieval` or `hybrid-search`; query rewriting - use `query-transformations`; agent decisions - use `agentic-rag`
28
affaan-m
Mle Workflow
Turn model work into a production ML system with data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
226k
alphagbm
Alphagbm Options Score
Score and rank options contracts for any ticker using a multi-factor model covering liquidity, IV attractiveness, Greeks balance, and risk/reward. Returns scored option chains with the best contracts highlighted.
1.2k
google-gemma
Gemma Dev
Selects the right Gemma model for a task, recommends deployment tooling (Gradio, Transformers.js, Vertex AI, MLX), and applies optimizations like MTP and QAT.
· bundle
sakamoto-family-smile
Mle Workflow
Turn model work into a production ML system with data contracts, reproducible training, quality gates, deployable artifacts, and monitoring.
0
lambenthan
Review
通用跨模型审查:Review LLM 对任意研究制品进行独立评审,输出结构化评分、wiki 实体映射与改进建议
77
michaelschecht
Model Evaluation
Evaluate model quality with task-appropriate metrics and systematic error analysis. Use when: (1) comparing models, (2) analyzing failures, (3) setting go/no-go thresholds. NOT for: production monitoring implementation.
0
aniruddhaadak80
Model Benchmark
Benchmark LLM performance across tasks — latency, quality, cost comparison.
0
lucassantana-dev
Eval
Evaluate LLM outputs systematically — benchmarks, automated metrics, human preference, and regression tracking
1 · bundle
samyakjhaveri
Model Route
Recommends the optimal Claude model tier (Opus, Sonnet, or Haiku) for a given task by analyzing reasoning depth, blast radius, domain expertise, output length, and correctness cost, and suggests parallelization opportunities.
0
machenjie
Threat Modeling
`analysis-agent`/`task-agent`/`review-agent`: use for changed assets, trust boundaries, reachable abuse paths, impact, or control placement; skip without a security delta.
4 · bundle
mattpocock
Domain Modeling
Build and sharpen a project's domain model. Use when discussing codebase terminology, writing or editing a CONTEXT.md, or recording or editing an ADR.
236k · bundle
bliss-fox
Project Review
针对 Modular RAG MCP Server 项目的老师式复习 Agent。按章节带领用户系统复习项目知识点,每道题互动问答、给出参考答案,复习结束后记录掌握进度,每次开始时回顾上次进度并建议继续或复习。Use when user says '复习项目', '帮我复习', '带我复习', '开始复习', '项目复习', 'review project', 'study review', '学习复习', '复盘', or wants to systematically review and study the project.
1 · bundle
dylanckawalec
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
3
qhjqhj00
Posh
Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
3