Packs

4 packs

Results for “evaluation”

396 skills
anantha-236
skill-stocktake
Use when auditing Claude skills and commands for quality. Supports Quick Scan (changed skills only) and Full Stocktake modes with sequential subagent batch evaluation.
1 · bundle
ecnu-icalk
skill-creator
Guides users through creating, refining, and evaluating agent skills, including drafting, testing, and optimizing descriptions for better triggering.
559 · bundle
auto-skiller
receiving-code-review
Guides technical evaluation of code review feedback, emphasizing verification before implementation and reasoned pushback over performative agreement.
1 · bundle
tools-only
085-aeon-556c1766
Provides guidance on using the Aeon library for time series forecasting, covering model selection, implementation, and evaluation.
7 · bundle
manu14357
mcp-builder
Build high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services. Covers design principles, architecture, implementation patterns, and evaluation-driven development.
16 · bundle
seaworld008
oracle
Designing and evaluating AI/ML systems: prompt engineering, RAG design, LLM application patterns, AI safety, evaluation frameworks, MLOps, cost optimization. Use for AI pipelines or eval harnesses.
65 · bundle
alirezarezvani
run
Execute the full AgentHub competition lifecycle in a single command: initialize, capture baseline, spawn agents, evaluate results, and merge the winner.
20.4k
nvidia
tao-run-on-brev
Manage NVIDIA Brev GPU instances for TAO training, evaluation, and inference using the Brev CLI and Docker.
2.2k · bundle
affaan-m
continuous-learning
Automatically evaluates Claude Code sessions to extract reusable patterns and save them as learned skills.
226k · bundle
tinh2
ux
Audits UX quality of a codebase or validates implementation against design mockups, then fixes issues and commits changes.
13
jorcan
agents
Evaluates execution transcripts and output files against a list of expectations, assigning pass/fail verdicts with cited evidence and critiquing the assertions themselves.
0 · bundle
tools-only
065-data-61a12d5f
Guides data protection impact assessments under GDPR Article 35, covering mandatory triggers, risk evaluation, and mitigation steps.
7 · bundle
kairyou
at-self-eval
Summarize a contributor's Git history, a provided work log, or both into a concise, review-friendly self-evaluation for quarterly, semi-annual, or promotion cycles.
167
vvieira010-pixel
metacognitive-prompt-library
Build a library of metacognitive prompts targeting planning, monitoring, or evaluation for a specific task. Use when developing students' thinking-about-thinking during independent work.
0
x3allamerican
train-eldt-btw-vendor-vetting
Use this skill when selecting a TPR-registered Training Provider for behind-the-wheel ELDT training. Covers the registry, evaluation criteria, and common pitfalls.
1
arustydev
openfeature-eng
Implement OpenFeature feature flags in software projects. Use when adding feature flags with OpenFeature SDKs, configuring providers, setting up evaluation context, or integrating the OpenFeature MCP Server.
8
alirezarezvani
cto-advisor
Provides technical leadership frameworks for architecture decisions, engineering team scaling, technology strategy, and technical debt assessment.
20.4k · bundle
nvidia
nemotron-retrieval-recipes
Plan, debug, tune, evaluate, export, or deploy public Nemotron embedding and reranking retrieval recipes using the current checkout.
2.2k · bundle
affaan-m
agent-eval
Compare coding agents head-to-head on reproducible tasks with pass rate, cost, time, and consistency metrics.
226k
majiayu000
rag
Builds Retrieval-Augmented Generation systems with document chunking, embedding generation, vector storage, and retrieval pipelines, including evaluation and optimization.
567 · bundle
nimoqup046-collab
langfuse
Instrument LLM applications with Langfuse to trace, score, and monitor cost, quality, and latency across OpenAI and LangChain integrations.
2
mhassan0000
mcp-builder
Guides the creation of high-quality MCP servers, covering design, implementation, testing, and evaluation for Python and TypeScript.
1 · bundle
mhassan0000
skill-creator
Guides users through creating, editing, and optimizing agent skills, including drafting, testing, evaluating, and improving skill descriptions for better triggering.
1 · bundle
qhjqhj00
fid
Measures distributional similarity between original GAN-generated images and their semantically manipulated counterparts using the Fréchet Inception Distance (FID) metric.
3
qhjqhj00
infolm
Computes the InfoLM metric from torchmetrics for evaluating text generation against ground truth, with configurable information measures and sentence-level scoring.
3
qhjqhj00
lambre
Scores generated text for morphosyntactic well-formedness by measuring how closely it adheres to language-specific dependency rules extracted from treebanks.
3
qhjqhj00
reflex
Evaluates machine-generated log summaries without human-written references, using LLM judgment and dense embeddings to score relevance, informativeness, and coherence.
3
livelybug
mle-workflow
Production machine-learning engineering workflow for data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback. Use when building, reviewing, or hardening ML systems beyond one-off notebooks.
0
rajanthar
mle-workflow
Production machine-learning engineering workflow for data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback. Use when building, reviewing, or hardening ML systems beyond one-off notebooks.
0
nvidia
tao-train-rtdetr
Train, evaluate, distill, quantize, export, and run inference for RT-DETR object detection models using NVIDIA TAO.
2.2k · bundle
jeffallan
architecture-designer
Design high-level system architecture, create Architecture Decision Records (ADRs), evaluate technology trade-offs, and plan for scalability.
10.4k · bundle
sirnosh
bmad-ml-cypher
Dataset analysis and data quality specialist. Use when the user asks to talk to Cypher, requests the data detective, or needs dataset assessment, bias analysis, and benchmark evaluation.
0 · bundle
tianhao909
langsmith-observability
LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.
1 · bundle
qcmuu
langsmith-observability
LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.
0 · bundle
alirezarezvani
agenthub
Spawns multiple parallel AI agents that compete on the same task using isolated git worktrees, evaluates results, and merges the best solution.
20.4k · bundle
microsoft
azure-ai-projects-ts
Build AI applications using the Azure AI Projects SDK for TypeScript, managing agents, connections, deployments, datasets, indexes, and evaluations.
2.7k · bundle