AI & ML
AI & ML agent skills cover the machine-learning workflow itself: writing and evaluating prompts, building RAG pipelines, running evals, and wiring up model APIs. Each one is a SKILL.md file your agent loads on demand, so the know-how travels across Claude Code, Cursor, and 60+ agents.
-
greyhaven-ai Skill Autocontext 2Iterative strategy generation and evaluation system. Use when the user wants to evaluate agent output quality, run improvement loops, queue tasks for background evaluation, check run status, inspect runtime artifacts and session branch lineage, or discover available scenarios. Provides LLM-based judging with rubric-driven scoring.
-
monte-carlo-data Bundle Monte Carlo Remediation 2Investigate and remediate data quality alerts using Monte Carlo MCP tools. Runs root cause analysis, assesses blast radius, discovers available tools (MCP/CLI/API), proposes and executes fixes, or escalates with full context when uncertain.
-
monte-carlo-data Bundle Automated Triage 2Triage Monte Carlo alerts interactively or build an automated workflow. Fetch, score, and troubleshoot alerts using MCP tools now, or design a reusable workflow that runs on a schedule.
-
helius-labs Bundle Dflow 2Build Solana trading applications combining DFlow trading APIs with Helius infrastructure. Covers spot swaps (imperative and declarative), prediction markets, real-time market streaming, Proof KYC, the DFlow Agent CLI for autonomous trading, transaction submission via Sender, fee optimization, shred-level streaming via LaserStream, and wallet intelligence.
-
data-goblin Bundle Pbir CLI 2This skill should be used whenever the user mentions "pbir", "pbir-cli", "Power BI reports", or "PBI reports", works with .pbir, .pbip, or .pbix files, or wants to refresh, screenshot, or visually verify a report that is open in Power BI Desktop. Covers creating, exploring, formatting, validating, and publishing Power BI reports through the pbir CLI and object model, plus driving Power BI Desktop (canvas reload, page screenshots) and querying connected or local semantic models.
-
scdenney Bundle Spawn 2Spawn full Claude Code peer sessions in their own terminal panes and git worktrees — real sessions, not subagents — each on a directed task with a contract brief, monitored and merged back by the spawning lead. Detects the environment and takes the strongest path — herdr first, then tmux, then a native claude background agent. Use when work must outlive or run beside the current session, needs its own worktree or its own permission settings, should stay steerable by the user in a visible pane, or when the user asks to spawn, hand off, or parallelize across full sessions. Also spawns Codex peers into the same panes. Not for bounded consults or work a subagent covers.
-
othmanadi Bundle Planning With Files 14Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
othmanadi Bundle Planning With Files 15Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
othmanadi Bundle Planning With Files 16Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
othmanadi Bundle Planning With Files 17Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
othmanadi Bundle Planning With Files 18Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
scdenney Bundle Advisor 2Consult Fable 5.1 as an independent second reviewer, always at max reasoning effort. The calling session is the main model — Opus 5 or Sonnet 5 — and Fable holds the advisor seat. Use before committing to an interpretation or a substantial piece of writing/analysis, when stuck (recurring errors, a non-converging approach, results that do not fit), when considering a change of approach, or when you believe a task is complete and want a check before finalizing. Fallback for when the native advisor tool is unavailable. Not a co-implementer — read-only advisory only, does not edit files.
-
scdenney Bundle Diverge 2Generate 3-5 conceptually distinct approaches labeled by creativity dimension (Novel, Surprising, Diverse, Conventional) and hold for selection instead of implementing the first idea; the --codex mode runs the same brainstorm on GPT-6 Astra via codex exec, then has Codex implement the selected approach. Use when a task has more than one non-obvious solution — creative, architectural, or analytical work — and before committing to an approach; use --codex when a second model family should widen the range of approaches.
-
scdenney Bundle Orchestrate 2Run a multi-model orchestration workflow led by the session's own strongest-available model — Fable 5.1 or Claude Opus 5 — delegating mechanical work (boilerplate, tests, formatting, bulk edits) to a fast-worker subagent (Sonnet), wide or parallel reasoning to deep-reasoner subagents (Opus, pinned high), and high-stakes or fresh-perspective calls to Codex, a different-vendor GPT-6 Astra peer (`gpt-6-astra` by default). The lead is detected from the model line in the session's own context, and `--lead fable` or `--lead opus` overrides it. Under a Fable lead the hard reasoning and the judgment calls stay in the lead and only mechanical or genuinely wide work goes out; under an Opus lead the lead is itself the deep reasoner, delegating to fan out or stay context-lean, running multi-agent phases as parallel Agent fan-outs and upgrading to a dynamic Workflow where the session actually has that tool. Use to orchestrate, delegate, fan out, get a decorrelated second opinion from Codex, run a blind Opus and Codex cros
-
othmanadi Bundle Planning With Files 19Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
othmanadi Bundle Planning With Files 20Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.
-
starchild-ai-agent Bundle Wallet 2Send transactions on Monad testnet or mainnet using Agent wallet, the transactions could be for smart contract deployment, onchain actions like swapping or smart contract calls and signing messages
-
scdenney Bundle Model Committee 2Runs a deliberative two-model committee — GPT-6 Astra and Claude Opus 5 as members, under a selectable chair. Chair defaults to Fable 5.1; `/model-committee-astra` selects Astra with a Sol member, `/model-committee-opus` selects Opus, and `/model-committee-sol` retains the Sol chair with a Terra member. Use when one consequential decision must come out of several defensible options and the two model families should propose independently, critique each other, revise, and cross-rank under a predeclared rubric before converging. Fits architecture, research design and interpretation, manuscript strategy, ambiguous diagnosis, evaluation design, plan reconciliation, and policy or standards tradeoffs. Not for factual lookups, independent-coder reliability (use model-council-voting), open-ended brainstorming, routine implementation, or final high-stakes professional judgment.
-
scdenney Bundle Conjoint Diagnostics 2Reviews an existing conjoint study for threats to inference and returns prioritized findings across five areas — design integrity (attributes, profile restrictions, task count and satisficing, randomization, power), estimation (estimand clarity, reference levels, subgroups, clustered standard errors, multiple testing), measurement error, external validity and behavioral benchmarking, and interpretation, including the guard-rail against reading an AMCE as a majority preference. Use when the user asks whether a conjoint design or analysis holds up, has referee comments on a conjoint, or wants a second opinion on estimation and interpretation choices. Building a design from scratch goes to conjoint-design, reshaping the data to conjoint-cleaning.
-
scdenney Skill LLM Calibration Logprobs 2Reads a model's own uncertainty off its token log-probabilities — collecting logprobs and aggregating multi-token labels, confidence tiers and margins for triage, calibration assessment with ECE, Brier scores, and reliability diagrams, using confidence downstream without laundering it into evidence, and what to archive for reproducibility. Use when the user asks how confident a classifier was on each item, asks about logprobs, top-k tokens, calibration, ECE, Brier, or reliability diagrams, or wants low-confidence cases routed to human review. This is within-model confidence — agreement across several independent model coders goes to model-council-voting, and codebook and validation design to text-classification.
-
scdenney Skill Pre Registration Writing 2Writes a pre-analysis plan before data collection — registry selection (OSF, AsPredicted, AEA, EGAP), PAP document structure, an analytical strategy specified down to the model and the decision rule, analysis code pre-registered against simulated data, contingency planning for attrition, failed manipulations, and exclusions, deviation documentation, and timeline. Operationalizes the pre-data-collection side of DA-RT. Use when the user asks to write or review a pre-registration or PAP, asks which registry to use, asks what to lock down versus leave exploratory, or asks how to handle a deviation later. Hypotheses and estimands come from hypothesis-building, post-hoc reporting from methods-reporting.
-
microsoft Bundle Apm Spec Guardian 2Use this skill to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-*.md), its inline / sidecar JSON Schemas (docs/src/content/docs/specs/schemas/*.schema.json), or the conformance fixture seed (tests/fixtures/spec-conformance/**). The panel fans out to four spec-ecosystem reviewers (swagger-openapi-editor, oci-distribution-editor, pkgmgr-registry-contract-editor, w3c-tag-architect), each running in its own agent thread, and a spec-editor synthesizer that produces a fold-now / defer-v0.1.1 / defer-v0.2 / reject list plus a ship decision keyed off a 1..10 shocked_meter scale. The orchestrator is the sole writer to the PR: ONE consolidated comment, no verdict labels, no merge gating. The panel is advisory -- it surfaces findings, prioritizes folds, and renders a ship recommendation that the maintainer weighs.
2.7k -
microsoft Bundle Docs Corpus Audit 2Use this skill to run a holistic regrounding pass on the entire microsoft/apm documentation corpus against current source code, page-by-page, and emit surgical fixes for stale claims. Activate when the maintainer wants a WHOLE-CORPUS audit (not per-PR review) -- typical triggers include "audit the docs", "reground the corpus", "check every page against code", "pre-release docs sweep", "the docs have drifted everywhere", or "we just reshaped the TOC, find dead links". Wave-batched and S7-verified; scales to the full ~112-page corpus in ~10 minutes wall-time. This is a SIBLING to docs-sync, not a replacement: docs-sync is per-PR (triggered by a diff); this skill is per-corpus (triggered by a maintainer ask). They share agent personas, schemas, and the docs index, but their triggers MUST NOT collide. Does NOT auto-merge, does NOT push without maintainer review, and does NOT replace per-PR drift detection.
2.7k -
athola Skill Ideate 2Generate diverse solution candidates with category-spanning ideation methods and rotation. Use when stuck on a design or fighting repetitive LLM output.
-
athola Bundle War Room 2Convenes a multi-LLM expert panel to pressure-test hard-to-reverse decisions. Use when reversibility score is low and adversarial review is warranted.
-
athola Bundle Utility 2Scores agent actions by expected gain, cost, uncertainty, and redundancy. Use when deciding whether to dispatch an agent or invoke a tool.
-
athola Bundle Latent Space Engineering 2Shapes agent behavior via instruction framing and style transfer. Use when composing dispatch prompts or writing skill instructions for parallel review agents.
-
athola Bundle Assisted Mastery 2Makes agent reasoning visible, surfaces tradeoffs, and fades help so humans build judgment. Use when reviewing or learning from agent-written code.
-
athola Bundle Damage Control 2Recovers broken agent state via crash recovery, context overflow, and merge conflict protocols. Use when an agent session fails or a worktree is corrupted.
-
athola Bundle Agent Teams 2Coordinates Claude agent teams via filesystem protocol. Use when orchestrating parallel agents with task dependencies. Do not use for single-agent tasks.
-
evrenverse Bundle PDF ToolInspect, search, read, render, create, fill, overlay, merge, split, batch-edit, and sign PDF files with the pdf-tool CLI. Use when an agent must work with PDF text, pages, tables, AcroForm fields, coordinates, or signatures while keeping reads bounded, emitting JSON, validating operations, and verifying the result.
-
athola Bundle Doc Consolidation 2Merges ephemeral report and analysis artifacts into permanent documentation. Use when LLM-generated markdown files have accumulated and need consolidation.
-
chinthakat Skill Agent InstructionsWrite and maintain the repository instruction file an AI coding agent loads every session (CLAUDE.md, AGENTS.md or equivalent), and decide what belongs there versus in an on-demand skill. Use when setting one up, when an agent keeps repeating a mistake the file was supposed to prevent, or when the file has grown past the point of being read.
-
chinthakat Skill Parallel WorktreesWork safely when several people or agent sessions share one repository through git worktrees or parallel checkouts. Use when a repo has sibling worktrees, when changes appear that you did not make, when a stash or branch goes missing, or when setting up a repo for concurrent agent sessions.
-
athola Bundle Agent Expenditure 2Tracks per-agent token usage and flags waste in parallel dispatch. Use when evaluating parallel agent efficiency or after a multi-agent run.
-
muhtasimm Bundle Multi Code ReviewReview a code change before merge with five independent specialist agents covering correctness, readability and simplicity, architecture, security, and performance. Use for pull requests, completed implementations, refactors, bug fixes, dependency changes, or human-authored and agent-authored diffs. Do not trigger for casual code explanation or when the user only asks to implement a change.
Frequently asked questions
What are AI & ML agent skills?
AI & ML agent skills cover the machine-learning workflow itself: writing and evaluating prompts, building RAG pipelines, running evals, and wiring up model APIs. Each one is a SKILL.md file your agent loads on demand, so the know-how travels across Claude Code, Cursor, and 60+ agents.
Which AI & ML skills are most installed?
Popular AI & ML skills on SkillMD right now include autocontext, monte-carlo-remediation, automated-triage. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do AI & ML skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.