AI & ML
AI & ML agent skills cover the machine-learning workflow itself: writing and evaluating prompts, building RAG pipelines, running evals, and wiring up model APIs. Each one is a SKILL.md file your agent loads on demand, so the know-how travels across Claude Code, Cursor, and 60+ agents.
-
williamlimasilva Bundle Arize LinkGenerate deep links to the Arize UI. Use when the user wants a clickable URL to open a specific trace, span, session, dataset, labeling queue, evaluator, or annotation config.
-
williamlimasilva Bundle Arize TraceINVOKE THIS SKILL when downloading or exporting Arize traces and spans. Covers exporting traces by ID, sessions by ID, and debugging LLM application issues using the ax CLI.
-
zwright8 Bundle Coalition Rapid Language Model Red Teaming CellRed-team coalition language models used in operational workflows for misuse, bias, and manipulation risks. Use when validating AI safety, reliability, and mission suitability under coalition constraints.
-
williamlimasilva Bundle Phoenix CLIDebug LLM applications using the Phoenix CLI. Fetch traces, analyze errors, review experiments, inspect datasets, and query the GraphQL API. Use when debugging AI/LLM applications, analyzing trace data, working with Phoenix observability, or investigating LLM performance issues.
-
williamlimasilva Bundle Arize DatasetINVOKE THIS SKILL when creating, managing, or querying Arize datasets and examples. Covers dataset CRUD, appending examples, exporting data, and file-based dataset creation using the ax CLI.
-
williamlimasilva Bundle Arize EvaluatorINVOKE THIS SKILL for LLM-as-judge evaluation workflows on Arize: creating/updating evaluators, running evaluations on spans or experiments, tasks, trigger-run, column mapping, and continuous monitoring. Use when the user says: create an evaluator, LLM judge, hallucination/faithfulness/correctness/relevance, run eval, score my spans or experiment, ax tasks, trigger-run, trigger eval, column mapping, continuous monitoring, query filter for evals, evaluator version, or improve an evaluator prompt.
-
williamlimasilva Bundle Eval Driven DevInstrument Python LLM apps, build golden datasets, write eval-based tests, run them, and root-cause failures — covering the full eval-driven development cycle. Make sure to use this skill whenever a user is developing, testing, QA-ing, evaluating, or benchmarking a Python project that calls an LLM, even if they don't say "evals" explicitly. Use for making sure an AI app works correctly, catching regressions after prompt changes, debugging why an agent started behaving differently, or validating output quality before shipping.
-
williamlimasilva Bundle Quality PlaybookExplore any codebase from scratch and generate six quality artifacts: a quality constitution (QUALITY.md), spec-traced functional tests, a code review protocol, an integration testing protocol, a multi-model spec audit (Council of Three), and an AI bootstrap file (AGENTS.md). Works with any language (Python, Java, Scala, TypeScript, Go, Rust, etc.). Use this skill whenever the user asks to set up a quality playbook, generate functional tests from specifications, create a quality constitution, build testing protocols, audit code against specs, or establish a repeatable quality system for a project. Also trigger when the user mentions 'quality playbook', 'spec audit', 'Council of Three', 'fitness-to-purpose', 'coverage theater', or wants to go beyond basic test generation to build a full quality system grounded in their actual codebase.
-
zwright8 Bundle Tactical Edge AI Model Assurance CellSupport U.S. warfighter planning and decision support for Tactical Edge AI Model Assurance Cell. Use when missions require tactical edge ai model assurance cell planning, integrated options, and protocol-aware staff outputs.
-
williamlimasilva Bundle Arize InstrumentationINVOKE THIS SKILL when adding Arize AX tracing to an application. Follow the Agent-Assisted Tracing two-phase flow: analyze the codebase (read-only), then implement instrumentation after user confirmation. When the app uses LLM tool/function calling, add manual CHAIN + TOOL spans so traces show each tool's input and output. Leverages https://arize.com/docs/ax/alyx/tracing-assistant and https://arize.com/docs/PROMPT.md.
-
williamlimasilva Skill MCP Security BaselineReview MCP (Model Context Protocol) server and client source code against a security baseline — authentication, sessions, rate limiting, input-schema validation, official-SDK usage, RCE vectors, and the OWASP MCP Top 10 — producing a report with file/line evidence. Use this skill when: - Reviewing an MCP server implementation for security before release - Checking a server against the baseline controls (MCP-01 to MCP-05) and the OWASP MCP Top 10 - Auditing tools for RCE vectors (command/code injection, unsafe deserialization, path traversal, SSTI, dependency hijacking, SSRF) - Verifying auth, session, rate-limiting, and input-validation controls on a network-exposed server - Reviewing MCP client code that handles untrusted server responses and session IDs - Requests like "review this MCP server for security" or "is my MCP server implementation secure?"
-
williamlimasilva Skill Onboard Context MaticInteractive onboarding tour for the context-matic MCP server. Walks the user through what the server does, shows all available APIs, lets them pick one to explore, explains it in their project language, demonstrates model_search and endpoint_search live, and ends with a menu of things the user can ask the agent to do. USE FOR: first-time setup; "what can this MCP do?"; "show me the available APIs"; "onboard me"; "how do I use the context-matic server"; "give me a tour". DO NOT USE FOR: actually integrating an API end-to-end (use integrate-context-matic instead).
-
williamlimasilva Skill Integrate Context MaticDiscovers and integrates third-party APIs using the context-matic MCP server. Uses `fetch_api` to find available API SDKs, `ask` for integration guidance, `model_search` and `endpoint_search` for SDK details. Use when the user asks to integrate a third-party API, add an API client, implement features with an external API, or work with any third-party API or SDK.
-
williamlimasilva Bundle Arize Prompt OptimizationINVOKE THIS SKILL when optimizing, improving, or debugging LLM prompts using production trace data, evaluations, and annotations. Covers extracting prompts from spans, gathering performance signal, and running a data-driven optimization loop using the ax CLI.
-
zwright8 Bundle Strategic Deterrence Escalation Management CellModel and communicate strategic deterrence signaling and escalation-control options. Use when commanders need option framing, threshold mapping, and cross-domain escalation risk controls under uncertainty.
-
williamlimasilva Skill Flowstudio Power Automate GovernanceGovern Power Automate flows and Power Apps at scale using the FlowStudio MCP cached store. Classify flows by business impact, detect orphaned resources, audit connector usage, enforce compliance standards, manage notification rules, and compute governance scores — all without Dataverse or the CoE Starter Kit. Load this skill when asked to: tag or classify flows, set business impact, assign ownership, detect orphans, audit connectors, check compliance, compute archive scores, manage notification rules, run a governance review, generate a compliance report, offboard a maker, or any task that involves writing governance metadata to flows. Requires a FlowStudio for Teams or MCP Pro+ subscription — see https://mcp.flowstudio.app
-
williamlimasilva Skill Flowstudio Power Automate MonitoringMonitor Power Automate flow health, track failure rates, and inventory tenant assets using the FlowStudio MCP cached store. The live API only returns top-level run status. Store tools surface aggregated stats, per-run failure details with remediation hints, maker activity, and Power Apps inventory — all from a fast cache with no rate-limit pressure on the PA API. Load this skill when asked to: check flow health, find failing flows, get failure rates, review error trends, list all flows with monitoring enabled, check who built a flow, find inactive makers, inventory Power Apps, see environment or connection counts, get a flow summary, or any tenant-wide health overview. Requires a FlowStudio for Teams or MCP Pro+ subscription — see https://mcp.flowstudio.app
-
bhaumikgohel Bundle Running Smoke TestsRuns automated smoke tests after deployment to validate critical paths and UI readiness. Use after deployments, before releases, or when verifying environment health.
-
bhaumikgohel Bundle Validating UI ContentValidates UI text content for spelling, grammar, formatting, and consistency issues. Use when testing web pages, reviewing UI copy, checking for content errors, or validating labels and messages.
-
bhaumikgohel Bundle Detecting Duplicate BugsDetects duplicate JIRA bugs before creation by comparing new bug summaries against existing issues. Use when QA is logging bugs, mentions potential duplicates, or before creating new JIRA tickets.
-
zwright8 Bundle U0542 Engineering Multi Agent Negotiation MediatorBuild and operate the "Engineering Multi-Agent Negotiation Mediator" capability for Software Engineering Automation. Trigger when this exact capability is needed in mission execution.
-
javiertarazon Bundle MCP BuilderGuide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate exte...
-
javiertarazon Skill Agent OrchestrationMulti-agent orchestration patterns using Gem/RUG Orchestrator style: decompose goals, delegate to sub-agents, validate and repeat until complete. Use when coordinating multiple specialized agents for complex autonomous tasks.
-
javiertarazon Skill Agent Safety GovernanceSafety and governance framework for autonomous AI agents. Validates agent actions before execution, enforces permission boundaries, detects dangerous patterns (file deletion, env modification, network calls), and maintains audit trails. Use before deploying autonomous agents or when reviewing agent-generated code.
-
zwright8 Bundle U0102 Memory Multi Agent Negotiation MediatorBuild and operate the "Memory Multi-Agent Negotiation Mediator" capability for Memory and Knowledge Operations. Trigger when this exact capability is needed in mission execution.
-
zwright8 Bundle U0262 Collab Multi Agent Negotiation MediatorBuild and operate the "Collab Multi-Agent Negotiation Mediator" capability for Collaboration and Negotiation. Trigger when this exact capability is needed in mission execution.
-
zwright8 Bundle U0662 Crisis Multi Agent Negotiation MediatorBuild and operate the "Crisis Multi-Agent Negotiation Mediator" capability for Crisis and Incident Response. Trigger when this exact capability is needed in mission execution.
-
activer007 Bundle Reviewing ChangesAndroid-specific code review workflow additions for Bitwarden Android. Provides change type refinements, checklist loading, and reference material organization. Complements bitwarden-code-reviewer agent's base review standards.
-
zwright8 Bundle U0702 Impact Multi Agent Negotiation MediatorBuild and operate the "Impact Multi-Agent Negotiation Mediator" capability for Social Impact Measurement. Trigger when this exact capability is needed in mission execution.
-
zwright8 Bundle U0902 Rights Multi Agent Negotiation MediatorBuild and operate the "Rights Multi-Agent Negotiation Mediator" capability for Legal, Rights, and Compliance. Trigger when this exact capability is needed in mission execution.
-
activer007 Bundle Moon Dev Trading AgentsMaster Moon Dev's Ai Agents Github with 48+ specialized agents, multi-exchange support, LLM abstraction, and autonomous trading capabilities across crypto markets
-
microck Bundle Hermes Tweet XquikUse Hermes Tweet and Xquik for X/Twitter agent workflows. Plan social listening, account and follower analysis, post research, monitors, webhook alerts, REST API calls, MCP client setup, and safe tweet actions.
-
duclm1x1 Skill Cellcog1 on DeepResearch Bench (Feb 2026). Any-to-Any AI for agents. Combines deep reasoning with all modalities through sophisticated multi-agent orchestration. Research, videos, images, audio, dashboards, presentations, spreadsheets, and more.
-
duclm1x1 Skill Instagram TeneoThe agent gives you the ability to extract data from instagram through different commands.
-
duclm1x1 Skill Command CenterMission control dashboard for OpenClaw - real-time session monitoring, LLM usage tracking, cost intelligence, and system vitals. View all your AI agents in one place.
-
duclm1x1 Skill Tiktok TeneoThe agent gives you the ability to extract data from tiktok through different commands.
Frequently asked questions
What are AI & ML agent skills?
AI & ML agent skills cover the machine-learning workflow itself: writing and evaluating prompts, building RAG pipelines, running evals, and wiring up model APIs. Each one is a SKILL.md file your agent loads on demand, so the know-how travels across Claude Code, Cursor, and 60+ agents.
Which AI & ML skills are most installed?
Popular AI & ML skills on SkillMD right now include quality-playbook, tactical-edge-ai-model-assurance-cell, reviewing-changes. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do AI & ML skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.