Global Master Standard – Claude Skills Specification
Document ID: 077-SPEC-MASTER-claude-skills-standard.md Version: 2.3.0 Status: AUTHORITATIVE - Single Source of Truth Created: 2025-12-06 Updated: 2025-12-08 Audited Against:
- Anthropic Engineering Blog (ENGINEERING SOURCE - oldest, deepest technical insights)
- Official Anthropic Skills Blog (PRIMARY SOURCE - newest product guidance)
- Lee Han Chung Deep Dive (implementation details)
Sources:
- Official Anthropic Skills Blog Post ⭐ PRIMARY SOURCE
- Anthropic Engineering Blog - Skills Deep Dive ⭐ ENGINEERING SOURCE
- Official Anthropic Agent Skills Overview
- Official Anthropic Best Practices
- Claude Code Skills Documentation
- Lee Han Chung Deep Dive
Executive Summary
What Is a Claude Skill?
A Claude Skill is a filesystem-based capability package containing instructions, executable code, and resources that Claude can discover and use automatically. Skills are prompt-based context modifiers—NOT executable plugins or slash commands.
Official Definition (Anthropic): "Specialized capability packages that extend Claude's functionality for specific tasks. Claude will only access a skill when it's relevant to the task at hand."
Mental Model: "Building a skill for an agent is like putting together an onboarding guide for a new hire."
Four Core Design Principles (Anthropic Official)
Skills are architecturally designed with four foundational attributes:
Composable
- Multiple skills work together seamlessly
- Claude automatically identifies and coordinates which skills are needed
- No manual orchestration required
- Nixtla Impact:
nixtla-schema-mapper+nixtla-experiment-architectcan chain automatically
Portable
- Same skill format works across ALL platforms:
- Claude apps (claude.ai web/mobile)
- Claude Code (CLI tool)
- Claude API (Messages API)
- Write once, deploy everywhere
- Nixtla Impact: Internal skills work for web users, API customers, and CLI developers
- Same skill format works across ALL platforms:
Efficient
- Loads only minimal necessary information when needed
- Progressive disclosure pattern (description → full instructions → resources)
- Avoids context window bloat
- Nixtla Impact: 40+ skill portfolio won't overwhelm context if designed properly
Powerful
- Can include executable code for reliable task performance
- Pre-written scripts eliminate non-determinism
- Combines prompts + execution for complete workflows
- Nixtla Impact: TimeGPT API calls, pandas transformations, validation scripts all executable
Code Execution Performance Economics (Anthropic Engineering):
"Sorting a list via token generation is far more expensive than simply running a sorting algorithm."
Deterministic code execution provides orders of magnitude cost/performance advantages:
| Operation | Token Generation | Code Execution | Advantage |
|---|---|---|---|
| Sort 1000 items | <10 tokens for script call (~$0.00003) | 200x cheaper | |
| Pandas transform | ~5,000 tokens loading data + code | Script runs without context load | No context cost |
| TimeGPT API call | Non-deterministic, requires retry logic | Deterministic, runs once | Reliable + cheap |
Nixtla Production Impact:
- TimeGPT forecasting loops: Execute via script, not token generation
- Schema validation: Run validation.py, don't describe validation in tokens
- Data transformations: pandas scripts consume ZERO context tokens
- Experiment harnesses: Pre-written Python orchestration
Production Implication for Nixtla: These four principles guide ALL architectural decisions. Violating any principle (non-composable, platform-locked, context-heavy, prompt-only) indicates poor skill design.
Why Use Skills Instead of Ad-Hoc Prompts?
| Aspect | Ad-Hoc Prompts | Skills |
|---|---|---|
| Reusability | One conversation | Persistent across all conversations |
| Discovery | Manual context provision | Automatic activation based on intent |
| Organization | Scattered knowledge | Structured packages |
| Context Management | Full context loaded | Progressive disclosure (on-demand) |
| Code Integration | Generated each time | Pre-written, deterministic scripts |
Where Skills Live
| Location | Scope | Priority |
|---|---|---|
~/.claude/skills/ |
Personal (all projects) | 1 (lowest) |
.claude/skills/ |
Project-specific | 2 |
Plugin skills/ directory |
Plugin-bundled | 3 |
| Built-in skills | Platform-provided | 4 (highest) |
Later sources override earlier ones when names conflict.
1. Core Concepts
Skill = What + When + How + Allowed Tools + Optional Model Override
Every skill answers:
- What: What capability does this provide?
- When: When should Claude activate it?
- How: Step-by-step instructions for Claude
- Allowed Tools: Which tools are pre-approved during execution?
- Model Override: Should a different model handle this? (optional)
The Skill Tool Architecture
Critical insight: Skills are NOT in the system prompt.
Skills live in a meta-tool called Skill within the tools array:
tools: [
{ name: "Read", ... },
{ name: "Write", ... },
{
name: "Skill", // Meta-tool (capital S)
inputSchema: { command: string },
description: "<available_skills>..." // Dynamic list of all skill descriptions
}
]
How Skills Are Discovered and Invoked
Model-Invoked (Automatic):
- At startup, Claude's system prompt includes metadata (name + description) for all skills
- Claude reads user request and matches intent to skill descriptions
- Claude invokes
Skilltool with matchingcommandparameter - No algorithmic routing, embeddings, or keyword matching—pure LLM reasoning
User-Invoked (Manual):
- Type
/skill-nameto explicitly invoke a skill - Required when
disable-model-invocation: true
Message Injection Architecture
CRITICAL FOR NIXTLA INTERNAL TEAMS: Understanding how skills inject into conversations ensures predictable behavior in production workflows.
When a skill is invoked, it injects two user messages into the conversation:
Message 1: Metadata Message (Visible to User)
{
role: "user",
isMeta: false, // Visible in UI
content: [
{
type: "text",
text: "<command-message>skill-name is loading...</command-message>"
}
]
}
Purpose: Provides transparent UI feedback that a skill is executing.
Message 2: Skill Prompt Message (Hidden from User)
{
role: "user",
isMeta: true, // Hidden from UI, sent to API only
content: [
{
type: "text",
text: "[Full SKILL.md content with {baseDir} substitutions]"
}
]
}
Purpose: Injects skill instructions directly into Claude's context for reasoning.
XML Tag Structure
Skills use specific XML tags for message formatting:
<command-message>- Wraps the visible status indicator<command-name>- Contains the skill name for tracking<available_skills>- Lists all discoverable skills in Skill tool description
Production Impact for Nixtla:
- Skills don't pollute conversation history visible to users
- Skill instructions consume context budget (tracked via
isMeta: truemessages) - Multiple skill invocations in one session stack context linearly
- Context budget management is critical (see Section 4 for limits)
Chain-of-Thought Visibility (UX Insight)
OFFICIAL ANTHROPIC: "Users can view skills in Claude's chain-of-thought output while it works."
When Claude invokes skills, users see them in the reasoning trace/thinking blocks. This provides:
- Transparency: Users know which skills are active
- Debuggability: Developers can trace skill activation patterns
- Trust: Users understand Claude's decision-making process
Nixtla Customer Impact:
- TimeGPT customers will SEE when
nixtla-experiment-architectactivates - Internal teams can debug skill selection by reviewing chain-of-thought
- Enterprise admins can audit which skills employees trigger
2. Folder & Discovery Layout
Standard Directory Structure
skill-name/
├── SKILL.md # REQUIRED - Instructions + YAML frontmatter
├── scripts/ # OPTIONAL - Executable Python/Bash scripts
│ ├── analyze.py
│ └── validate.py
├── references/ # OPTIONAL - Docs loaded into context
│ ├── API_REFERENCE.md
│ └── EXAMPLES.md
├── assets/ # OPTIONAL - Templates referenced by path
│ └── report_template.md
└── LICENSE.txt # OPTIONAL - License terms
Naming Conventions
Folder names must match the name field exactly.
Recommended: Use gerund form (verb + -ing) for clarity:
processing-pdfsanalyzing-spreadsheetsgenerating-commit-messages
Acceptable alternatives:
- Noun phrases:
pdf-processing,data-analysis - Action-oriented:
process-pdfs,analyze-data
Avoid:
- Vague names:
helper,utils,tools - Generic names:
documents,data,files - Reserved words:
anthropic-*,claude-*
Directory Purposes
| Directory | Purpose | Loaded Into Context? | Token Cost |
|---|---|---|---|
004-scripts/ |
Executable code (deterministic operations) | No (executed via Bash) | None |
references/ |
Documentation (API docs, examples) | Yes (via Read tool) | High |
assets/ |
Templates, configs, static files | No (path reference only) | None |
Key Insight: Scripts execute without loading code into context. Only script OUTPUT consumes tokens.
Platform-Specific Implementation & Distribution
Claude Apps (claude.ai, mobile)
Availability: Pro, Max, Team, and Enterprise users
Installation Methods:
Official Skill-Creator Tool (Anthropic Recommended)
- Use the built-in
skill-creatorskill for interactive guidance - No manual file editing required
- Provides step-by-step skill creation workflow
- Available via slash command:
/skill-creator
- Use the built-in
Manual Installation:
- Create
.claude/skills/skill-name/SKILL.mdstructure - Upload to project or personal workspace
- Create
Admin Controls (Team/Enterprise):
- Admins must enable skills organization-wide via Settings
- Organization-level toggle controls all member access
- NIXTLA ENTERPRISE NOTE: Customers need admin approval before deploying skills
Official Pre-built Skills (Anthropic-provided):
- Excel processing
- PowerPoint manipulation
- Word document editing
- Fillable PDF handling
Nixtla Strategy: Leverage pre-built skills where possible; avoid reinventing document processing.
Claude API (Messages API)
⚠️ CRITICAL FOR NIXTLA API INTEGRATIONS:
New /v1/skills Endpoint:
- Programmatic skill control, versioning, and management
- Separate from Messages API calls
- Enables automated skill deployment and updates
Code Execution Tool Requirement:
- Skills with executable scripts require Code Execution Tool beta
- Must be enabled in API configuration
- Provides secure sandbox for script execution
- BLOCKER: Nixtla's TimeGPT/pandas/validation scripts won't work without this beta feature
Implementation Pattern (Nixtla API Integration):
# 1. Enable Code Execution Tool beta in API settings
# 2. Deploy skills via /v1/skills endpoint
# 3. Reference skills in Messages API calls
import anthropic
client = anthropic.Anthropic(api_key="...")
# Deploy skill (one-time)
skill_response = client.skills.create(
skill_content=skill_md_content,
version="1.0.0"
)
# Use skill in conversation
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=4096,
messages=[{"role": "user", "content": "Map my data to Nixtla schema"}],
# Skills automatically discovered and invoked
)
Production Requirements for Nixtla:
- Request Code Execution Tool beta access from Anthropic
- Implement
/v1/skillsdeployment automation - Version all skills semantically (1.0.0, 1.1.0, etc.)
- Test scripts in sandbox environment before production
Claude Code (CLI)
Installation Methods:
Marketplace Plugins (Anthropic Official):
- Install via
anthropics/skillsmarketplace - Managed updates and versioning
- NIXTLA DISTRIBUTION: Consider publishing nixtla-skills as marketplace plugin
- Install via
Manual Installation:
- Place skills in
~/.claude/skills/(personal, all projects) - Place skills in
.claude/skills/(project-specific) - Version control friendly: Commit
.claude/skills/to git repos
- Place skills in
Priority Hierarchy:
~/.claude/skills/ Priority 1 (lowest)
.claude/skills/ Priority 2
Plugin skills/ Priority 3
Built-in skills Priority 4 (highest)
Later sources override earlier ones when names conflict.
Nixtla Distribution Strategy:
- Primary: Project-specific (
.claude/skills/nixtla-*/) - Fallback: Personal installation (
~/.claude/skills/nixtla-*/) - Future: Marketplace plugin for external distribution
3. SKILL.md Specification
Complete Structure
---
name: skill-name
description: What this skill does. Use when [conditions]. Trigger with "[phrases]".
---
# Skill Name
Brief purpose statement (1-2 sentences).
## Overview
What this skill does, when to use it, key capabilities.
## Prerequisites
Required tools, APIs, environment variables, packages.
## Instructions
### Step 1: [Action Verb]
[Imperative instructions]
### Step 2: [Action Verb]
[More instructions]
## Output
What artifacts this skill produces.
## Error Handling
Common failures and solutions.
## Examples
Concrete usage examples with input/output.
## Resources
Links to bundled files using {baseDir} variable.
4. YAML Frontmatter Fields
Required Fields
name
Type: string Required: YES Max Length: 64 characters Constraints:
- Lowercase letters, numbers, and hyphens only
- No XML tags
- Cannot contain reserved words:
"anthropic","claude"
Purpose: Serves as the command identifier when Claude invokes the Skill tool.
Examples:
name: processing-pdfs # Good - gerund form
name: pdf-processing # Good - noun phrase
name: PDF_Processing # Bad - uppercase
name: claude-helper # Bad - reserved word
description
Type: string Required: YES Max Length: 1024 characters per skill Constraints:
- Must be non-empty
- No XML tags
- Must use third person voice (injected into system prompt)
- CRITICAL BUDGET LIMIT: All skill descriptions combined have a 15,000-character context budget
Purpose: Primary signal for Claude's skill selection. Claude uses this to decide when to activate the skill.
⚠️ PRODUCTION CONSTRAINT FOR NIXTLA TEAMS:
The Skill tool's description field has a 15,000-character token budget across ALL skills in the workspace. If your combined skill descriptions exceed this limit, Claude will silently filter out skills, causing unpredictable skill discovery failures.
Critical Formula:
(Number of Skills) × (Avg Chars per Description) < 15,000 chars
Best Practices for Scaling Skill Portfolios:
- Target 300-400 characters per skill (not the 1024 maximum)
- Monitor total character count across all skill descriptions in workspace
- Verbose descriptions don't improve intent matching—specificity does
- If skills stop activating unexpectedly, audit total description length first
Scaling Examples:
10 skills × 400 chars = 4,000 chars (✅ safe, 11,000 chars headroom)
20 skills × 400 chars = 8,000 chars (✅ safe, 7,000 chars headroom)
30 skills × 400 chars = 12,000 chars (⚠️ risky, approaching limit)
40 skills × 400 chars = 16,000 chars (❌ exceeds budget, filtering likely)
20 skills × 750 chars = 15,000 chars (⚠️ at limit exactly, no headroom)
30 skills × 500 chars = 15,000 chars (⚠️ at limit exactly, no headroom)
Production Monitoring for Nixtla:
# Audit total description length across all skills
find .claude/skills/nixtla-*/SKILL.md -exec grep -A 5 '^description:' {} \; | wc -c
Formula:
[Primary capabilities]. [Secondary features]. Use when [scenarios]. Trigger with "[phrases]".
Good Examples:
description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction.
description: Generate descriptive commit messages by analyzing git diffs. Use when the user asks for help writing commit messages or reviewing staged changes.
description: Analyze Polymarket prediction market contracts using TimeGPT forecasting. Fetches contract odds, transforms to time series, generates price predictions with confidence intervals. Use when analyzing prediction markets, forecasting contract prices, or comparing platform odds. Trigger with 'forecast Polymarket', 'analyze prediction market'.
Bad Examples:
description: Helps with documents # Too vague
description: I can process your PDFs # Wrong voice (first person)
description: You can use this for data # Wrong voice (second person)
Optional Fields
allowed-tools
Type: CSV string Required: No Default: No pre-approved tools (user prompted for each)
Purpose: Pre-approves tools scoped to skill execution only. Tools revert to normal permissions after skill completes.
Syntax Examples:
# Multiple tools (comma-separated)
allowed-tools: "Read,Write,Glob,Grep,Edit"
# Scoped bash commands (restrict to specific commands)
allowed-tools: "Bash(git status:*),Bash(git diff:*),Read,Grep"
# NPM-scoped operations
allowed-tools: "Bash(npm:*),Bash(npx:*),Read,Write"
# Read-only audit
allowed-tools: "Read,Glob,Grep"
Security Principle: Grant ONLY tools the skill actually requires. Over-specifying creates unnecessary attack surface.
NOTE: Only supported in Claude Code, not claude.ai web version.
model
Type: string
Required: No
Default: "inherit" (use session model)
Purpose: Override the session model for skill execution.
Examples:
model: inherit # Use current session model (default)
model: "claude-opus-4-20250514" # Force specific model
model: "claude-sonnet-4-20250514" # Use Sonnet
Guidance: Reserve model overrides for genuinely complex tasks. Higher-capability models increase cost and latency.
version
Type: string (semver) Required: No Purpose: Version tracking for skill evolution.
Examples:
version: "1.0.0" # Initial release
version: "1.1.0" # New features
version: "2.0.0" # Breaking changes
license
Type: string Required: No Purpose: License terms reference.
Examples:
license: "MIT"
license: "Proprietary - See LICENSE.txt"
license: "Apache-2.0"
mode
Type: boolean
Required: No
Default: false
Purpose: When true, categorizes the skill as a "mode command" appearing in a prominent UI section separate from utility skills.
Use Case: Skills that fundamentally transform Claude's behavior for an extended session.
mode: true # Appears in "Mode Commands" section
mode: false # Appears in regular skills list (default)
disable-model-invocation
Type: boolean
Required: No
Default: false
Purpose: When true, removes the skill from the <available_skills> list. Users can still invoke manually via /skill-name.
Use Cases:
- Dangerous operations requiring explicit user action
- Infrastructure/deployment skills
- Skills that should never auto-activate
disable-model-invocation: true # Manual invocation only
disable-model-invocation: false # Auto-discovery enabled (default)
Undocumented/Experimental Fields
when_to_use
Status: UNDOCUMENTED - Avoid in production
Behavior: Appends to description with hyphen separator.
Recommendation: Do NOT use. Rely on detailed description field instead. This field may change or be removed without notice.
5. Instruction-Body Best Practices
Recommended Markdown Layout
# [Skill Name]
[1-2 sentence purpose statement]
## Overview
[What this skill does, when to use it, key capabilities - 3-5 sentences]
## Prerequisites
**Required**:
- [Tool/API/package 1]
- [Tool/API/package 2]
**Environment Variables**:
- `API_KEY_NAME`: [Description]
**Optional**:
- [Nice-to-have dependency]
## Instructions
### Step 1: [Action Verb]
[Clear, imperative instructions]
### Step 2: [Action Verb]
[More instructions]
## Output
This skill produces:
- [File/artifact 1]
- [File/artifact 2]
## Error Handling
**Common Failures**:
1. **Error**: [Error message or condition]
**Solution**: [How to fix]
2. **Error**: [Another failure]
**Solution**: [Resolution]
## Examples
### Example 1: [Scenario]
**Input**:
[Example input]
**Output**:
[Example output]
### Example 2: [Advanced Scenario]
[Another example]
## Resources
- Advanced patterns: `{baseDir}/references/ADVANCED.md`
- API reference: `{baseDir}/references/API_DOCS.md`
- Utility script: `{baseDir}/scripts/validate.py`
Content Guidelines
| Guideline | Requirement |
|---|---|
| Size Limit | Keep SKILL.md body under 500 lines |
| Word Count | Target ~1,500-2,500 words, max 5,000 words to avoid context saturation |
| Token Budget | Target ~2,500 tokens, max 5,000 tokens |
| Language | Use imperative voice ("Analyze data", not "You should analyze") |
| Paths | Always use {baseDir} variable, NEVER hardcode absolute paths |
| Examples | Include at least 2-3 concrete examples with input/output |
| Error Handling | Document 4+ common failures with solutions |
| Voice | Third person in descriptions, imperative in instructions |
⚠️ NIXTLA PRODUCTION GUIDANCE: The 5,000-word limit prevents context window saturation. Symptoms of oversized SKILL.md files include:
- Skills randomly failing to complete workflows
- Partial instruction execution
- Claude "forgetting" later sections of skill instructions
- Increased token costs per skill invocation
For production skill portfolios at scale: Aim for 1,500-2,000 words per SKILL.md to leave headroom for:
- Multiple skill invocations in one session
- User conversation context
- Tool output accumulation
- Growing skill count over time
Progressive Disclosure Patterns
⚠️ CRITICAL ENGINEERING INSIGHT (Anthropic Engineering Blog):
"The amount of context that can be bundled into a skill is effectively unbounded"
Three-Tier Disclosure Architecture:
- Tier 1: Metadata (name/description) - Pre-loaded in system prompt at startup (~100 chars)
- Tier 2: Full SKILL.md - Loaded when skill activates (~2,000 tokens)
- Tier 3: Referenced files - Loaded on-demand only when contextually necessary (0 tokens until needed)
Production Implication for Nixtla:
You can bundle MASSIVE reference materials without context penalty:
- Entire TimeGPT API documentation (10,000+ words)
- Complete model comparison tables
- Exhaustive troubleshooting guides
- Full example library
Key: Only Tier 1 (description) counts against the 15,000-char budget. Tier 2 and Tier 3 load dynamically, so total skill content can be "effectively unbounded."
Example (nixtla-timegpt-lab):
Tier 1: 400-char description (always loaded)
Tier 2: 2,000-word SKILL.md (loaded when activated)
Tier 3: references/
├── TIMEGPT_API_COMPLETE.md (15,000 words - loaded only if API questions)
├── TROUBLESHOOTING_GUIDE.md (8,000 words - loaded only if errors)
└── EXAMPLES_LIBRARY.md (20,000 words - loaded only if user asks for examples)
Total potential: 43,000+ words, but context cost = 400 chars + dynamic loading
When SKILL.md exceeds 400 lines, split content:
Pattern 1: High-level guide with references
# PDF Processing
## Quick start
[Basic instructions]
## Advanced features
**Form filling**: See [FORMS.md](FORMS.md)
**API reference**: See [REFERENCE.md](REFERENCE.md)
Pattern 2: Domain-specific organization
bigquery-skill/
├── SKILL.md (overview)
└── reference/
├── finance.md
├── sales.md
└── product.md
Pattern 3: Conditional details
For basic edits, modify XML directly.
**For tracked changes**: See [REDLINING.md](REDLINING.md)
Pattern 4: Mutually Exclusive Contexts (Anthropic Engineering)
Engineering Insight:
"This separation reduces token usage for mutually exclusive contexts."
When skill capabilities have contexts that are NEVER used together, split them into separate reference files. Only the contextually relevant file loads.
Example (PDF Skill from Anthropic):
pdf-skill/
├── SKILL.md (core capabilities)
├── reference.md (general PDF manipulation)
└── forms.md (form-filling workflows - ONLY loads if task involves forms)
If user asks "extract text from PDF", forms.md never loads (mutually exclusive context).
Nixtla Production Pattern:
nixtla-timegpt-lab/
├── SKILL.md (core orchestration)
├── references/
│ ├── TIMEGPT_FORECASTING.md (forecasting workflows)
│ ├── TIMEGPT_FINETUNING.md (fine-tuning workflows - mutually exclusive)
│ ├── TIMEGPT_ANOMALY_DETECTION.md (anomaly detection - mutually exclusive)
│ └── TIMEGPT_TROUBLESHOOTING.md (debugging - loaded only on errors)
Key Insight: If a user is fine-tuning, they're NOT doing anomaly detection. Don't waste context loading both.
Decision Tree:
Task: "Run TimeGPT forecast"
→ Loads: SKILL.md + TIMEGPT_FORECASTING.md
→ Skips: FINETUNING.md, ANOMALY_DETECTION.md, TROUBLESHOOTING.md
Task: "Fine-tune TimeGPT model"
→ Loads: SKILL.md + TIMEGPT_FINETUNING.md
→ Skips: FORECASTING.md, ANOMALY_DETECTION.md, TROUBLESHOOTING.md
Error encountered:
→ Additionally loads: TIMEGPT_TROUBLESHOOTING.md
Context Savings:
Without pattern: 43,000 words loaded (all reference files)
With pattern: 2,000 (SKILL.md) + 8,000 (relevant context) = 10,000 words
Savings: 76% context reduction
Critical Rule: One-Level-Deep References
AVOID deeply nested references. Claude may only partially read nested files.
Bad:
SKILL.md → advanced.md → details.md → actual_info.md
Good:
SKILL.md → advanced.md
SKILL.md → reference.md
SKILL.md → examples.md
6. Security & Safety Guidance
Choosing allowed-tools Conservatively
Principle of Least Privilege: Grant ONLY tools the skill actually needs.
⚠️ CRITICAL FOR NIXTLA PRODUCTION SKILLS: Tool permissions are scoped to skill execution only and automatically revert when the skill completes. This temporary escalation pattern ensures:
- Skills can't permanently expand Claude's attack surface
- Tool permissions return to user-controlled defaults after execution
- Multiple skill invocations in one session maintain isolation
Lifecycle Example:
1. User session starts → Standard tool permissions active
2. nixtla-experiment-architect invokes → allowed-tools: "Read,Write,Bash(python:*),Grep,Glob"
3. Skill executes experiments → Pre-approved tools available without prompts
4. Skill completes → Permissions revert to standard session defaults
5. Next skill invocation → New temporary permission scope
Good Examples:
# Read-only audit skill
allowed-tools: "Read,Glob,Grep"
# File transformation skill
allowed-tools: "Read,Write,Edit"
# Git operations only
allowed-tools: "Bash(git:*),Read,Grep"
Bad Examples:
# Overly permissive - unnecessary attack surface
allowed-tools: "Bash,Read,Write,Edit,Glob,Grep,WebSearch,Task,Agent"
# Unscoped bash - allows any command
allowed-tools: "Bash"
When to Use disable-model-invocation: true
Set this flag for skills that:
- Perform destructive operations (delete files, drop databases)
- Deploy to production environments
- Access sensitive credentials
- Run irreversible commands
- Should NEVER auto-activate
---
name: deploy-production
description: Deploy application to production. Dangerous - requires explicit invocation.
disable-model-invocation: true
allowed-tools: "Bash(deploy:*),Read,Glob"
---
Security Considerations
CRITICAL: Only use Skills from trusted sources.
Before using an untrusted skill:
- Review all bundled files (SKILL.md, scripts, resources)
- Check for unusual network calls
- Inspect scripts for malicious code
- Verify tool invocations match stated purpose
- Validate external URLs (if any)
Malicious skills could:
- Exfiltrate data via network calls
- Access unauthorized files
- Misuse tools (Bash for system manipulation)
- Inject instructions overriding safety guidelines
7. Model Selection Guidance
When to Inherit vs Override
| Scenario | Recommendation |
|---|---|
| Most skills | model: inherit or omit field |
| Complex reasoning required | Consider claude-opus-4-* |
| Fast, simple tasks | claude-haiku-* |
| Balanced performance | claude-sonnet-4-* |
Trade-offs
| Model | Speed | Cost | Capability |
|---|---|---|---|
| Haiku | Fast | Low | Basic tasks |
| Sonnet | Balanced | Medium | Most tasks |
| Opus | Slower | High | Complex reasoning |
Testing Across Models
Always test skills with all models you plan to use:
- Haiku: Does the skill provide sufficient guidance?
- Sonnet: Is content clear and efficient?
- Opus: Are instructions avoiding over-explanation?
What works for Opus may need more detail for Haiku.
8. Production-Readiness Checklist
Naming & Description
-
namematches folder name (lowercase + hyphens) -
nameis under 64 characters -
descriptionunder 1024 characters -
descriptionuses third person voice -
descriptionincludes what + when + trigger phrases - No reserved words (
anthropic,claude)
Structure & Tools
- SKILL.md at root of skill folder
- Body under 500 lines
- Uses
{baseDir}for all paths - No hardcoded absolute paths
-
allowed-toolsincludes only necessary tools - Forward slashes in all paths (not backslashes)
Instructions Quality
- Has all required sections (Overview, Prerequisites, Instructions, Output, Error Handling, Examples, Resources)
- Uses imperative voice
- 2-3 concrete examples with input/output
- 4+ common errors documented with solutions
- One-level-deep file references only
Testing
- Tested with Haiku, Sonnet, and Opus
- Trigger phrases activate skill correctly
- Scripts execute without errors
- Examples produce expected output
- No false positive activations
9. Versioning & Evolution
Semantic Versioning
MAJOR.MINOR.PATCH
│ │ └── Bug fixes, clarifications
│ └──────── New features, additive changes
└────────────── Breaking changes to interface
Examples:
1.0.0→ Initial release1.1.0→ Added new workflow step1.0.1→ Fixed typo in instructions2.0.0→ Changed output format (breaking)
Changelog Notes
Include version history in SKILL.md:
## Version History
- **v2.0.0** (2025-12-01): Breaking - Changed output format to JSON
- **v1.1.0** (2025-11-15): Added batch processing support
- **v1.0.0** (2025-11-01): Initial release
Deprecation Strategy
When deprecating a skill:
Add deprecation notice to description:
description: "[DEPRECATED - Use new-skill instead] Original description..."Set
disable-model-invocation: trueto prevent auto-activationKeep skill available for manual invocation during transition
Remove entirely in next major version
10. Canonical SKILL.md Template
---
name: your-skill-name
description: |
[Primary capabilities as action verbs]. [Secondary features].
Use when [3-4 trigger scenarios].
Trigger with "[phrase 1]", "[phrase 2]", "[phrase 3]".
allowed-tools: "Read,Write,Glob,Grep,Edit"
version: "1.0.0"
---
# [Skill Name]
[1-2 sentence purpose statement explaining what this skill does.]
## Overview
[3-5 sentences covering:]
- What this skill does
- When to use it
- Key capabilities
- What it produces
## Prerequisites
**Required**:
- [Tool/API/package 1]: [Brief purpose]
- [Tool/API/package 2]: [Brief purpose]
**Environment Variables**:
- `ENV_VAR_NAME`: [Description and how to obtain]
**Optional**:
- [Nice-to-have dependency]: [When needed]
## Instructions
### Step 1: [Action Verb - e.g., "Analyze Input"]
[Clear, imperative instructions for this step]
```bash
# Example command if applicable
python {baseDir}/scripts/step1.py --input data.json
Expected result: [What should happen]
Step 2: [Action Verb - e.g., "Transform Data"]
[Instructions for next step]
Step 3: [Action Verb - e.g., "Generate Output"]
[Final step instructions]
Output
This skill produces:
- [Artifact 1]: [Description and format]
- [Artifact 2]: [Description and format]
- [Report/Summary]: [Description]
Error Handling
Common Failures
Error:
[Error message or condition]Cause: [Why this happens] Solution: [How to fix]Error:
[Another error]Cause: [Reason] Solution: [Resolution]Error:
[Third error]Cause: [Reason] Solution: [Fix]Error:
[Fourth error]Cause: [Reason] Solution: [Fix]
Examples
Example 1: [Basic Scenario]
User Request: "[What user says]"
Input:
[Example input data]
Output:
[Expected output]
Example 2: [Advanced Scenario]
User Request: "[More complex request]"
Input:
[Input data]
Output:
[Expected result]
Resources
Reference Documentation:
- API reference:
{baseDir}/references/API_REFERENCE.md - Advanced patterns:
{baseDir}/references/ADVANCED.md
Utility Scripts:
- Data processor:
{baseDir}/scripts/process.py - Validator:
{baseDir}/scripts/validate.py
Templates:
- Report template:
{baseDir}/assets/report_template.md
Version History
- v1.0.0 (YYYY-MM-DD): Initial release
---
## 11. Minimal Example Skill
### Structured PR Review Helper
```yaml
---
name: reviewing-pull-requests
description: |
Analyze pull request diffs and generate structured code reviews.
Checks for bugs, security issues, performance problems, and style violations.
Use when reviewing PRs, analyzing code changes, or checking diffs.
Trigger with "review this PR", "check my code changes", "analyze diff".
allowed-tools: "Read,Grep,Glob,Bash(git:*)"
version: "1.0.0"
---
# Structured PR Review Helper
Generate comprehensive, structured code reviews from git diffs.
## Overview
This skill analyzes code changes and produces structured review feedback covering:
- Bug detection and edge cases
- Security vulnerabilities
- Performance considerations
- Code style and maintainability
- Test coverage gaps
## Prerequisites
**Required**:
- Git repository with staged or committed changes
- Read access to codebase
**Optional**:
- Project-specific style guide in `.github/STYLE_GUIDE.md`
## Instructions
### Step 1: Get the Diff
```bash
# For staged changes
git diff --staged
# For specific PR/branch
git diff main...feature-branch
Step 2: Analyze Each Changed File
For each modified file:
- Read the full file for context
- Identify the nature of changes (new feature, bug fix, refactor)
- Check for issues in each category
Step 3: Generate Structured Review
Produce review in this format:
## PR Review: [Brief Title]
### Summary
[1-2 sentence overview of changes]
### Findings
#### Critical Issues
- [ ] [Issue description with file:line reference]
#### Suggestions
- [ ] [Improvement suggestion]
#### Questions
- [ ] [Clarification needed]
### Recommendation
[APPROVE / REQUEST_CHANGES / COMMENT]
Output
- Structured markdown review
- Categorized findings (Critical, Suggestions, Questions)
- Clear recommendation
Error Handling
Error: No diff available Solution: Ensure changes are staged or specify correct branch comparison
Error: Binary files in diff Solution: Skip binary files, note in review that manual check needed
Error: Large diff (>1000 lines) Solution: Focus on most critical files, suggest splitting PR
Error: Missing context (file deleted) Solution: Note inability to review deleted code, focus on remaining changes
Examples
Example 1: Simple Bug Fix Review
User Request: "Review my bug fix PR"
Output:
## PR Review: Fix null pointer in user validation
### Summary
Adds null check before accessing user.email property.
### Findings
#### Critical Issues
- None
#### Suggestions
- [ ] Consider adding unit test for null user case (src/validation.ts:45)
#### Questions
- [ ] Should we also check for empty string?
### Recommendation
APPROVE - Good defensive fix, minor test suggestion
Example 2: Security-Sensitive Change
User Request: "Check this authentication change"
Output:
## PR Review: Update JWT token validation
### Summary
Modifies token expiration handling and adds refresh token support.
### Findings
#### Critical Issues
- [ ] Token secret exposed in error message (src/auth.ts:78)
- [ ] Missing rate limiting on refresh endpoint (src/routes/auth.ts:112)
#### Suggestions
- [ ] Add token rotation on refresh
- [ ] Log authentication failures for monitoring
### Recommendation
REQUEST_CHANGES - Security issues must be addressed
---
## 12. Author Checklist
Run through this checklist every time you create or update a skill:
### Before Writing
- [ ] Identified clear, single-purpose capability
- [ ] Confirmed no existing skill handles this
- [ ] Gathered all necessary reference materials
### Frontmatter
- [ ] `name`: lowercase, hyphens, under 64 chars, matches folder
- [ ] `description`: third person, under 1024 chars, includes what + when + triggers
- [ ] `allowed-tools`: minimal necessary tools only
- [ ] `version`: semver format
### Content
- [ ] Body under 500 lines
- [ ] All required sections present
- [ ] Imperative voice throughout instructions
- [ ] `{baseDir}` used for all paths
- [ ] 2-3 concrete examples with input/output
- [ ] 4+ errors documented with solutions
- [ ] One-level-deep references only
### Testing
- [ ] Triggers correctly on intended phrases
- [ ] Does NOT trigger on unrelated requests
- [ ] Scripts execute successfully
- [ ] Tested with multiple models (Haiku, Sonnet, Opus)
- [ ] Team review completed (if applicable)
### Security
- [ ] No secrets or credentials in skill
- [ ] Tools appropriately scoped
- [ ] Dangerous operations require explicit invocation
- [ ] External dependencies audited
---
## 13. Open Questions / Potentially Out-of-Date Areas
### Confirmed Speculative or Unclear
1. **`when_to_use` field**: Exists in codebase but undocumented. Behavior may ch
…(truncated)