# Global Master Standard – Claude Skills Specification

> A Claude Skill is a filesystem-based capability package containing instructions, executable code, and resources that Claude can discover and use automatically.

- Skill: `tools-only/global-master-standard-claude-skills-specification-2` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/global-master-standard-claude-skills-specification-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/global-master-standard-claude-skills-specification-2/raw
- Safety review: pending (external: skill-scanner PASS, skillspector CAUTION)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/global-master-standard-claude-skills-specification-2

---

# Global Master Standard – Claude Skills Specification

**Document ID**: 077-SPEC-MASTER-claude-skills-standard.md
**Version**: 2.3.0
**Status**: AUTHORITATIVE - Single Source of Truth
**Created**: 2025-12-06
**Updated**: 2025-12-08
**Audited Against**:
- Anthropic Engineering Blog (ENGINEERING SOURCE - oldest, deepest technical insights)
- Official Anthropic Skills Blog (PRIMARY SOURCE - newest product guidance)
- Lee Han Chung Deep Dive (implementation details)

**Sources**:
- [Official Anthropic Skills Blog Post](https://claude.com/blog/skills) ⭐ **PRIMARY SOURCE**
- [Anthropic Engineering Blog - Skills Deep Dive](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) ⭐ **ENGINEERING SOURCE**
- [Official Anthropic Agent Skills Overview](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview)
- [Official Anthropic Best Practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices)
- [Claude Code Skills Documentation](https://code.claude.com/docs/en/skills)
- [Lee Han Chung Deep Dive](https://leehanchung.github.io/blogs/2025/10/26/claude-skills-deep-dive/)

---

## Executive Summary

### What Is a Claude Skill?

A Claude Skill is a **filesystem-based capability package** containing instructions, executable code, and resources that Claude can discover and use automatically. Skills are prompt-based context modifiers—NOT executable plugins or slash commands.

**Official Definition** (Anthropic): "Specialized capability packages that extend Claude's functionality for specific tasks. Claude will only access a skill when it's relevant to the task at hand."

**Mental Model**: "Building a skill for an agent is like putting together an onboarding guide for a new hire."

### Four Core Design Principles (Anthropic Official)

Skills are architecturally designed with four foundational attributes:

1. **Composable**
   - Multiple skills work together seamlessly
   - Claude automatically identifies and coordinates which skills are needed
   - No manual orchestration required
   - **Nixtla Impact**: `nixtla-schema-mapper` + `nixtla-experiment-architect` can chain automatically

2. **Portable**
   - Same skill format works across ALL platforms:
     - Claude apps (claude.ai web/mobile)
     - Claude Code (CLI tool)
     - Claude API (Messages API)
   - Write once, deploy everywhere
   - **Nixtla Impact**: Internal skills work for web users, API customers, and CLI developers

3. **Efficient**
   - Loads only minimal necessary information when needed
   - Progressive disclosure pattern (description → full instructions → resources)
   - Avoids context window bloat
   - **Nixtla Impact**: 40+ skill portfolio won't overwhelm context if designed properly

4. **Powerful**
   - Can include executable code for reliable task performance
   - Pre-written scripts eliminate non-determinism
   - Combines prompts + execution for complete workflows
   - **Nixtla Impact**: TimeGPT API calls, pandas transformations, validation scripts all executable

**Code Execution Performance Economics** (Anthropic Engineering):

> **"Sorting a list via token generation is far more expensive than simply running a sorting algorithm."**

Deterministic code execution provides **orders of magnitude** cost/performance advantages:

| Operation | Token Generation | Code Execution | Advantage |
|-----------|------------------|----------------|-----------|
| **Sort 1000 items** | ~2,000 tokens (~$0.006) | <10 tokens for script call (~$0.00003) | **200x cheaper** |
| **Pandas transform** | ~5,000 tokens loading data + code | Script runs without context load | **No context cost** |
| **TimeGPT API call** | Non-deterministic, requires retry logic | Deterministic, runs once | **Reliable + cheap** |

**Nixtla Production Impact**:
- TimeGPT forecasting loops: Execute via script, not token generation
- Schema validation: Run validation.py, don't describe validation in tokens
- Data transformations: pandas scripts consume ZERO context tokens
- Experiment harnesses: Pre-written Python orchestration

**Production Implication for Nixtla**: These four principles guide ALL architectural decisions. Violating any principle (non-composable, platform-locked, context-heavy, prompt-only) indicates poor skill design.

### Why Use Skills Instead of Ad-Hoc Prompts?

| Aspect | Ad-Hoc Prompts | Skills |
|--------|---------------|--------|
| Reusability | One conversation | Persistent across all conversations |
| Discovery | Manual context provision | Automatic activation based on intent |
| Organization | Scattered knowledge | Structured packages |
| Context Management | Full context loaded | Progressive disclosure (on-demand) |
| Code Integration | Generated each time | Pre-written, deterministic scripts |

### Where Skills Live

| Location | Scope | Priority |
|----------|-------|----------|
| `~/.claude/skills/` | Personal (all projects) | 1 (lowest) |
| `.claude/skills/` | Project-specific | 2 |
| Plugin `skills/` directory | Plugin-bundled | 3 |
| Built-in skills | Platform-provided | 4 (highest) |

Later sources override earlier ones when names conflict.

---

## 1. Core Concepts

### Skill = What + When + How + Allowed Tools + Optional Model Override

Every skill answers:
- **What**: What capability does this provide?
- **When**: When should Claude activate it?
- **How**: Step-by-step instructions for Claude
- **Allowed Tools**: Which tools are pre-approved during execution?
- **Model Override**: Should a different model handle this? (optional)

### The Skill Tool Architecture

**Critical insight**: Skills are NOT in the system prompt.

Skills live in a meta-tool called `Skill` within the `tools` array:

```javascript
tools: [
  { name: "Read", ... },
  { name: "Write", ... },
  {
    name: "Skill",                    // Meta-tool (capital S)
    inputSchema: { command: string },
    description: "<available_skills>..." // Dynamic list of all skill descriptions
  }
]
```

### How Skills Are Discovered and Invoked

**Model-Invoked (Automatic)**:
1. At startup, Claude's system prompt includes metadata (name + description) for all skills
2. Claude reads user request and matches intent to skill descriptions
3. Claude invokes `Skill` tool with matching `command` parameter
4. No algorithmic routing, embeddings, or keyword matching—**pure LLM reasoning**

**User-Invoked (Manual)**:
- Type `/skill-name` to explicitly invoke a skill
- Required when `disable-model-invocation: true`

### Message Injection Architecture

**CRITICAL FOR NIXTLA INTERNAL TEAMS**: Understanding how skills inject into conversations ensures predictable behavior in production workflows.

When a skill is invoked, it injects **two user messages** into the conversation:

#### Message 1: Metadata Message (Visible to User)
```javascript
{
  role: "user",
  isMeta: false,  // Visible in UI
  content: [
    {
      type: "text",
      text: "<command-message>skill-name is loading...</command-message>"
    }
  ]
}
```

**Purpose**: Provides transparent UI feedback that a skill is executing.

#### Message 2: Skill Prompt Message (Hidden from User)
```javascript
{
  role: "user",
  isMeta: true,   // Hidden from UI, sent to API only
  content: [
    {
      type: "text",
      text: "[Full SKILL.md content with {baseDir} substitutions]"
    }
  ]
}
```

**Purpose**: Injects skill instructions directly into Claude's context for reasoning.

#### XML Tag Structure

Skills use specific XML tags for message formatting:
- `<command-message>` - Wraps the visible status indicator
- `<command-name>` - Contains the skill name for tracking
- `<available_skills>` - Lists all discoverable skills in Skill tool description

**Production Impact for Nixtla**:
- Skills don't pollute conversation history visible to users
- Skill instructions consume context budget (tracked via `isMeta: true` messages)
- Multiple skill invocations in one session stack context linearly
- **Context budget management is critical** (see Section 4 for limits)

#### Chain-of-Thought Visibility (UX Insight)

**OFFICIAL ANTHROPIC**: "Users can view skills in Claude's chain-of-thought output while it works."

When Claude invokes skills, users see them in the reasoning trace/thinking blocks. This provides:
- **Transparency**: Users know which skills are active
- **Debuggability**: Developers can trace skill activation patterns
- **Trust**: Users understand Claude's decision-making process

**Nixtla Customer Impact**:
- TimeGPT customers will SEE when `nixtla-experiment-architect` activates
- Internal teams can debug skill selection by reviewing chain-of-thought
- Enterprise admins can audit which skills employees trigger

---

## 2. Folder & Discovery Layout

### Standard Directory Structure

```
skill-name/
├── SKILL.md              # REQUIRED - Instructions + YAML frontmatter
├── scripts/              # OPTIONAL - Executable Python/Bash scripts
│   ├── analyze.py
│   └── validate.py
├── references/           # OPTIONAL - Docs loaded into context
│   ├── API_REFERENCE.md
│   └── EXAMPLES.md
├── assets/               # OPTIONAL - Templates referenced by path
│   └── report_template.md
└── LICENSE.txt           # OPTIONAL - License terms
```

### Naming Conventions

**Folder names must match the `name` field exactly.**

**Recommended**: Use **gerund form** (verb + -ing) for clarity:
- `processing-pdfs`
- `analyzing-spreadsheets`
- `generating-commit-messages`

**Acceptable alternatives**:
- Noun phrases: `pdf-processing`, `data-analysis`
- Action-oriented: `process-pdfs`, `analyze-data`

**Avoid**:
- Vague names: `helper`, `utils`, `tools`
- Generic names: `documents`, `data`, `files`
- Reserved words: `anthropic-*`, `claude-*`

### Directory Purposes

| Directory | Purpose | Loaded Into Context? | Token Cost |
|-----------|---------|---------------------|------------|
| `004-scripts/` | Executable code (deterministic operations) | No (executed via Bash) | None |
| `references/` | Documentation (API docs, examples) | Yes (via Read tool) | High |
| `assets/` | Templates, configs, static files | No (path reference only) | None |

**Key Insight**: Scripts execute without loading code into context. Only script OUTPUT consumes tokens.

---

### Platform-Specific Implementation & Distribution

#### Claude Apps (claude.ai, mobile)

**Availability**: Pro, Max, Team, and Enterprise users

**Installation Methods**:
1. **Official Skill-Creator Tool** (Anthropic Recommended)
   - Use the built-in `skill-creator` skill for interactive guidance
   - No manual file editing required
   - Provides step-by-step skill creation workflow
   - Available via slash command: `/skill-creator`

2. **Manual Installation**:
   - Create `.claude/skills/skill-name/SKILL.md` structure
   - Upload to project or personal workspace

**Admin Controls** (Team/Enterprise):
- Admins must **enable skills organization-wide** via Settings
- Organization-level toggle controls all member access
- **NIXTLA ENTERPRISE NOTE**: Customers need admin approval before deploying skills

**Official Pre-built Skills** (Anthropic-provided):
- Excel processing
- PowerPoint manipulation
- Word document editing
- Fillable PDF handling

**Nixtla Strategy**: Leverage pre-built skills where possible; avoid reinventing document processing.

---

#### Claude API (Messages API)

**⚠️ CRITICAL FOR NIXTLA API INTEGRATIONS**:

**New `/v1/skills` Endpoint**:
- Programmatic skill control, versioning, and management
- Separate from Messages API calls
- Enables automated skill deployment and updates

**Code Execution Tool Requirement**:
- Skills with executable scripts **require Code Execution Tool beta**
- Must be enabled in API configuration
- Provides secure sandbox for script execution
- **BLOCKER**: Nixtla's TimeGPT/pandas/validation scripts won't work without this beta feature

**Implementation Pattern** (Nixtla API Integration):
```python
# 1. Enable Code Execution Tool beta in API settings
# 2. Deploy skills via /v1/skills endpoint
# 3. Reference skills in Messages API calls

import anthropic

client = anthropic.Anthropic(api_key="...")

# Deploy skill (one-time)
skill_response = client.skills.create(
    skill_content=skill_md_content,
    version="1.0.0"
)

# Use skill in conversation
message = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=4096,
    messages=[{"role": "user", "content": "Map my data to Nixtla schema"}],
    # Skills automatically discovered and invoked
)
```

**Production Requirements for Nixtla**:
1. Request Code Execution Tool beta access from Anthropic
2. Implement `/v1/skills` deployment automation
3. Version all skills semantically (1.0.0, 1.1.0, etc.)
4. Test scripts in sandbox environment before production

---

#### Claude Code (CLI)

**Installation Methods**:
1. **Marketplace Plugins** (Anthropic Official):
   - Install via `anthropics/skills` marketplace
   - Managed updates and versioning
   - **NIXTLA DISTRIBUTION**: Consider publishing nixtla-skills as marketplace plugin

2. **Manual Installation**:
   - Place skills in `~/.claude/skills/` (personal, all projects)
   - Place skills in `.claude/skills/` (project-specific)
   - **Version control friendly**: Commit `.claude/skills/` to git repos

**Priority Hierarchy**:
```
~/.claude/skills/          Priority 1 (lowest)
.claude/skills/            Priority 2
Plugin skills/             Priority 3
Built-in skills            Priority 4 (highest)
```

Later sources override earlier ones when names conflict.

**Nixtla Distribution Strategy**:
- Primary: Project-specific (`.claude/skills/nixtla-*/`)
- Fallback: Personal installation (`~/.claude/skills/nixtla-*/`)
- Future: Marketplace plugin for external distribution

---

## 3. SKILL.md Specification

### Complete Structure

```yaml
---
name: skill-name
description: What this skill does. Use when [conditions]. Trigger with "[phrases]".
---

# Skill Name

Brief purpose statement (1-2 sentences).

## Overview

What this skill does, when to use it, key capabilities.

## Prerequisites

Required tools, APIs, environment variables, packages.

## Instructions

### Step 1: [Action Verb]
[Imperative instructions]

### Step 2: [Action Verb]
[More instructions]

## Output

What artifacts this skill produces.

## Error Handling

Common failures and solutions.

## Examples

Concrete usage examples with input/output.

## Resources

Links to bundled files using {baseDir} variable.
```

---

## 4. YAML Frontmatter Fields

### Required Fields

#### `name`

**Type**: string
**Required**: YES
**Max Length**: 64 characters
**Constraints**:
- Lowercase letters, numbers, and hyphens only
- No XML tags
- Cannot contain reserved words: `"anthropic"`, `"claude"`

**Purpose**: Serves as the command identifier when Claude invokes the Skill tool.

**Examples**:
```yaml
name: processing-pdfs          # Good - gerund form
name: pdf-processing           # Good - noun phrase
name: PDF_Processing           # Bad - uppercase
name: claude-helper            # Bad - reserved word
```

#### `description`

**Type**: string
**Required**: YES
**Max Length**: 1024 characters per skill
**Constraints**:
- Must be non-empty
- No XML tags
- Must use **third person** voice (injected into system prompt)
- **CRITICAL BUDGET LIMIT**: All skill descriptions combined have a 15,000-character context budget

**Purpose**: Primary signal for Claude's skill selection. Claude uses this to decide when to activate the skill.

**⚠️ PRODUCTION CONSTRAINT FOR NIXTLA TEAMS**:

The Skill tool's description field has a **15,000-character token budget** across ALL skills in the workspace. If your combined skill descriptions exceed this limit, Claude will silently filter out skills, causing unpredictable skill discovery failures.

**Critical Formula**:
```
(Number of Skills) × (Avg Chars per Description) < 15,000 chars
```

**Best Practices for Scaling Skill Portfolios**:
- **Target 300-400 characters per skill** (not the 1024 maximum)
- **Monitor total character count** across all skill descriptions in workspace
- Verbose descriptions don't improve intent matching—specificity does
- If skills stop activating unexpectedly, audit total description length first

**Scaling Examples**:
```
10 skills × 400 chars = 4,000 chars   (✅ safe, 11,000 chars headroom)
20 skills × 400 chars = 8,000 chars   (✅ safe, 7,000 chars headroom)
30 skills × 400 chars = 12,000 chars  (⚠️ risky, approaching limit)
40 skills × 400 chars = 16,000 chars  (❌ exceeds budget, filtering likely)

20 skills × 750 chars = 15,000 chars  (⚠️ at limit exactly, no headroom)
30 skills × 500 chars = 15,000 chars  (⚠️ at limit exactly, no headroom)
```

**Production Monitoring for Nixtla**:
```bash
# Audit total description length across all skills
find .claude/skills/nixtla-*/SKILL.md -exec grep -A 5 '^description:' {} \; | wc -c
```

**Formula**:
```
[Primary capabilities]. [Secondary features]. Use when [scenarios]. Trigger with "[phrases]".
```

**Good Examples**:
```yaml
description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction.

description: Generate descriptive commit messages by analyzing git diffs. Use when the user asks for help writing commit messages or reviewing staged changes.

description: Analyze Polymarket prediction market contracts using TimeGPT forecasting. Fetches contract odds, transforms to time series, generates price predictions with confidence intervals. Use when analyzing prediction markets, forecasting contract prices, or comparing platform odds. Trigger with 'forecast Polymarket', 'analyze prediction market'.
```

**Bad Examples**:
```yaml
description: Helps with documents          # Too vague
description: I can process your PDFs       # Wrong voice (first person)
description: You can use this for data     # Wrong voice (second person)
```

### Optional Fields

#### `allowed-tools`

**Type**: CSV string
**Required**: No
**Default**: No pre-approved tools (user prompted for each)

**Purpose**: Pre-approves tools **scoped to skill execution only**. Tools revert to normal permissions after skill completes.

**Syntax Examples**:
```yaml
# Multiple tools (comma-separated)
allowed-tools: "Read,Write,Glob,Grep,Edit"

# Scoped bash commands (restrict to specific commands)
allowed-tools: "Bash(git status:*),Bash(git diff:*),Read,Grep"

# NPM-scoped operations
allowed-tools: "Bash(npm:*),Bash(npx:*),Read,Write"

# Read-only audit
allowed-tools: "Read,Glob,Grep"
```

**Security Principle**: Grant ONLY tools the skill actually requires. Over-specifying creates unnecessary attack surface.

**NOTE**: Only supported in Claude Code, not claude.ai web version.

#### `model`

**Type**: string
**Required**: No
**Default**: `"inherit"` (use session model)

**Purpose**: Override the session model for skill execution.

**Examples**:
```yaml
model: inherit                           # Use current session model (default)
model: "claude-opus-4-20250514"          # Force specific model
model: "claude-sonnet-4-20250514"        # Use Sonnet
```

**Guidance**: Reserve model overrides for genuinely complex tasks. Higher-capability models increase cost and latency.

#### `version`

**Type**: string (semver)
**Required**: No
**Purpose**: Version tracking for skill evolution.

**Examples**:
```yaml
version: "1.0.0"    # Initial release
version: "1.1.0"    # New features
version: "2.0.0"    # Breaking changes
```

#### `license`

**Type**: string
**Required**: No
**Purpose**: License terms reference.

**Examples**:
```yaml
license: "MIT"
license: "Proprietary - See LICENSE.txt"
license: "Apache-2.0"
```

#### `mode`

**Type**: boolean
**Required**: No
**Default**: `false`

**Purpose**: When `true`, categorizes the skill as a "mode command" appearing in a prominent UI section separate from utility skills.

**Use Case**: Skills that fundamentally transform Claude's behavior for an extended session.

```yaml
mode: true     # Appears in "Mode Commands" section
mode: false    # Appears in regular skills list (default)
```

#### `disable-model-invocation`

**Type**: boolean
**Required**: No
**Default**: `false`

**Purpose**: When `true`, removes the skill from the `<available_skills>` list. Users can still invoke manually via `/skill-name`.

**Use Cases**:
- Dangerous operations requiring explicit user action
- Infrastructure/deployment skills
- Skills that should never auto-activate

```yaml
disable-model-invocation: true    # Manual invocation only
disable-model-invocation: false   # Auto-discovery enabled (default)
```

### Undocumented/Experimental Fields

#### `when_to_use`

**Status**: UNDOCUMENTED - Avoid in production

**Behavior**: Appends to `description` with hyphen separator.

**Recommendation**: Do NOT use. Rely on detailed `description` field instead. This field may change or be removed without notice.

---

## 5. Instruction-Body Best Practices

### Recommended Markdown Layout

```markdown
# [Skill Name]

[1-2 sentence purpose statement]

## Overview

[What this skill does, when to use it, key capabilities - 3-5 sentences]

## Prerequisites

**Required**:
- [Tool/API/package 1]
- [Tool/API/package 2]

**Environment Variables**:
- `API_KEY_NAME`: [Description]

**Optional**:
- [Nice-to-have dependency]

## Instructions

### Step 1: [Action Verb]

[Clear, imperative instructions]

### Step 2: [Action Verb]

[More instructions]

## Output

This skill produces:
- [File/artifact 1]
- [File/artifact 2]

## Error Handling

**Common Failures**:

1. **Error**: [Error message or condition]
   **Solution**: [How to fix]

2. **Error**: [Another failure]
   **Solution**: [Resolution]

## Examples

### Example 1: [Scenario]

**Input**:
[Example input]

**Output**:
[Example output]

### Example 2: [Advanced Scenario]

[Another example]

## Resources

- Advanced patterns: `{baseDir}/references/ADVANCED.md`
- API reference: `{baseDir}/references/API_DOCS.md`
- Utility script: `{baseDir}/scripts/validate.py`
```

### Content Guidelines

| Guideline | Requirement |
|-----------|-------------|
| **Size Limit** | Keep SKILL.md body under **500 lines** |
| **Word Count** | Target ~1,500-2,500 words, **max 5,000 words** to avoid context saturation |
| **Token Budget** | Target ~2,500 tokens, max 5,000 tokens |
| **Language** | Use **imperative voice** ("Analyze data", not "You should analyze") |
| **Paths** | Always use `{baseDir}` variable, NEVER hardcode absolute paths |
| **Examples** | Include at least **2-3 concrete examples** with input/output |
| **Error Handling** | Document **4+ common failures** with solutions |
| **Voice** | Third person in descriptions, imperative in instructions |

**⚠️ NIXTLA PRODUCTION GUIDANCE**: The 5,000-word limit prevents context window saturation. Symptoms of oversized SKILL.md files include:
- Skills randomly failing to complete workflows
- Partial instruction execution
- Claude "forgetting" later sections of skill instructions
- Increased token costs per skill invocation

**For production skill portfolios at scale**: Aim for 1,500-2,000 words per SKILL.md to leave headroom for:
- Multiple skill invocations in one session
- User conversation context
- Tool output accumulation
- Growing skill count over time

### Progressive Disclosure Patterns

**⚠️ CRITICAL ENGINEERING INSIGHT** (Anthropic Engineering Blog):

> **"The amount of context that can be bundled into a skill is effectively unbounded"**

**Three-Tier Disclosure Architecture**:
1. **Tier 1**: Metadata (name/description) - Pre-loaded in system prompt at startup (~100 chars)
2. **Tier 2**: Full SKILL.md - Loaded when skill activates (~2,000 tokens)
3. **Tier 3**: Referenced files - Loaded on-demand only when contextually necessary (0 tokens until needed)

**Production Implication for Nixtla**:

You can bundle MASSIVE reference materials without context penalty:
- Entire TimeGPT API documentation (10,000+ words)
- Complete model comparison tables
- Exhaustive troubleshooting guides
- Full example library

**Key**: Only Tier 1 (description) counts against the 15,000-char budget. Tier 2 and Tier 3 load dynamically, so total skill content can be "effectively unbounded."

**Example** (nixtla-timegpt-lab):
```
Tier 1: 400-char description (always loaded)
Tier 2: 2,000-word SKILL.md (loaded when activated)
Tier 3: references/
  ├── TIMEGPT_API_COMPLETE.md (15,000 words - loaded only if API questions)
  ├── TROUBLESHOOTING_GUIDE.md (8,000 words - loaded only if errors)
  └── EXAMPLES_LIBRARY.md (20,000 words - loaded only if user asks for examples)

Total potential: 43,000+ words, but context cost = 400 chars + dynamic loading
```

---

**When SKILL.md exceeds 400 lines, split content:**

**Pattern 1: High-level guide with references**
```markdown
# PDF Processing

## Quick start
[Basic instructions]

## Advanced features
**Form filling**: See [FORMS.md](FORMS.md)
**API reference**: See [REFERENCE.md](REFERENCE.md)
```

**Pattern 2: Domain-specific organization**
```
bigquery-skill/
├── SKILL.md (overview)
└── reference/
    ├── finance.md
    ├── sales.md
    └── product.md
```

**Pattern 3: Conditional details**
```markdown
For basic edits, modify XML directly.
**For tracked changes**: See [REDLINING.md](REDLINING.md)
```

**Pattern 4: Mutually Exclusive Contexts** (Anthropic Engineering)

**Engineering Insight**:
> "This separation reduces token usage for mutually exclusive contexts."

When skill capabilities have contexts that are NEVER used together, split them into separate reference files. Only the contextually relevant file loads.

**Example** (PDF Skill from Anthropic):
```
pdf-skill/
├── SKILL.md (core capabilities)
├── reference.md (general PDF manipulation)
└── forms.md (form-filling workflows - ONLY loads if task involves forms)
```

If user asks "extract text from PDF", `forms.md` never loads (mutually exclusive context).

**Nixtla Production Pattern**:
```
nixtla-timegpt-lab/
├── SKILL.md (core orchestration)
├── references/
│   ├── TIMEGPT_FORECASTING.md (forecasting workflows)
│   ├── TIMEGPT_FINETUNING.md (fine-tuning workflows - mutually exclusive)
│   ├── TIMEGPT_ANOMALY_DETECTION.md (anomaly detection - mutually exclusive)
│   └── TIMEGPT_TROUBLESHOOTING.md (debugging - loaded only on errors)
```

**Key Insight**: If a user is fine-tuning, they're NOT doing anomaly detection. Don't waste context loading both.

**Decision Tree**:
```
Task: "Run TimeGPT forecast"
→ Loads: SKILL.md + TIMEGPT_FORECASTING.md
→ Skips: FINETUNING.md, ANOMALY_DETECTION.md, TROUBLESHOOTING.md

Task: "Fine-tune TimeGPT model"
→ Loads: SKILL.md + TIMEGPT_FINETUNING.md
→ Skips: FORECASTING.md, ANOMALY_DETECTION.md, TROUBLESHOOTING.md

Error encountered:
→ Additionally loads: TIMEGPT_TROUBLESHOOTING.md
```

**Context Savings**:
```
Without pattern: 43,000 words loaded (all reference files)
With pattern: 2,000 (SKILL.md) + 8,000 (relevant context) = 10,000 words
Savings: 76% context reduction
```

### Critical Rule: One-Level-Deep References

**AVOID deeply nested references**. Claude may only partially read nested files.

**Bad**:
```
SKILL.md → advanced.md → details.md → actual_info.md
```

**Good**:
```
SKILL.md → advanced.md
SKILL.md → reference.md
SKILL.md → examples.md
```

---

## 6. Security & Safety Guidance

### Choosing `allowed-tools` Conservatively

**Principle of Least Privilege**: Grant ONLY tools the skill actually needs.

**⚠️ CRITICAL FOR NIXTLA PRODUCTION SKILLS**: Tool permissions are **scoped to skill execution only** and **automatically revert** when the skill completes. This temporary escalation pattern ensures:
- Skills can't permanently expand Claude's attack surface
- Tool permissions return to user-controlled defaults after execution
- Multiple skill invocations in one session maintain isolation

**Lifecycle Example**:
```
1. User session starts → Standard tool permissions active
2. nixtla-experiment-architect invokes → allowed-tools: "Read,Write,Bash(python:*),Grep,Glob"
3. Skill executes experiments → Pre-approved tools available without prompts
4. Skill completes → Permissions revert to standard session defaults
5. Next skill invocation → New temporary permission scope
```

**Good Examples**:
```yaml
# Read-only audit skill
allowed-tools: "Read,Glob,Grep"

# File transformation skill
allowed-tools: "Read,Write,Edit"

# Git operations only
allowed-tools: "Bash(git:*),Read,Grep"
```

**Bad Examples**:
```yaml
# Overly permissive - unnecessary attack surface
allowed-tools: "Bash,Read,Write,Edit,Glob,Grep,WebSearch,Task,Agent"

# Unscoped bash - allows any command
allowed-tools: "Bash"
```

### When to Use `disable-model-invocation: true`

Set this flag for skills that:
- Perform destructive operations (delete files, drop databases)
- Deploy to production environments
- Access sensitive credentials
- Run irreversible commands
- Should NEVER auto-activate

```yaml
---
name: deploy-production
description: Deploy application to production. Dangerous - requires explicit invocation.
disable-model-invocation: true
allowed-tools: "Bash(deploy:*),Read,Glob"
---
```

### Security Considerations

**CRITICAL**: Only use Skills from trusted sources.

Before using an untrusted skill:
- [ ] Review all bundled files (SKILL.md, scripts, resources)
- [ ] Check for unusual network calls
- [ ] Inspect scripts for malicious code
- [ ] Verify tool invocations match stated purpose
- [ ] Validate external URLs (if any)

**Malicious skills could**:
- Exfiltrate data via network calls
- Access unauthorized files
- Misuse tools (Bash for system manipulation)
- Inject instructions overriding safety guidelines

---

## 7. Model Selection Guidance

### When to Inherit vs Override

| Scenario | Recommendation |
|----------|----------------|
| Most skills | `model: inherit` or omit field |
| Complex reasoning required | Consider `claude-opus-4-*` |
| Fast, simple tasks | `claude-haiku-*` |
| Balanced performance | `claude-sonnet-4-*` |

### Trade-offs

| Model | Speed | Cost | Capability |
|-------|-------|------|------------|
| Haiku | Fast | Low | Basic tasks |
| Sonnet | Balanced | Medium | Most tasks |
| Opus | Slower | High | Complex reasoning |

### Testing Across Models

**Always test skills with all models you plan to use:**

- **Haiku**: Does the skill provide sufficient guidance?
- **Sonnet**: Is content clear and efficient?
- **Opus**: Are instructions avoiding over-explanation?

What works for Opus may need more detail for Haiku.

---

## 8. Production-Readiness Checklist

### Naming & Description

- [ ] `name` matches folder name (lowercase + hyphens)
- [ ] `name` is under 64 characters
- [ ] `description` under 1024 characters
- [ ] `description` uses third person voice
- [ ] `description` includes what + when + trigger phrases
- [ ] No reserved words (`anthropic`, `claude`)

### Structure & Tools

- [ ] SKILL.md at root of skill folder
- [ ] Body under 500 lines
- [ ] Uses `{baseDir}` for all paths
- [ ] No hardcoded absolute paths
- [ ] `allowed-tools` includes only necessary tools
- [ ] Forward slashes in all paths (not backslashes)

### Instructions Quality

- [ ] Has all required sections (Overview, Prerequisites, Instructions, Output, Error Handling, Examples, Resources)
- [ ] Uses imperative voice
- [ ] 2-3 concrete examples with input/output
- [ ] 4+ common errors documented with solutions
- [ ] One-level-deep file references only

### Testing

- [ ] Tested with Haiku, Sonnet, and Opus
- [ ] Trigger phrases activate skill correctly
- [ ] Scripts execute without errors
- [ ] Examples produce expected output
- [ ] No false positive activations

---

## 9. Versioning & Evolution

### Semantic Versioning

```
MAJOR.MINOR.PATCH
  │     │     └── Bug fixes, clarifications
  │     └──────── New features, additive changes
  └────────────── Breaking changes to interface
```

**Examples**:
- `1.0.0` → Initial release
- `1.1.0` → Added new workflow step
- `1.0.1` → Fixed typo in instructions
- `2.0.0` → Changed output format (breaking)

### Changelog Notes

Include version history in SKILL.md:

```markdown
## Version History

- **v2.0.0** (2025-12-01): Breaking - Changed output format to JSON
- **v1.1.0** (2025-11-15): Added batch processing support
- **v1.0.0** (2025-11-01): Initial release
```

### Deprecation Strategy

When deprecating a skill:

1. Add deprecation notice to description:
   ```yaml
   description: "[DEPRECATED - Use new-skill instead] Original description..."
   ```

2. Set `disable-model-invocation: true` to prevent auto-activation

3. Keep skill available for manual invocation during transition

4. Remove entirely in next major version

---

## 10. Canonical SKILL.md Template

```yaml
---
name: your-skill-name
description: |
  [Primary capabilities as action verbs]. [Secondary features].
  Use when [3-4 trigger scenarios].
  Trigger with "[phrase 1]", "[phrase 2]", "[phrase 3]".
allowed-tools: "Read,Write,Glob,Grep,Edit"
version: "1.0.0"
---

# [Skill Name]

[1-2 sentence purpose statement explaining what this skill does.]

## Overview

[3-5 sentences covering:]
- What this skill does
- When to use it
- Key capabilities
- What it produces

## Prerequisites

**Required**:
- [Tool/API/package 1]: [Brief purpose]
- [Tool/API/package 2]: [Brief purpose]

**Environment Variables**:
- `ENV_VAR_NAME`: [Description and how to obtain]

**Optional**:
- [Nice-to-have dependency]: [When needed]

## Instructions

### Step 1: [Action Verb - e.g., "Analyze Input"]

[Clear, imperative instructions for this step]

```bash
# Example command if applicable
python {baseDir}/scripts/step1.py --input data.json
```

**Expected result**: [What should happen]

### Step 2: [Action Verb - e.g., "Transform Data"]

[Instructions for next step]

### Step 3: [Action Verb - e.g., "Generate Output"]

[Final step instructions]

## Output

This skill produces:

- **[Artifact 1]**: [Description and format]
- **[Artifact 2]**: [Description and format]
- **[Report/Summary]**: [Description]

## Error Handling

### Common Failures

1. **Error**: `[Error message or condition]`
   **Cause**: [Why this happens]
   **Solution**: [How to fix]

2. **Error**: `[Another error]`
   **Cause**: [Reason]
   **Solution**: [Resolution]

3. **Error**: `[Third error]`
   **Cause**: [Reason]
   **Solution**: [Fix]

4. **Error**: `[Fourth error]`
   **Cause**: [Reason]
   **Solution**: [Fix]

## Examples

### Example 1: [Basic Scenario]

**User Request**: "[What user says]"

**Input**:
```
[Example input data]
```

**Output**:
```
[Expected output]
```

### Example 2: [Advanced Scenario]

**User Request**: "[More complex request]"

**Input**:
```
[Input data]
```

**Output**:
```
[Expected result]
```

## Resources

**Reference Documentation**:
- API reference: `{baseDir}/references/API_REFERENCE.md`
- Advanced patterns: `{baseDir}/references/ADVANCED.md`

**Utility Scripts**:
- Data processor: `{baseDir}/scripts/process.py`
- Validator: `{baseDir}/scripts/validate.py`

**Templates**:
- Report template: `{baseDir}/assets/report_template.md`

## Version History

- **v1.0.0** (YYYY-MM-DD): Initial release
```

---

## 11. Minimal Example Skill

### Structured PR Review Helper

```yaml
---
name: reviewing-pull-requests
description: |
  Analyze pull request diffs and generate structured code reviews.
  Checks for bugs, security issues, performance problems, and style violations.
  Use when reviewing PRs, analyzing code changes, or checking diffs.
  Trigger with "review this PR", "check my code changes", "analyze diff".
allowed-tools: "Read,Grep,Glob,Bash(git:*)"
version: "1.0.0"
---

# Structured PR Review Helper

Generate comprehensive, structured code reviews from git diffs.

## Overview

This skill analyzes code changes and produces structured review feedback covering:
- Bug detection and edge cases
- Security vulnerabilities
- Performance considerations
- Code style and maintainability
- Test coverage gaps

## Prerequisites

**Required**:
- Git repository with staged or committed changes
- Read access to codebase

**Optional**:
- Project-specific style guide in `.github/STYLE_GUIDE.md`

## Instructions

### Step 1: Get the Diff

```bash
# For staged changes
git diff --staged

# For specific PR/branch
git diff main...feature-branch
```

### Step 2: Analyze Each Changed File

For each modified file:
1. Read the full file for context
2. Identify the nature of changes (new feature, bug fix, refactor)
3. Check for issues in each category

### Step 3: Generate Structured Review

Produce review in this format:

```markdown
## PR Review: [Brief Title]

### Summary
[1-2 sentence overview of changes]

### Findings

#### Critical Issues
- [ ] [Issue description with file:line reference]

#### Suggestions
- [ ] [Improvement suggestion]

#### Questions
- [ ] [Clarification needed]

### Recommendation
[APPROVE / REQUEST_CHANGES / COMMENT]
```

## Output

- Structured markdown review
- Categorized findings (Critical, Suggestions, Questions)
- Clear recommendation

## Error Handling

1. **Error**: No diff available
   **Solution**: Ensure changes are staged or specify correct branch comparison

2. **Error**: Binary files in diff
   **Solution**: Skip binary files, note in review that manual check needed

3. **Error**: Large diff (>1000 lines)
   **Solution**: Focus on most critical files, suggest splitting PR

4. **Error**: Missing context (file deleted)
   **Solution**: Note inability to review deleted code, focus on remaining changes

## Examples

### Example 1: Simple Bug Fix Review

**User Request**: "Review my bug fix PR"

**Output**:
```markdown
## PR Review: Fix null pointer in user validation

### Summary
Adds null check before accessing user.email property.

### Findings

#### Critical Issues
- None

#### Suggestions
- [ ] Consider adding unit test for null user case (src/validation.ts:45)

#### Questions
- [ ] Should we also check for empty string?

### Recommendation
APPROVE - Good defensive fix, minor test suggestion
```

### Example 2: Security-Sensitive Change

**User Request**: "Check this authentication change"

**Output**:
```markdown
## PR Review: Update JWT token validation

### Summary
Modifies token expiration handling and adds refresh token support.

### Findings

#### Critical Issues
- [ ] Token secret exposed in error message (src/auth.ts:78)
- [ ] Missing rate limiting on refresh endpoint (src/routes/auth.ts:112)

#### Suggestions
- [ ] Add token rotation on refresh
- [ ] Log authentication failures for monitoring

### Recommendation
REQUEST_CHANGES - Security issues must be addressed
```
```

---

## 12. Author Checklist

Run through this checklist every time you create or update a skill:

### Before Writing

- [ ] Identified clear, single-purpose capability
- [ ] Confirmed no existing skill handles this
- [ ] Gathered all necessary reference materials

### Frontmatter

- [ ] `name`: lowercase, hyphens, under 64 chars, matches folder
- [ ] `description`: third person, under 1024 chars, includes what + when + triggers
- [ ] `allowed-tools`: minimal necessary tools only
- [ ] `version`: semver format

### Content

- [ ] Body under 500 lines
- [ ] All required sections present
- [ ] Imperative voice throughout instructions
- [ ] `{baseDir}` used for all paths
- [ ] 2-3 concrete examples with input/output
- [ ] 4+ errors documented with solutions
- [ ] One-level-deep references only

### Testing

- [ ] Triggers correctly on intended phrases
- [ ] Does NOT trigger on unrelated requests
- [ ] Scripts execute successfully
- [ ] Tested with multiple models (Haiku, Sonnet, Opus)
- [ ] Team review completed (if applicable)

### Security

- [ ] No secrets or credentials in skill
- [ ] Tools appropriately scoped
- [ ] Dangerous operations require explicit invocation
- [ ] External dependencies audited

---

## 13. Open Questions / Potentially Out-of-Date Areas

### Confirmed Speculative or Unclear

1. **`when_to_use` field**: Exists in codebase but undocumented. Behavior may ch

…(truncated)
