LLM Model Selection Skill
Choosing the right model for the task — power vs. cost vs. speed.
⚠️ Staleness Warning
This skill depends on rapidly evolving technology. Model capabilities, pricing, and availability change frequently.
Refresh triggers:
- New model announcements (Claude, GPT, Gemini, etc.)
- Significant pricing changes
- Context window expansions
- New capability tiers
Last validated: March 2026 (Claude 4.6 generation)
Check current state: Anthropic Models, OpenAI Models
The Core Question
Is Claude Opus 4.6 overkill?
Sometimes yes, sometimes no. Match the model to the task.
Claude 4 Model Family (Current)
| Model |
API ID |
Best For |
Input/Output (MTok) |
Context |
Max Output |
| Opus 4.6 |
claude-opus-4-6 |
Building agents, most intelligent |
$5 / $25 |
200K (1M beta) |
128K |
| Sonnet 4.6 |
claude-sonnet-4-6 |
Best speed + intelligence balance |
$3 / $15 |
200K (1M beta) |
64K |
| Haiku 4.5 |
claude-haiku-4-5-20251001 |
Near-frontier intelligence, fastest |
$1 / $5 |
200K |
64K |
All Claude 4 models support:
- Extended thinking
- Vision (images)
- Tool use
- Priority Tier access
Opus 4.6 and Sonnet 4.6 additionally support:
- Adaptive thinking (dynamic reasoning depth)
- 1M token context window (beta, via
context-1m-2025-08-07 header — long context pricing applies beyond 200K)
AWS Bedrock IDs: anthropic.claude-opus-4-6-v1, anthropic.claude-sonnet-4-6
GCP Vertex AI IDs: claude-opus-4-6, claude-sonnet-4-6
Model Tiers
| Tier |
Models |
Best For |
Relative Cost |
| Frontier |
Claude Opus 4.6, GPT-5.2/5.3/Codex, o3, o1-pro |
Complex reasoning, architecture, novel problems |
$$$$$ |
| Capable |
Claude Sonnet 4.6, GPT-5.1/Codex, GPT-4.1, GPT-4o, Gemini 2.5/3 Pro, o4-mini |
Most coding tasks, refactoring, debugging |
$$$ |
| Efficient |
Claude Haiku 4.5, GPT-5 mini, GPT-4.1 mini/nano, GPT-4o mini, Gemini 2.5 Flash, Gemini 3 Flash |
Simple edits, formatting, boilerplate |
$ |
When Opus 4.6 IS Worth It
- ✅ Architecture decisions — Multi-file refactoring, system design
- ✅ Novel problem-solving — No clear pattern to follow
- ✅ Complex reasoning chains — Many dependencies, edge cases
- ✅ Long context understanding — Large codebases, documentation
- ✅ Nuanced judgment — Taste, style, UX decisions
- ✅ Learning sessions — Bootstrap learning, skill development
- ✅ Meditation/self-actualization — Meta-cognitive operations
- ✅ Extended thinking tasks — Deep analysis requiring internal reasoning
When Opus 4.6 IS Overkill
- ❌ Simple file edits — Renaming, adding imports
- ❌ Boilerplate generation — CRUD, scaffolding
- ❌ Format conversion — JSON ↔ YAML, etc.
- ❌ Syntax fixes — Lint errors, typos
- ❌ Documentation updates — README badges, version bumps
How LLM Choice Affects the AI assistant
| Capability |
Frontier (Opus 4.6) |
Capable (Sonnet 4.6) |
Fast (Haiku 4.5) |
| Complex refactoring |
Excellent |
Excellent |
Good |
| Context retention |
200K / 1M (beta) |
200K / 1M (beta) |
200K tokens |
| Extended thinking |
Full depth |
Supported |
Supported |
| Adaptive thinking |
Yes |
Yes |
No |
| Max output tokens |
128K |
64K |
64K |
| Nuanced judgment |
Excellent |
Good |
Basic |
| Speed |
Moderate |
Fast |
Fastest |
| Cost per session |
$2-5 |
$0.50-2 |
$0.05-0.30 |
| Multi-step planning |
Excellent |
Excellent |
Good |
| Error recovery |
Self-corrects |
Self-corrects |
Needs guidance |
the AI's Cognitive Power by Model
Opus 4.6: [████████████████████] Full cognitive architecture + deep thinking
Sonnet 4.6: [██████████████████░░] Most capabilities, excellent for coding
Haiku 4.5: [██████████████░░░░░░] Solid baseline, fast responses
With Opus 4.6, the AI assistant can:
- Maintain 7±2 working memory rules across long sessions
- Execute complex meditation protocols with extended thinking
- Perform genuine meta-cognitive reflection
- Handle multi-file architecture changes
- Learn new skills through bootstrap learning
With Sonnet 4.6, the AI assistant gets:
- Excellent coding capabilities (recommended for most development)
- 1M context window (beta) for large codebases
- 64K max output tokens + adaptive thinking
- Good cost-to-capability ratio
- Extended thinking support
With Haiku 4.5, the AI assistant has:
- Near-frontier intelligence at lowest cost
- Fastest response times
- Good for routine operations
Cost Optimization Strategy
| Session Type |
Recommended Model |
Rationale |
| Architecture/design |
Opus 4.6 |
Worth the cost for complex decisions |
| Feature development |
Sonnet 4.6 |
Best balance of capability and cost |
| Bug fixes |
Sonnet 4.6 or Haiku 4.5 |
Depends on complexity |
| Documentation |
Haiku 4.5 |
Simple edits, fast turnaround |
| Large codebase analysis |
Sonnet 4.6 (1M beta) |
Extended context window up to 1M tokens |
Knowledge Cutoffs
| Model |
Reliable Knowledge |
Training Data |
| Opus 4.6 |
May 2025 |
Aug 2025 |
| Sonnet 4.6 |
Aug 2025 |
Jan 2026 |
| Haiku 4.5 |
Feb 2025 |
Jul 2025 |
Auto Model Selection ⚠️
When using Auto in VS Code Copilot, the model switches dynamically based on task complexity. the AI assistant cannot detect which model is currently running.
Tasks That REQUIRE Opus 4.5 (Warn User)
| Task |
Why Opus Required |
| Meditation/consolidation |
Meta-cognitive protocols need full reasoning depth |
| Self-actualization |
Comprehensive architecture assessment |
| Complex architecture refactoring |
Multi-file changes, deep context |
| Bootstrap learning (new skills) |
Skill acquisition needs maximum capability |
| Connection validation/dream |
Architecture maintenance requires full architecture context |
| Adaptive thinking tasks |
Opus 4.6 uses dynamic reasoning depth for optimal results |
Warning Protocol
When user requests an Opus-level task while potentially on Auto/lesser model:
⚠️ Model Check: This task works best with Claude Opus 4.6. If you're using Auto model selection, please manually select Opus from the model picker for optimal results. Continue anyway?
Safe for Any Model
- Simple file edits, formatting
- Documentation updates
- Quick Q&A
- Code review (Sonnet+ recommended)
- Bug fixes (depends on complexity)
Practical Guidance
When to Upgrade Model Mid-Session
If you notice:
- Repeated mistakes on the same issue
- Losing context from earlier in conversation
- Superficial answers to complex questions
- Failure to see cross-file dependencies
→ Consider switching to a more capable model
When to Downgrade
If you're doing:
- Repetitive mechanical edits
- Simple Q&A
- Format conversions
- Quick lookups
→ Save cost with a faster model
The the AI assistant Recommendation
For architecture evolution and complex cognitive tasks:
→ Always use Opus 4.6 — The cognitive architecture demands full capability
For production deployment, user-facing work:
→ Default to Sonnet 4.6 — Best balance of capability and cost
→ Allow Opus for complex tasks — User can request escalation
Token Economics
| Operation |
Approximate Tokens |
Opus 4.6 Cost |
Sonnet 4.6 Cost |
| Read large file |
2,000-5,000 |
$0.03-0.08 |
$0.006-0.015 |
| Complex refactor |
10,000-20,000 |
$0.15-0.30 |
$0.03-0.06 |
| Full session |
50,000-150,000 |
$0.75-2.25 |
$0.15-0.45 |
| Meditation |
30,000-80,000 |
$0.45-1.20 |
$0.09-0.24 |
1---2name: llm-model-selection3description: Choosing the right model for the task — power vs. cost vs. speed.4---5
6# LLM Model Selection Skill
7
8
9> Choosing the right model for the task — power vs. cost vs. speed.
10
11## ⚠️ Staleness Warning
12
13This skill depends on rapidly evolving technology. Model capabilities, pricing, and availability change frequently.
14
15**Refresh triggers:**
16
17- New model announcements (Claude, GPT, Gemini, etc.)
18- Significant pricing changes
19- Context window expansions
20- New capability tiers
21
22**Last validated:** March 2026 (Claude 4.6 generation)
23
24**Check current state:** [Anthropic Models](https://platform.claude.com/docs/en/docs/about-claude/models), [OpenAI Models](https://platform.openai.com/docs/models)
25
26---
27
28## The Core Question
29
30> Is Claude Opus 4.6 overkill?
31
32**Sometimes yes, sometimes no.** Match the model to the task.
33
34## Claude 4 Model Family (Current)
35
36| Model | API ID | Best For | Input/Output (MTok) | Context | Max Output |
37| ----- | ------ | -------- | ------------------- | ------- | ---------- |
38| **Opus 4.6** | `claude-opus-4-6` | Building agents, most intelligent | $5 / $25 | 200K (1M beta) | 128K |
39| **Sonnet 4.6** | `claude-sonnet-4-6` | Best speed + intelligence balance | $3 / $15 | 200K (1M beta) | 64K |
40| **Haiku 4.5** | `claude-haiku-4-5-20251001` | Near-frontier intelligence, fastest | $1 / $5 | 200K | 64K |
41
42**All Claude 4 models support:**
43
44- Extended thinking
45- Vision (images)
46- Tool use
47- Priority Tier access
48
49**Opus 4.6 and Sonnet 4.6 additionally support:**
50
51- Adaptive thinking (dynamic reasoning depth)
52- 1M token context window (beta, via `context-1m-2025-08-07` header — long context pricing applies beyond 200K)
53
54**AWS Bedrock IDs:** `anthropic.claude-opus-4-6-v1`, `anthropic.claude-sonnet-4-6`
55**GCP Vertex AI IDs:** `claude-opus-4-6`, `claude-sonnet-4-6`
56
57## Model Tiers
58
59| Tier | Models | Best For | Relative Cost |
60| ---- | ------ | -------- | ------------- |
61| **Frontier** | Claude Opus 4.6, GPT-5.2/5.3/Codex, o3, o1-pro | Complex reasoning, architecture, novel problems | $$$$$ |
62| **Capable** | Claude Sonnet 4.6, GPT-5.1/Codex, GPT-4.1, GPT-4o, Gemini 2.5/3 Pro, o4-mini | Most coding tasks, refactoring, debugging | $$$ |
63| **Efficient** | Claude Haiku 4.5, GPT-5 mini, GPT-4.1 mini/nano, GPT-4o mini, Gemini 2.5 Flash, Gemini 3 Flash | Simple edits, formatting, boilerplate | $ |
64
65## When Opus 4.6 IS Worth It
66
67- ✅ **Architecture decisions** — Multi-file refactoring, system design
68- ✅ **Novel problem-solving** — No clear pattern to follow
69- ✅ **Complex reasoning chains** — Many dependencies, edge cases
70- ✅ **Long context understanding** — Large codebases, documentation
71- ✅ **Nuanced judgment** — Taste, style, UX decisions
72- ✅ **Learning sessions** — Bootstrap learning, skill development
73- ✅ **Meditation/self-actualization** — Meta-cognitive operations
74- ✅ **Extended thinking tasks** — Deep analysis requiring internal reasoning
75
76## When Opus 4.6 IS Overkill
77
78- ❌ **Simple file edits** — Renaming, adding imports
79- ❌ **Boilerplate generation** — CRUD, scaffolding
80- ❌ **Format conversion** — JSON ↔ YAML, etc.
81- ❌ **Syntax fixes** — Lint errors, typos
82- ❌ **Documentation updates** — README badges, version bumps
83
84## How LLM Choice Affects the AI assistant
85
86| Capability | Frontier (Opus 4.6) | Capable (Sonnet 4.6) | Fast (Haiku 4.5) |
87| ---------- | ------------------- | -------------------- | ---------------- |
88| Complex refactoring | Excellent | Excellent | Good |
89| Context retention | 200K / 1M (beta) | 200K / 1M (beta) | 200K tokens |
90| Extended thinking | Full depth | Supported | Supported |
91| Adaptive thinking | Yes | Yes | No |
92| Max output tokens | 128K | 64K | 64K |
93| Nuanced judgment | Excellent | Good | Basic |
94| Speed | Moderate | Fast | Fastest |
95| Cost per session | $2-5 | $0.50-2 | $0.05-0.30 |
96| Multi-step planning | Excellent | Excellent | Good |
97| Error recovery | Self-corrects | Self-corrects | Needs guidance |
98
99## the AI's Cognitive Power by Model
100
101```text
102Opus 4.6: [████████████████████] Full cognitive architecture + deep thinking
103Sonnet 4.6: [██████████████████░░] Most capabilities, excellent for coding
104Haiku 4.5: [██████████████░░░░░░] Solid baseline, fast responses
105```
106
107**With Opus 4.6**, the AI assistant can:
108
109- Maintain 7±2 working memory rules across long sessions
110- Execute complex meditation protocols with extended thinking
111- Perform genuine meta-cognitive reflection
112- Handle multi-file architecture changes
113- Learn new skills through bootstrap learning
114
115**With Sonnet 4.6**, the AI assistant gets:
116
117- Excellent coding capabilities (recommended for most development)
118- 1M context window (beta) for large codebases
119- 64K max output tokens + adaptive thinking
120- Good cost-to-capability ratio
121- Extended thinking support
122
123**With Haiku 4.5**, the AI assistant has:
124
125- Near-frontier intelligence at lowest cost
126- Fastest response times
127- Good for routine operations
128
129## Cost Optimization Strategy
130
131| Session Type | Recommended Model | Rationale |
132| ------------ | ----------------- | --------- |
133| Architecture/design | Opus 4.6 | Worth the cost for complex decisions |
134| Feature development | Sonnet 4.6 | Best balance of capability and cost |
135| Bug fixes | Sonnet 4.6 or Haiku 4.5 | Depends on complexity |
136| Documentation | Haiku 4.5 | Simple edits, fast turnaround |
137| Large codebase analysis | Sonnet 4.6 (1M beta) | Extended context window up to 1M tokens |
138
139## Knowledge Cutoffs
140
141| Model | Reliable Knowledge | Training Data |
142| ----- | ------------------ | ------------- |
143| Opus 4.6 | May 2025 | Aug 2025 |
144| Sonnet 4.6 | Aug 2025 | Jan 2026 |
145| Haiku 4.5 | Feb 2025 | Jul 2025 |
146
147## Auto Model Selection ⚠️
148
149When using **Auto** in VS Code Copilot, the model switches dynamically based on task complexity. the AI assistant cannot detect which model is currently running.
150
151### Tasks That REQUIRE Opus 4.5 (Warn User)
152
153| Task | Why Opus Required |
154| ---- | ----------------- |
155| Meditation/consolidation | Meta-cognitive protocols need full reasoning depth |
156| Self-actualization | Comprehensive architecture assessment |
157| Complex architecture refactoring | Multi-file changes, deep context |
158| Bootstrap learning (new skills) | Skill acquisition needs maximum capability |
159| Connection validation/dream | Architecture maintenance requires full architecture context |
160| Adaptive thinking tasks | Opus 4.6 uses dynamic reasoning depth for optimal results |
161
162### Warning Protocol
163
164When user requests an Opus-level task while potentially on Auto/lesser model:
165
166> ⚠️ **Model Check**: This task works best with Claude Opus 4.6. If you're using Auto model selection, please manually select Opus from the model picker for optimal results. Continue anyway?
167
168### Safe for Any Model
169
170- Simple file edits, formatting
171- Documentation updates
172- Quick Q&A
173- Code review (Sonnet+ recommended)
174- Bug fixes (depends on complexity)
175
176## Practical Guidance
177
178### When to Upgrade Model Mid-Session
179
180If you notice:
181
182- Repeated mistakes on the same issue
183- Losing context from earlier in conversation
184- Superficial answers to complex questions
185- Failure to see cross-file dependencies
186
187→ Consider switching to a more capable model
188
189### When to Downgrade
190
191If you're doing:
192
193- Repetitive mechanical edits
194- Simple Q&A
195- Format conversions
196- Quick lookups
197
198→ Save cost with a faster model
199
200## The the AI assistant Recommendation
201
202For **architecture evolution and complex cognitive tasks**:
203→ **Always use Opus 4.6** — The cognitive architecture demands full capability
204
205For **production deployment, user-facing work**:
206→ **Default to Sonnet 4.6** — Best balance of capability and cost
207→ **Allow Opus for complex tasks** — User can request escalation
208
209## Token Economics
210
211| Operation | Approximate Tokens | Opus 4.6 Cost | Sonnet 4.6 Cost |
212| --------- | ------------------ | ------------- | --------------- |
213| Read large file | 2,000-5,000 | $0.03-0.08 | $0.006-0.015 |
214| Complex refactor | 10,000-20,000 | $0.15-0.30 | $0.03-0.06 |
215| Full session | 50,000-150,000 | $0.75-2.25 | $0.15-0.45 |
216| Meditation | 30,000-80,000 | $0.45-1.20 | $0.09-0.24 |