Anthropic Performance Tuning
Overview
Optimize Claude API latency and throughput via prompt caching, model selection, streaming, and request optimization. The biggest wins come from prompt caching (90% input cost reduction) and model selection (Haiku is 4x faster than Sonnet).
Prompt Caching (Biggest Win)
import anthropic
client = anthropic.Anthropic()
# Mark long, reusable content with cache_control
# Cached content: 90% cheaper on subsequent requests, near-zero latency for cached portion
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are an expert on the following 50-page document: ...<long document>...",
"cache_control": {"type": "ephemeral"} # Cache this block
}
],
messages=[{"role": "user", "content": "What does section 3.2 say?"}]
)
# Check cache performance
print(f"Cache read tokens: {message.usage.cache_read_input_tokens}") # Free/cheap
print(f"Cache creation tokens: {message.usage.cache_creation_input_tokens}") # First call only
print(f"Uncached input tokens: {message.usage.input_tokens}")
Cache requirements: Minimum 1,024 tokens for Sonnet/Opus, 2,048 for Haiku. Cache lives for 5 minutes (refreshed on each hit).
Model Selection for Speed
| Model |
Speed |
Cost (per MTok in/out) |
Best For |
| Claude Haiku |
Fastest |
$0.80 / $4.00 |
Classification, extraction, routing |
| Claude Sonnet |
Balanced |
$3.00 / $15.00 |
General tasks, tool use, code |
| Claude Opus |
Deepest |
$15.00 / $75.00 |
Complex reasoning, research |
# Route by task complexity
def select_model(task_type: str) -> str:
routing = {
"classify": "claude-haiku-4-20250514",
"extract": "claude-haiku-4-20250514",
"summarize": "claude-sonnet-4-20250514",
"code": "claude-sonnet-4-20250514",
"research": "claude-opus-4-20250514",
}
return routing.get(task_type, "claude-sonnet-4-20250514")
Streaming for Perceived Speed
# Streaming reduces time-to-first-token from seconds to ~200ms
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=2048,
messages=[{"role": "user", "content": prompt}]
) as stream:
for text in stream.text_stream:
yield text # User sees response immediately
Reduce Token Count
# 1. Set max_tokens to what you actually need (not max)
msg = client.messages.create(
model="claude-haiku-4-20250514",
max_tokens=128, # Not 4096 — smaller = faster generation
messages=[{"role": "user", "content": "Classify as positive/negative: 'Great product!'"}]
)
# 2. Use prefill to skip preamble
msg = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=64,
messages=[
{"role": "user", "content": "Classify sentiment: 'Great product!'"},
{"role": "assistant", "content": "Sentiment:"} # Skip "Sure, I'd be happy to..."
]
)
# 3. Pre-check token count for large inputs
count = client.messages.count_tokens(
model="claude-sonnet-4-20250514",
messages=[{"role": "user", "content": large_document}]
)
if count.input_tokens > 100_000:
# Chunk or summarize first
pass
Parallel Requests
import Anthropic from '@anthropic-ai/sdk';
import PQueue from 'p-queue';
const client = new Anthropic();
const queue = new PQueue({ concurrency: 10 });
// Process multiple prompts in parallel (within rate limits)
const results = await Promise.all(
prompts.map(p => queue.add(() =>
client.messages.create({
model: 'claude-haiku-4-20250514',
max_tokens: 256,
messages: [{ role: 'user', content: p }],
})
))
);
Performance Benchmarks
| Optimization |
Latency Impact |
Cost Impact |
| Prompt caching |
-50% (cached portion) |
-90% input cost |
| Haiku over Sonnet |
-75% TTFT |
-73% cost |
| Streaming |
-80% TTFT (perceived) |
Same cost |
| Lower max_tokens |
-10-30% total time |
Same cost |
| Prefill technique |
-20% output tokens |
Proportional savings |
Prerequisites
- Define latency, throughput, quality, token, and error SLOs plus the owner-approved model, cache, concurrency, and retry policy.
- Use pinned model IDs, synthetic prompts, an isolated workspace, and representative non-sensitive fixtures; do not benchmark with customer content or production credentials.
- Configure aggregate-only telemetry, bounded concurrency, rate-limit awareness, and a tested rollback configuration.
Instructions
- Establish a baseline for time-to-first-token, completion latency, tokens, cache hit rate, throughput, quality, and errors using repeated synthetic runs.
- Change one lever at a time: model, prompt/cache layout, token budget, streaming, batching, or concurrency. Keep prompt content out of logs and verify cache eligibility for sensitive data before enabling it.
- Enforce request scope,
max_tokens, timeout, retry, and concurrency limits. Stop the run when rate limits, quality, or data-policy checks fail rather than increasing access or disabling controls.
- Canary the selected configuration in a sandbox or internal workspace, compare against baseline, and obtain approval before production rollout. Monitor p95/p99 latency, error rate, token use, and spend.
- Restore the prior configuration on regression, invalidate temporary cache/test artifacts according to retention policy, and retain a redacted benchmark receipt.
Output
Produce a performance receipt containing configuration and model IDs, benchmark fixture class, sample size, latency/throughput/token/cache aggregates, quality and error outcomes, workspace/canary scope, approval, retention, and rollback reference. Exclude prompts, responses, user identifiers, and secrets.
Error Handling
| Failure |
Response |
| Rate limit or queue saturation |
Reduce bounded concurrency, honor retry guidance, and stop the canary if the SLO remains breached. |
| Quality falls after model/token change |
Restore the baseline configuration and quarantine the comparison until reviewed. |
| Cache miss or policy-ineligible content |
Disable caching for that path and use the approved uncached flow. |
| Timeout or streaming disconnect |
Apply bounded retry/idempotency handling, return a safe partial-state result, and investigate without logging content. |
Examples
Benchmark 500 synthetic fixture-prompt-* requests in a staging workspace with Haiku and the current route, assert content_logged=0, p99_latency<approved_limit, and quality=pass, then canary the winner to internal traffic. If p99 or error thresholds fail, emit canary=halted; rollback=perf-baseline.
Resources
Next Steps
For cost optimization, see anth-cost-tuning.
1---2name: anth-performance-tuning3description: Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques. Use when experiencing slow responses, optimizing token usage, or reducing time-to-first-token in production. Trigger with phrases like "anthropic performance", "claude speed", "optimize claude latency", "anthropic caching", "faster claude responses".4license: MIT5---6# Anthropic Performance Tuning
7
8## Overview
9
10Optimize Claude API latency and throughput via prompt caching, model selection, streaming, and request optimization. The biggest wins come from prompt caching (90% input cost reduction) and model selection (Haiku is 4x faster than Sonnet).
11
12## Prompt Caching (Biggest Win)
13
14```python
15import anthropic
16
17client = anthropic.Anthropic()
18
19# Mark long, reusable content with cache_control
20# Cached content: 90% cheaper on subsequent requests, near-zero latency for cached portion
21message = client.messages.create(
22 model="claude-sonnet-4-20250514",
23 max_tokens=1024,
24 system=[
25 {
26 "type": "text",
27 "text": "You are an expert on the following 50-page document: ...<long document>...",
28 "cache_control": {"type": "ephemeral"} # Cache this block
29 }
30 ],
31 messages=[{"role": "user", "content": "What does section 3.2 say?"}]
32)
33
34# Check cache performance
35print(f"Cache read tokens: {message.usage.cache_read_input_tokens}") # Free/cheap
36print(f"Cache creation tokens: {message.usage.cache_creation_input_tokens}") # First call only
37print(f"Uncached input tokens: {message.usage.input_tokens}")
38```
39
40**Cache requirements:** Minimum 1,024 tokens for Sonnet/Opus, 2,048 for Haiku. Cache lives for 5 minutes (refreshed on each hit).
41
42## Model Selection for Speed
43
44| Model | Speed | Cost (per MTok in/out) | Best For |
45|-------|-------|----------------------|----------|
46| Claude Haiku | Fastest | $0.80 / $4.00 | Classification, extraction, routing |
47| Claude Sonnet | Balanced | $3.00 / $15.00 | General tasks, tool use, code |
48| Claude Opus | Deepest | $15.00 / $75.00 | Complex reasoning, research |
49
50```python
51# Route by task complexity
52def select_model(task_type: str) -> str:
53 routing = {
54 "classify": "claude-haiku-4-20250514",
55 "extract": "claude-haiku-4-20250514",
56 "summarize": "claude-sonnet-4-20250514",
57 "code": "claude-sonnet-4-20250514",
58 "research": "claude-opus-4-20250514",
59 }
60 return routing.get(task_type, "claude-sonnet-4-20250514")
61```
62
63## Streaming for Perceived Speed
64
65```python
66# Streaming reduces time-to-first-token from seconds to ~200ms
67with client.messages.stream(
68 model="claude-sonnet-4-20250514",
69 max_tokens=2048,
70 messages=[{"role": "user", "content": prompt}]
71) as stream:
72 for text in stream.text_stream:
73 yield text # User sees response immediately
74```
75
76## Reduce Token Count
77
78```python
79# 1. Set max_tokens to what you actually need (not max)
80msg = client.messages.create(
81 model="claude-haiku-4-20250514",
82 max_tokens=128, # Not 4096 — smaller = faster generation
83 messages=[{"role": "user", "content": "Classify as positive/negative: 'Great product!'"}]
84)
85
86# 2. Use prefill to skip preamble
87msg = client.messages.create(
88 model="claude-sonnet-4-20250514",
89 max_tokens=64,
90 messages=[
91 {"role": "user", "content": "Classify sentiment: 'Great product!'"},
92 {"role": "assistant", "content": "Sentiment:"} # Skip "Sure, I'd be happy to..."
93 ]
94)
95
96# 3. Pre-check token count for large inputs
97count = client.messages.count_tokens(
98 model="claude-sonnet-4-20250514",
99 messages=[{"role": "user", "content": large_document}]
100)
101if count.input_tokens > 100_000:
102 # Chunk or summarize first
103 pass
104```
105
106## Parallel Requests
107
108```typescript
109import Anthropic from '@anthropic-ai/sdk';
110import PQueue from 'p-queue';
111
112const client = new Anthropic();
113const queue = new PQueue({ concurrency: 10 });
114
115// Process multiple prompts in parallel (within rate limits)
116const results = await Promise.all(
117 prompts.map(p => queue.add(() =>
118 client.messages.create({
119 model: 'claude-haiku-4-20250514',
120 max_tokens: 256,
121 messages: [{ role: 'user', content: p }],
122 })
123 ))
124);
125```
126
127## Performance Benchmarks
128
129| Optimization | Latency Impact | Cost Impact |
130|-------------|----------------|-------------|
131| Prompt caching | -50% (cached portion) | -90% input cost |
132| Haiku over Sonnet | -75% TTFT | -73% cost |
133| Streaming | -80% TTFT (perceived) | Same cost |
134| Lower max_tokens | -10-30% total time | Same cost |
135| Prefill technique | -20% output tokens | Proportional savings |
136
137## Prerequisites
138
139- Define latency, throughput, quality, token, and error SLOs plus the owner-approved model, cache, concurrency, and retry policy.
140- Use pinned model IDs, synthetic prompts, an isolated workspace, and representative non-sensitive fixtures; do not benchmark with customer content or production credentials.
141- Configure aggregate-only telemetry, bounded concurrency, rate-limit awareness, and a tested rollback configuration.
142
143## Instructions
144
1451. Establish a baseline for time-to-first-token, completion latency, tokens, cache hit rate, throughput, quality, and errors using repeated synthetic runs.
1462. Change one lever at a time: model, prompt/cache layout, token budget, streaming, batching, or concurrency. Keep prompt content out of logs and verify cache eligibility for sensitive data before enabling it.
1473. Enforce request scope, `max_tokens`, timeout, retry, and concurrency limits. Stop the run when rate limits, quality, or data-policy checks fail rather than increasing access or disabling controls.
1484. Canary the selected configuration in a sandbox or internal workspace, compare against baseline, and obtain approval before production rollout. Monitor p95/p99 latency, error rate, token use, and spend.
1495. Restore the prior configuration on regression, invalidate temporary cache/test artifacts according to retention policy, and retain a redacted benchmark receipt.
150
151## Output
152
153Produce a performance receipt containing configuration and model IDs, benchmark fixture class, sample size, latency/throughput/token/cache aggregates, quality and error outcomes, workspace/canary scope, approval, retention, and rollback reference. Exclude prompts, responses, user identifiers, and secrets.
154
155## Error Handling
156
157| Failure | Response |
158|---|---|
159| Rate limit or queue saturation | Reduce bounded concurrency, honor retry guidance, and stop the canary if the SLO remains breached. |
160| Quality falls after model/token change | Restore the baseline configuration and quarantine the comparison until reviewed. |
161| Cache miss or policy-ineligible content | Disable caching for that path and use the approved uncached flow. |
162| Timeout or streaming disconnect | Apply bounded retry/idempotency handling, return a safe partial-state result, and investigate without logging content. |
163
164## Examples
165
166Benchmark 500 synthetic `fixture-prompt-*` requests in a staging workspace with Haiku and the current route, assert `content_logged=0`, `p99_latency<approved_limit`, and `quality=pass`, then canary the winner to internal traffic. If p99 or error thresholds fail, emit `canary=halted; rollback=perf-baseline`.
167
168## Resources
169
170- [Prompt Caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching)
171- [Token Counting](https://docs.anthropic.com/en/docs/build-with-claude/token-counting)
172- [Pricing](https://docs.anthropic.com/en/docs/about-claude/pricing)
173
174## Next Steps
175
176For cost optimization, see `anth-cost-tuning`.