CoreWeave Performance Tuning
Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Overview
Tune GPU inference or training only against measured throughput, latency, quality,
availability, and cost targets. A higher utilization figure is not a success if it
causes queueing, memory pressure, or a customer-facing SLO regression.
Prerequisites
- A baseline for p95/p99 latency, throughput, error rate, GPU memory, and utilization.
- A representative non-sensitive evaluation set and a named owner for the SLO.
- A staging lane and a rollback manifest for every resource or serving change.
Instructions
- Change one variable at a time—batching, GPU class, replicas, or memory target.
- Run the agreed load and quality evaluation in staging, then compare with baseline.
- Promote a canary only when all SLO and quality thresholds pass for the observation window.
- Revert to the prior manifest when latency, errors, or quality crosses the agreed limit.
GPU Selection by Workload
| Workload |
Recommended GPU |
Why |
| LLM inference (7-13B) |
A100 80GB |
Good balance of memory and cost |
| LLM inference (70B+) |
8xH100 |
NVLink for tensor parallelism |
| Image generation |
L40 |
Good for diffusion models |
| Training (large models) |
8xH100 SXM5 |
Fastest interconnect |
| Batch processing |
A100 40GB |
Cost-effective |
Inference Optimization
# Continuous batching with vLLM
containers:
- name: vllm
args:
- "--model=meta-llama/Llama-3.1-8B-Instruct"
- "--max-num-batched-tokens=8192"
- "--max-num-seqs=256"
- "--gpu-memory-utilization=0.90"
- "--enable-prefix-caching"
- "--dtype=float16"
Autoscaling Tuning
# HPA based on GPU utilization
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: inference-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: inference-server
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: DCGM_FI_DEV_GPU_UTIL
target:
type: AverageValue
averageValue: "70"
Performance Benchmarks
| Metric |
A100-80GB |
H100-80GB |
| Llama-8B tokens/sec |
~2,000 |
~4,500 |
| Llama-70B tokens/sec |
~200 (4x) |
~500 (4x) |
| Cold start (vLLM) |
30-60s |
20-40s |
Output
- A measured performance baseline and a single reviewed tuning recommendation.
- A canary result covering throughput, latency, error rate, GPU memory, and quality.
- A versioned rollback manifest with a named decision owner.
Error Handling
| Condition |
Safe response |
| GPU memory exceeds the guardrail |
Restore the previous batch or memory setting and investigate the request distribution. |
| Latency rises after batching |
Reduce concurrency or restore replica count; do not raise timeouts to hide the regression. |
| Evaluation quality drops |
Route the canary back to the baseline configuration and preserve aggregate results. |
| Autoscaler oscillates |
Restore stable bounds and tune from a longer measured window. |
Examples
Run a staging canary and save only aggregate measurements for review:
kubectl -n inference-staging apply -f inference-tuned.yaml
kubectl -n inference-staging rollout status deployment/inference-server --timeout=10m
./scripts/load-test --target staging --duration 15m --report aggregate.json
If the report breaches the signed SLO or quality threshold, apply the previous
manifest immediately and attach aggregate.json to the change record.
Resources
Next Steps
For cost optimization, see coreweave-cost-tuning.
1---2name: coreweave-performance-tuning3description: Optimize CoreWeave GPU inference latency and throughput. Use when reducing inference latency, maximizing GPU utilization, or tuning batch sizes and concurrency. Trigger with phrases like "coreweave performance", "coreweave latency", "coreweave throughput", "optimize coreweave inference".4license: MIT5---6# CoreWeave Performance Tuning
7
8> **Community-contributed.** Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
9
10## Overview
11
12Tune GPU inference or training only against measured throughput, latency, quality,
13availability, and cost targets. A higher utilization figure is not a success if it
14causes queueing, memory pressure, or a customer-facing SLO regression.
15
16## Prerequisites
17
18- A baseline for p95/p99 latency, throughput, error rate, GPU memory, and utilization.
19- A representative non-sensitive evaluation set and a named owner for the SLO.
20- A staging lane and a rollback manifest for every resource or serving change.
21
22## Instructions
23
241. Change one variable at a time—batching, GPU class, replicas, or memory target.
252. Run the agreed load and quality evaluation in staging, then compare with baseline.
263. Promote a canary only when all SLO and quality thresholds pass for the observation window.
274. Revert to the prior manifest when latency, errors, or quality crosses the agreed limit.
28
29## GPU Selection by Workload
30
31| Workload | Recommended GPU | Why |
32|----------|----------------|-----|
33| LLM inference (7-13B) | A100 80GB | Good balance of memory and cost |
34| LLM inference (70B+) | 8xH100 | NVLink for tensor parallelism |
35| Image generation | L40 | Good for diffusion models |
36| Training (large models) | 8xH100 SXM5 | Fastest interconnect |
37| Batch processing | A100 40GB | Cost-effective |
38
39## Inference Optimization
40
41```yaml
42# Continuous batching with vLLM
43containers:
44 - name: vllm
45 args:
46 - "--model=meta-llama/Llama-3.1-8B-Instruct"
47 - "--max-num-batched-tokens=8192"
48 - "--max-num-seqs=256"
49 - "--gpu-memory-utilization=0.90"
50 - "--enable-prefix-caching"
51 - "--dtype=float16"
52```
53
54## Autoscaling Tuning
55
56```yaml
57# HPA based on GPU utilization
58apiVersion: autoscaling/v2
59kind: HorizontalPodAutoscaler
60metadata:
61 name: inference-hpa
62spec:
63 scaleTargetRef:
64 apiVersion: apps/v1
65 kind: Deployment
66 name: inference-server
67 minReplicas: 2
68 maxReplicas: 10
69 metrics:
70 - type: Pods
71 pods:
72 metric:
73 name: DCGM_FI_DEV_GPU_UTIL
74 target:
75 type: AverageValue
76 averageValue: "70"
77```
78
79## Performance Benchmarks
80
81| Metric | A100-80GB | H100-80GB |
82|--------|-----------|-----------|
83| Llama-8B tokens/sec | ~2,000 | ~4,500 |
84| Llama-70B tokens/sec | ~200 (4x) | ~500 (4x) |
85| Cold start (vLLM) | 30-60s | 20-40s |
86
87## Output
88
89- A measured performance baseline and a single reviewed tuning recommendation.
90- A canary result covering throughput, latency, error rate, GPU memory, and quality.
91- A versioned rollback manifest with a named decision owner.
92
93## Error Handling
94
95| Condition | Safe response |
96|---|---|
97| GPU memory exceeds the guardrail | Restore the previous batch or memory setting and investigate the request distribution. |
98| Latency rises after batching | Reduce concurrency or restore replica count; do not raise timeouts to hide the regression. |
99| Evaluation quality drops | Route the canary back to the baseline configuration and preserve aggregate results. |
100| Autoscaler oscillates | Restore stable bounds and tune from a longer measured window. |
101
102## Examples
103
104Run a staging canary and save only aggregate measurements for review:
105
106```bash
107kubectl -n inference-staging apply -f inference-tuned.yaml
108kubectl -n inference-staging rollout status deployment/inference-server --timeout=10m
109./scripts/load-test --target staging --duration 15m --report aggregate.json
110```
111
112If the report breaches the signed SLO or quality threshold, apply the previous
113manifest immediately and attach `aggregate.json` to the change record.
114
115## Resources
116
117- [CoreWeave Inference](https://www.coreweave.com/solutions/ai-inference)
118- [vLLM Documentation](https://docs.vllm.ai)
119
120## Next Steps
121
122For cost optimization, see `coreweave-cost-tuning`.