Rate Limiting Design
Phase 1: Requirements Gathering
Define the goals and constraints for rate limiting.
- Identify services or endpoints requiring rate limiting
- Determine the primary goal: abuse prevention / fair usage / cost control / stability
- List client types and their expected usage patterns:
| Client Type | Expected RPS | Burst Tolerance | SLA Tier |
|---|---|---|---|
- Define acceptable latency overhead for rate limit checks: ___ms
- Identify regulatory or contractual rate limit requirements
Phase 2: Algorithm Selection
Evaluate and select the rate limiting algorithm.
| Algorithm | Pros | Cons | Fit |
|---|---|---|---|
| Fixed Window | Simple, low memory | Burst at window edges | |
| Sliding Window Log | Accurate | High memory for high-volume | |
| Sliding Window Counter | Good accuracy, low memory | Slight approximation | |
| Token Bucket | Allows controlled bursts | Slightly complex | |
| Leaky Bucket | Smooth output rate | No burst tolerance |
- Selected algorithm: ___
- Justification: ___
Phase 3: Limit Configuration
Define rate limits per client tier and endpoint.
| Endpoint / Group | Tier | Requests per Window | Window Size | Burst Limit |
|---|---|---|---|---|
Graduated Limits:
- Unauthenticated requests: ___ req/min
- Authenticated (free tier): ___ req/min
- Authenticated (paid tier): ___ req/min
- Internal services: ___ req/min (or exempt)
Phase 4: Implementation Design
- Enforcement point: API gateway / middleware / application layer / sidecar
- State storage: in-memory / Redis / distributed cache
- Key strategy: API key / IP address / user ID / composite
- Distributed coordination approach (if multi-instance): ___
- Response headers to include:
-
X-RateLimit-Limit(max requests) -
X-RateLimit-Remaining(requests left) -
X-RateLimit-Reset(reset timestamp) -
Retry-After(on 429 responses)
-
- HTTP 429 response body format defined
- Graceful degradation if rate limit store is unavailable
Phase 5: Monitoring and Alerting
- Dashboard showing current rate limit utilization per client/tier
- Alert when clients consistently hit rate limits (potential legitimate need)
- Alert when rate limit infrastructure latency exceeds threshold
- Log all rate-limited requests with client identifier and endpoint
- Track rate limit bypass attempts or anomalous patterns
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |
Output Format
Summary
- Service: ___
- Algorithm: ___
- Enforcement point: ___
- Client tiers: ___
- Endpoints covered: ___
Action Items
- Implement rate limiting at selected enforcement point
- Configure limits per tier and endpoint
- Add standard rate limit response headers
- Set up monitoring dashboard and alerts
- Document rate limits in API documentation
- Communicate limits to existing clients with migration timeline