# Rate Limiting Design

> Use when performing rate limiting design — guides teams through designing and implementing rate limiting strategies for APIs and services. This template covers algorithm selection, limit configuration, client communication, and monitoring to protect backend systems from abuse and overload while maintaining a good developer experience.

- Skill: `cloudthinker-ai/rate-limiting-design` (Agent Skill)
- Install (CLI): `npx skillmds@latest add cloudthinker-ai/rate-limiting-design`
- Raw SKILL.md: https://api.skillmd.com/api/skills/cloudthinker-ai/rate-limiting-design/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: cloudthinker-ai (https://skillmd.com/u/cloudthinker-ai)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/cloudthinker-ai/rate-limiting-design

---


# Rate Limiting Design

## Phase 1: Requirements Gathering

Define the goals and constraints for rate limiting.

- [ ] Identify services or endpoints requiring rate limiting
- [ ] Determine the primary goal: abuse prevention / fair usage / cost control / stability
- [ ] List client types and their expected usage patterns:

| Client Type | Expected RPS | Burst Tolerance | SLA Tier |
|-------------|-------------|-----------------|----------|
|             |             |                 |          |

- [ ] Define acceptable latency overhead for rate limit checks: ___ms
- [ ] Identify regulatory or contractual rate limit requirements

## Phase 2: Algorithm Selection

Evaluate and select the rate limiting algorithm.

| Algorithm | Pros | Cons | Fit |
|-----------|------|------|-----|
| Fixed Window | Simple, low memory | Burst at window edges | |
| Sliding Window Log | Accurate | High memory for high-volume | |
| Sliding Window Counter | Good accuracy, low memory | Slight approximation | |
| Token Bucket | Allows controlled bursts | Slightly complex | |
| Leaky Bucket | Smooth output rate | No burst tolerance | |

- [ ] Selected algorithm: ___
- [ ] Justification: ___

## Phase 3: Limit Configuration

Define rate limits per client tier and endpoint.

| Endpoint / Group | Tier | Requests per Window | Window Size | Burst Limit |
|-----------------|------|--------------------:|-------------|------------:|
|                 |      |                     |             |             |

**Graduated Limits:**

- [ ] Unauthenticated requests: ___ req/min
- [ ] Authenticated (free tier): ___ req/min
- [ ] Authenticated (paid tier): ___ req/min
- [ ] Internal services: ___ req/min (or exempt)

## Phase 4: Implementation Design

- [ ] Enforcement point: API gateway / middleware / application layer / sidecar
- [ ] State storage: in-memory / Redis / distributed cache
- [ ] Key strategy: API key / IP address / user ID / composite
- [ ] Distributed coordination approach (if multi-instance): ___
- [ ] Response headers to include:
  - [ ] `X-RateLimit-Limit` (max requests)
  - [ ] `X-RateLimit-Remaining` (requests left)
  - [ ] `X-RateLimit-Reset` (reset timestamp)
  - [ ] `Retry-After` (on 429 responses)
- [ ] HTTP 429 response body format defined
- [ ] Graceful degradation if rate limit store is unavailable

## Phase 5: Monitoring and Alerting

- [ ] Dashboard showing current rate limit utilization per client/tier
- [ ] Alert when clients consistently hit rate limits (potential legitimate need)
- [ ] Alert when rate limit infrastructure latency exceeds threshold
- [ ] Log all rate-limited requests with client identifier and endpoint
- [ ] Track rate limit bypass attempts or anomalous patterns

## Counter-Rationalizations

| Shortcut | Counter | Why |
|----------|---------|-----|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |

## Output Format

### Summary

- **Service:** ___
- **Algorithm:** ___
- **Enforcement point:** ___
- **Client tiers:** ___
- **Endpoints covered:** ___

### Action Items

- [ ] Implement rate limiting at selected enforcement point
- [ ] Configure limits per tier and endpoint
- [ ] Add standard rate limit response headers
- [ ] Set up monitoring dashboard and alerts
- [ ] Document rate limits in API documentation
- [ ] Communicate limits to existing clients with migration timeline

