Backend Skill — Production Backend Engineering (2025–2026)
This skill covers framework-agnostic backend engineering principles: API design, database patterns, caching, security, observability, architecture decisions, and deployment. Combine with framework-specific skills (FastAPI, Express, Django, Spring, etc.) for implementation details.
1. API Design Fundamentals
REST URL conventions
Resources are nouns (plural). HTTP methods are the verbs.
GET /api/v1/users # List users
POST /api/v1/users # Create a user
GET /api/v1/users/123 # Get user 123
PATCH /api/v1/users/123 # Partial update
PUT /api/v1/users/123 # Full replace
DELETE /api/v1/users/123 # Delete user 123
GET /api/v1/users/123/orders # User's orders (shallow nesting)
GET /api/v1/users/123/orders/456 # Specific order
# ❌ Anti-patterns
GET /api/v1/getAllUsers
POST /api/v1/createUser
POST /api/v1/deleteUser
GET /api/v1/user?id=123
URL design rules
- Use plural nouns:
/users, /orders, /products.
- Use lowercase with hyphens:
/order-items, not /orderItems or /order_items.
- Keep nesting shallow (max 2 levels):
/users/123/orders is fine; /users/123/orders/456/items/789/reviews is too deep — flatten to /reviews?item_id=789.
- Use query parameters for filtering, sorting, pagination:
/users?role=admin&sort=-created_at&page=2.
- Avoid exposing internal IDs when possible — use UUIDs or slugs for public-facing resources.
HTTP methods and semantics
| Method |
Idempotent |
Safe |
Use for |
| GET |
Yes |
Yes |
Read resources |
| POST |
No |
No |
Create resources, trigger actions |
| PUT |
Yes |
No |
Full replace of a resource |
| PATCH |
Yes* |
No |
Partial update |
| DELETE |
Yes |
No |
Remove a resource |
| HEAD |
Yes |
Yes |
Like GET but headers only |
| OPTIONS |
Yes |
Yes |
CORS preflight, capabilities |
*PATCH is idempotent when using merge-patch semantics (the common case).
HTTP status codes
Use them correctly — they are part of your API contract.
2xx — Success
200 OK — Standard success response
201 Created — Resource created (include Location header)
204 No Content — Success with no response body (DELETE, PUT)
3xx — Redirection
301 Moved Permanently — Resource URL changed permanently
304 Not Modified — Cached version is still valid
4xx — Client Errors
400 Bad Request — Malformed syntax or invalid input
401 Unauthorized — Missing or invalid authentication
403 Forbidden — Authenticated but insufficient permissions
404 Not Found — Resource doesn't exist
409 Conflict — State conflict (e.g., duplicate email)
422 Unprocessable Entity — Valid syntax but semantic errors
429 Too Many Requests — Rate limit exceeded (include Retry-After header)
5xx — Server Errors
500 Internal Server Error — Unhandled server error
502 Bad Gateway — Upstream service failure
503 Service Unavailable — Temporary overload or maintenance
504 Gateway Timeout — Upstream service timeout
Consistent error response format
Follow RFC 9457 (Problem Details for HTTP APIs):
{
"type": "https://api.example.com/errors/validation",
"title": "Validation Error",
"status": 422,
"detail": "The email field is required and must be a valid email address.",
"instance": "/api/v1/users",
"errors": [
{
"field": "email",
"message": "This field is required.",
"code": "required"
}
]
}
Rules for error responses:
- Always return a consistent structure across all endpoints.
- Include machine-readable error codes alongside human-readable messages.
- Include field-level validation errors so clients can fix all issues in one round-trip.
- Never leak internal details (stack traces, SQL queries, file paths) in production errors.
- Log the full error server-side; return a sanitized version to the client.
2. API Versioning
Strategies
| Strategy |
Example |
Pros |
Cons |
| URL path |
/api/v1/users |
Explicit, easy to route |
URL pollution |
| Header |
Accept: application/vnd.api+json;version=2 |
Clean URLs |
Hidden, harder to test |
| Query param |
/api/users?version=2 |
Simple |
Inconsistent |
Recommendation: Use URL path versioning (/api/v1/). It's the most explicit, easiest to route, cache, and document. When introducing a breaking change, create a new version (/api/v2/) and run both in parallel with a deprecation timeline.
Deprecation protocol
- Announce deprecation with a
Deprecation header and Sunset header on old endpoints.
- Provide a migration guide.
- Give clients at least 6–12 months before removal.
- Monitor usage of deprecated endpoints before turning them off.
3. Pagination
Cursor-based pagination (recommended)
Best for real-time data, large datasets, and infinite scroll. Immune to offset drift when data changes.
GET /api/v1/orders?limit=20&after=eyJpZCI6MTIzfQ
Response:
{
"data": [...],
"pagination": {
"has_next": true,
"next_cursor": "eyJpZCI6MTQzfQ",
"has_previous": true,
"previous_cursor": "eyJpZCI6MTI0fQ"
}
}
Offset-based pagination (simpler)
Fine for admin dashboards, small datasets, or when users need to jump to specific pages.
GET /api/v1/users?page=3&per_page=25
Response:
{
"data": [...],
"pagination": {
"page": 3,
"per_page": 25,
"total": 482,
"total_pages": 20
}
}
Pagination rules
- Always set a maximum page size (e.g.,
per_page capped at 100).
- Default to a reasonable page size (20–50).
- Include pagination metadata in every list response.
- For very large datasets (millions of rows), cursor-based is the only performant option.
4. Filtering, Sorting, and Field Selection
# Filtering
GET /api/v1/products?category=electronics&price_min=100&price_max=500&in_stock=true
# Sorting (prefix with - for descending)
GET /api/v1/products?sort=-price,name
# Field selection (reduce payload size)
GET /api/v1/users?fields=id,name,email
# Full-text search
GET /api/v1/products?q=wireless+headphones
Be consistent across all endpoints. Document every supported filter, sort field, and selection parameter.
5. Authentication & Authorization
Authentication methods by use case
| Use case |
Method |
| User-facing apps |
OAuth 2.0 / OIDC |
| Server-to-server |
API keys + HMAC signing |
| Microservice-to-microservice |
Mutual TLS (mTLS) or JWT |
| Simple internal tools |
API keys in headers |
JWT best practices
Authorization: Bearer eyJhbGciOiJSUzI1NiIs...
- Use short-lived access tokens (5–15 minutes) with long-lived refresh tokens (days/weeks).
- Store JWTs in httpOnly, Secure, SameSite=Strict cookies for browser clients — never in localStorage.
- Use asymmetric signing (RS256/ES256) for distributed systems so services can verify without the private key.
- Include minimal claims:
sub (user ID), exp, iat, iss, roles/scopes.
- Validate
exp, iss, and aud on every request.
- Implement token revocation via a short blocklist or by keeping access tokens very short-lived.
OAuth 2.0 / OIDC
- Use Authorization Code flow with PKCE for all client types (web, mobile, SPA).
- Never use the Implicit flow — it's deprecated.
- Use established providers (Auth0, Clerk, Keycloak, Supabase Auth) rather than rolling your own.
- Always validate the
state parameter to prevent CSRF.
- Store client secrets server-side only.
Authorization patterns
- RBAC (Role-Based Access Control): Assign roles (admin, editor, viewer) with predefined permissions. Works for most applications.
- ABAC (Attribute-Based Access Control): Policies based on user attributes, resource attributes, and context. More flexible but more complex.
- Check authorization at every layer: Middleware for route-level checks, service layer for business-logic checks. Never rely on a single layer.
6. Database Design
Schema design principles
- Normalize first, denormalize intentionally. Start with 3NF; denormalize specific tables for read performance when you have evidence (not speculation) of a bottleneck.
- Use appropriate data types: Don't store dates as strings, don't use TEXT for short fixed-length fields, use DECIMAL/NUMERIC for money (never FLOAT).
- Always have a primary key: Prefer auto-incrementing integers for internal IDs; use UUIDs (v7 for sortability) for public-facing IDs.
- Add created_at and updated_at to every table (with timezone).
- Use foreign keys for referential integrity in OLTP databases.
- Index strategically: Index columns used in WHERE, JOIN, ORDER BY. Composite indexes follow the leftmost prefix rule. Don't over-index — each index slows writes.
Query optimization
-- ✅ Use EXPLAIN ANALYZE to understand query plans
EXPLAIN ANALYZE
SELECT u.id, u.name, COUNT(o.id) as order_count
FROM users u
LEFT JOIN orders o ON o.user_id = u.id
WHERE u.created_at > '2026-01-01'
GROUP BY u.id, u.name
ORDER BY order_count DESC
LIMIT 20;
-- ✅ Avoid SELECT * — fetch only needed columns
SELECT id, name, email FROM users WHERE active = true;
-- ❌ N+1 queries — the most common performance killer
-- Fetching users, then looping to fetch each user's orders separately
-- ✅ Fix: JOIN or batch fetch in a single query
Common query pitfalls
| Pitfall |
Fix |
| N+1 queries |
Use JOINs, eager loading, or batch fetching |
| Missing indexes on foreign keys |
Always index FK columns |
| Full table scans on large tables |
Add appropriate indexes; use LIMIT |
| Using OFFSET for deep pagination |
Switch to cursor/keyset pagination |
| Locking issues on hot tables |
Use row-level locking; consider read replicas |
| Not using connection pooling |
Use PgBouncer, HikariCP, or framework-level pools |
SQL vs NoSQL decision
| Choose SQL (PostgreSQL, MySQL) when |
Choose NoSQL when |
| Complex relationships and JOINs |
Document-shaped data (MongoDB) |
| ACID transactions required |
High write throughput at scale (Cassandra) |
| Structured, well-known schema |
Key-value lookups / caching (Redis) |
| Reporting and aggregation |
Graph relationships (Neo4j) |
| Most CRUD applications |
Time-series data (InfluxDB, TimescaleDB) |
Default to PostgreSQL for most applications. It handles JSON (jsonb), full-text search, geospatial data, and scales to millions of rows before you need anything exotic.
7. Caching
Caching layers
Client ──→ CDN ──→ API Gateway ──→ Application Cache ──→ Database
↑ ↑ ↑
Static Response Redis/Memcached
assets caching or in-memory
Caching strategies
| Strategy |
How it works |
Best for |
| Cache-aside (Lazy loading) |
App checks cache → miss → query DB → write to cache |
General-purpose; most common |
| Write-through |
App writes to cache AND DB simultaneously |
Strong consistency requirements |
| Write-behind (Write-back) |
App writes to cache; cache async writes to DB |
High write throughput |
| Read-through |
Cache itself fetches from DB on miss |
Simpler app code |
Cache invalidation
Cache invalidation is one of the two hard problems in CS. Strategies:
- TTL (Time-to-Live): Set expiration. Simple but can serve stale data.
- Event-driven invalidation: Invalidate cache when the underlying data changes (e.g., after a write).
- Tag-based invalidation: Tag cache entries; invalidate all entries with a given tag.
- Versioned keys: Include a version number in the cache key; bump it on changes.
Redis patterns
# Cache-aside pattern (pseudocode)
def get_user(user_id):
cached = redis.get(f"user:{user_id}")
if cached:
return deserialize(cached)
user = db.query("SELECT * FROM users WHERE id = %s", user_id)
redis.setex(f"user:{user_id}", 300, serialize(user)) # TTL: 5 min
return user
# Distributed locking (prevent thundering herd)
lock = redis.set(f"lock:user:{user_id}", "1", nx=True, ex=10)
if lock:
# Only one process rebuilds the cache
...
What to cache
- Cache: Expensive query results, computed aggregations, API responses from third parties, session data, frequently-read and rarely-changed data.
- Don't cache: User-specific sensitive data without proper scoping, rapidly-changing data where staleness is unacceptable, one-time reads.
8. Rate Limiting
Protect your API from abuse, ensure fair usage, and prevent cascading failures.
Algorithms
| Algorithm |
Behavior |
Best for |
| Token bucket |
Tokens refill at a steady rate; requests consume tokens |
Smooth rate limiting with burst allowance |
| Sliding window |
Counts requests in a rolling time window |
Precise rate limiting |
| Fixed window |
Counts requests per fixed time window |
Simplest to implement |
Implementation
# Response headers for rate limiting
X-RateLimit-Limit: 1000 # Max requests per window
X-RateLimit-Remaining: 847 # Requests remaining
X-RateLimit-Reset: 1712800000 # Unix timestamp when window resets
Retry-After: 30 # Seconds to wait (on 429 response)
- Apply rate limits per API key/user, not just per IP (IPs are shared behind NATs/VPNs).
- Use different limits for different tiers (free: 100/hr, pro: 10,000/hr).
- Return
429 Too Many Requests with a Retry-After header.
- Implement at the API gateway or reverse proxy level for efficiency.
- Use Redis or a distributed counter for multi-instance deployments.
9. Background Jobs & Async Processing
When to use background jobs
- Email/notification sending
- PDF/report generation
- Image/video processing
- Data import/export
- Webhook delivery and retries
- Scheduled tasks (cron-like)
- Any operation that takes > 1–2 seconds
Architecture
API Server ──→ Message Queue ──→ Workers
│ (Redis, RabbitMQ, │
│ Kafka, SQS) │
│ ↓
└── Returns 202 Accepted Process job
with job status URL Update status
Notify on completion
Job design principles
- Idempotent: Jobs should be safe to retry. If a job runs twice, the result should be the same.
- Small and focused: Each job does one thing. Chain jobs for complex workflows.
- Observable: Log job start/end/failure. Track duration, success rate, and queue depth.
- Retry with backoff: Use exponential backoff with jitter (e.g., 1s, 2s, 4s, 8s + random jitter).
- Dead letter queue: Failed jobs after max retries go to a DLQ for manual inspection.
- Timeout: Set maximum execution time to prevent zombie jobs.
Tools
| Language/Ecosystem |
Tool |
| Python |
Celery + Redis/RabbitMQ, ARQ, Dramatiq |
| Node.js |
BullMQ + Redis, Temporal |
| Java/Kotlin |
Spring Batch, Quartz |
| Language-agnostic |
Kafka, RabbitMQ, AWS SQS, Temporal |
10. Security
OWASP Top 10 awareness (2025–2026)
The most critical backend security risks. Every backend engineer should know these:
- Broken Access Control — Enforce authorization at every endpoint and data access layer.
- Cryptographic Failures — Use TLS everywhere. Hash passwords with bcrypt/argon2. Encrypt sensitive data at rest.
- Injection — Use parameterized queries / ORMs. Never concatenate user input into SQL/commands.
- Insecure Design — Threat model during design, not after deployment.
- Security Misconfiguration — Disable debug endpoints in production. Remove default credentials. Set security headers.
- Vulnerable Components — Keep dependencies updated. Audit with
npm audit, pip-audit, Snyk, or Dependabot.
- Authentication Failures — Implement account lockout, rate limiting on login, MFA.
- Data Integrity Failures — Verify software updates and CI/CD pipelines. Use SRI for CDN assets.
- Logging & Monitoring Failures — Log security events. Set up alerts for anomalies.
- SSRF — Validate and sanitize URLs the server fetches. Allowlist permitted domains.
Security headers
Strict-Transport-Security: max-age=63072000; includeSubDomains; preload
Content-Security-Policy: default-src 'self'; script-src 'self'
X-Content-Type-Options: nosniff
X-Frame-Options: DENY
Referrer-Policy: strict-origin-when-cross-origin
Permissions-Policy: camera=(), microphone=(), geolocation=()
Input validation
- Validate all input: request body, query params, path params, headers.
- Validate on the server — client validation is for UX, not security.
- Use schema validation libraries (Pydantic, Zod, Joi, JSON Schema).
- Reject unexpected fields (strict mode).
- Sanitize output when rendering user content to prevent XSS.
Password storage
- Never store passwords in plain text or with reversible encryption.
- Use bcrypt (cost factor ≥ 12) or argon2id (preferred for new systems).
- Salt is built into bcrypt/argon2 — don't roll your own salting scheme.
11. Observability
Observability = Logs + Metrics + Traces. Design it alongside your application, not as an afterthought.
The three pillars
Logs — Structured event records.
{
"timestamp": "2026-04-10T14:32:01Z",
"level": "error",
"service": "order-service",
"request_id": "req_abc123",
"user_id": "usr_456",
"message": "Payment processing failed",
"error": "stripe_timeout",
"duration_ms": 30042
}
- Use structured logging (JSON) — not unstructured text.
- Include request IDs / correlation IDs that propagate across services.
- Log at appropriate levels: DEBUG (development only), INFO (significant events), WARN (recoverable issues), ERROR (failures requiring attention).
- Never log passwords, tokens, PII, or credit card numbers.
Metrics — Aggregated numerical measurements.
Key metrics to track (the RED method):
- Rate: Requests per second
- Errors: Error rate (4xx, 5xx)
- Duration: Response time (p50, p95, p99)
Additional: CPU/memory usage, queue depth, active connections, cache hit ratio, DB query time.
Traces — Request flow across services.
- Use OpenTelemetry (the industry standard) for instrumentation.
- Propagate trace context (
traceparent header) across service boundaries.
- Trace critical paths: API → service → database → external API.
Alerting rules
- Alert on symptoms (high error rate, slow response times), not causes.
- Set thresholds based on SLOs (Service Level Objectives), e.g., "99.9% of requests complete in < 500ms."
- Use severity levels: critical (paging), warning (next business day), info (dashboard only).
- Avoid alert fatigue — every alert should be actionable.
Tools
| Category |
Tools |
| Logging |
ELK Stack, Loki + Grafana, Datadog |
| Metrics |
Prometheus + Grafana, Datadog, CloudWatch |
| Tracing |
Jaeger, Tempo, Datadog APM, Honeycomb |
| All-in-one |
Datadog, New Relic, Grafana Cloud |
| Instrumentation |
OpenTelemetry (universal standard) |
12. Architecture Patterns
Monolith vs Microservices
|
Modular Monolith |
Microservices |
| Start with |
✅ Always start here |
Migrate to when needed |
| Deploy |
Single artifact |
Independent per service |
| Complexity |
Lower operational overhead |
Higher (networking, observability, consistency) |
| Team size |
< 20–30 engineers |
Large orgs with team boundaries |
| Data |
Single database |
Database per service |
Start with a well-structured monolith. Extract services only when you have clear domain boundaries and team/scaling needs that justify the operational complexity.
Layered architecture (for any backend)
┌─────────────────────────────────┐
│ Presentation (Routes/Handlers) │ ← HTTP layer, input validation
├─────────────────────────────────┤
│ Service / Business Logic │ ← Domain rules, orchestration
├─────────────────────────────────┤
│ Repository / Data Access │ ← Database queries, external APIs
├─────────────────────────────────┤
│ Infrastructure │ ← DB connections, caches, queues
└─────────────────────────────────┘
Rules:
- Each layer only calls the layer directly below it.
- Routes should be thin — delegate to services.
- Services contain business logic — no HTTP concepts (no request/response objects).
- Repositories abstract data access — services don't write SQL directly.
Communication patterns
| Pattern |
Protocol |
Use when |
| Synchronous request/response |
HTTP/REST, gRPC |
Direct queries, CRUD |
| Async messaging |
RabbitMQ, Kafka, SQS |
Decoupled processing, event-driven workflows |
| Streaming |
gRPC streams, WebSockets, SSE |
Real-time updates, live feeds |
| Webhooks |
HTTP callbacks |
Notifying external systems of events |
Resilience patterns
- Circuit breaker: Stop calling a failing service; fail fast and recover gracefully.
- Retry with exponential backoff + jitter: Retry transient failures without thundering herd.
- Timeout: Set timeouts on all external calls. A missing timeout is a bug.
- Bulkhead: Isolate failures — don't let one slow dependency bring down the entire system.
- Health checks: Expose
/health (liveness) and /ready (readiness) endpoints.
13. Webhooks
Sending webhooks
- Sign payloads with HMAC-SHA256 so receivers can verify authenticity.
- Use exponential backoff for retries (e.g., 5 attempts: 1m, 5m, 30m, 2h, 24h).
- Consider successful delivery as 2xx response within a timeout (e.g., 30s).
- Log all delivery attempts for debugging.
- Provide a webhook event log / dashboard in your API.
- Send a minimal payload with an event type and resource ID; let receivers fetch details via API.
Receiving webhooks
- Verify the signature before processing.
- Respond with
200 OK quickly; process asynchronously.
- Handle duplicate deliveries gracefully (idempotent processing).
- Use a queue for processing to avoid blocking the webhook endpoint.
14. Testing Backend Systems
Testing pyramid
╱ E2E Tests ╲ ← Few: Full user flows
╱ Integration ╲ ← Some: API endpoints with real DB
╱ Unit Tests ╲ ← Many: Business logic, utilities
╱────────────────────╲
What to test at each level
Unit tests (fast, isolated):
- Business logic functions
- Validation rules
- Data transformations
- Edge cases and error paths
Integration tests (with real dependencies):
- API endpoints end-to-end (HTTP → handler → service → DB → response)
- Database queries against a test database
- Cache behavior
- Authentication/authorization flows
Contract tests (for API consumers):
- Verify your API responses match the documented schema.
- Use tools like Pact or Schemathesis.
Load tests (before production):
- Use k6, Locust, or Gatling.
- Test at 2–3× expected peak traffic.
- Identify bottlenecks: slow queries, memory leaks, connection exhaustion.
Testing rules
- Test behavior, not implementation.
- Use factories (Factory Boy, Faker) to generate test data — not fixtures.
- Isolate tests — each test should set up and tear down its own data.
- Test error paths as thoroughly as happy paths.
- Run tests in CI on every pull request.
15. Deployment & Infrastructure
Container best practices (Docker)
# Multi-stage build for smaller images
FROM python:3.12-slim AS builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
FROM python:3.12-slim
WORKDIR /app
COPY --from=builder /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packages
COPY . .
# Run as non-root user
RUN adduser --disabled-password appuser
USER appuser
EXPOSE 8000
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]
- Use multi-stage builds for smaller images.
- Run as a non-root user.
- Pin exact versions for base images and dependencies.
- Use
.dockerignore to exclude unnecessary files.
- Scan images for vulnerabilities (Trivy, Snyk).
CI/CD pipeline
Push → Lint → Unit Tests → Build Image → Integration Tests → Security Scan → Deploy to Staging → E2E Tests → Deploy to Production
- Every commit triggers lint + unit tests.
- Every PR merge triggers full pipeline including integration tests.
- Use blue-green or canary deployments for zero-downtime releases.
- Rollback must be a single command / automated on error spike.
Health checks
// GET /health (liveness — is the process alive?)
{ "status": "ok" }
// GET /ready (readiness — can the service handle traffic?)
{
"status": "ok",
"checks": {
"database": "connected",
"redis": "connected",
"queue": "connected"
}
}
16. API Documentation
- Use OpenAPI 3.1 (Swagger) for REST APIs.
- Generate docs from code (not maintained separately) — most frameworks support this.
- Include: authentication instructions, endpoint references, request/response examples, error format, rate limits, pagination details.
- Provide runnable examples (curl, HTTPie, or language-specific snippets).
- Keep docs in sync — stale docs are worse than no docs.
- Consider publishing an
llms.txt for AI agent consumption.
17. Common Pitfalls
| Pitfall |
Fix |
| No input validation |
Validate everything server-side with schema libraries |
| Leaking internal errors to clients |
Return generic errors; log details server-side |
| N+1 database queries |
Use JOINs, eager loading, or DataLoader patterns |
| No rate limiting |
Implement at API gateway level; return 429 with Retry-After |
| OFFSET-based pagination on large datasets |
Switch to cursor/keyset pagination |
| Rolling your own auth |
Use established providers or battle-tested libraries |
| No timeouts on external calls |
Set timeouts on every HTTP client, DB connection, and queue consumer |
| Monolithic error handling |
Use structured error types with consistent response format |
| No request/correlation IDs |
Generate and propagate IDs for cross-service tracing |
| Testing only happy paths |
Test error cases, edge cases, and auth failures |
| Skipping database migrations |
Use Alembic, Flyway, or Prisma Migrate — never ALTER in production manually |
| No health check endpoints |
Expose /health and /ready for orchestrator probes |
| Storing secrets in code/env files in git |
Use secret managers (Vault, AWS Secrets Manager, etc.) |
| Ignoring CORS until frontend integration |
Configure CORS from day one with explicit origins |
18. Production Readiness Checklist
Before shipping to production, verify:
API Design:
Security:
Database:
Observability:
Infrastructure:
1---2name: backend3description: Use this skill whenever building, designing, reviewing, or debugging backend systems, APIs, services, or server-side architecture. Triggers include any mention of REST API design, GraphQL, gRPC, database design, SQL optimization, caching (Redis, Memcached), authentication/authorization (JWT, OAuth), message queues (Kafka, RabbitMQ), microservices, monolith architecture, observability (logging, metrics, tracing), rate limiting, pagination, error handling patterns, API versioning, background jobs, webhooks, CI/CD, Docker, or deployment. Also use when the user asks to "design an API", "structure a backend", "optimize queries", "add caching", "set up auth", "handle errors", or any server-side architecture question. This skill is framework-agnostic — for framework-specific guidance (FastAPI, Express, Django, etc.), combine this with the relevant framework skill. Covers production patterns for 2025–2026: API design, database architecture, caching, security, observability, scalability, and deployment.4---5
6# Backend Skill — Production Backend Engineering (2025–2026)
7
8This skill covers framework-agnostic backend engineering principles: API design, database patterns, caching, security, observability, architecture decisions, and deployment. Combine with framework-specific skills (FastAPI, Express, Django, Spring, etc.) for implementation details.
9
10---
11
12## 1. API Design Fundamentals
13
14### REST URL conventions
15
16Resources are **nouns** (plural). HTTP methods are the **verbs**.
17
18```
19GET /api/v1/users # List users
20POST /api/v1/users # Create a user
21GET /api/v1/users/123 # Get user 123
22PATCH /api/v1/users/123 # Partial update
23PUT /api/v1/users/123 # Full replace
24DELETE /api/v1/users/123 # Delete user 123
25
26GET /api/v1/users/123/orders # User's orders (shallow nesting)
27GET /api/v1/users/123/orders/456 # Specific order
28
29# ❌ Anti-patterns
30GET /api/v1/getAllUsers
31POST /api/v1/createUser
32POST /api/v1/deleteUser
33GET /api/v1/user?id=123
34```
35
36### URL design rules
37
38- Use **plural nouns**: `/users`, `/orders`, `/products`.
39- Use **lowercase with hyphens**: `/order-items`, not `/orderItems` or `/order_items`.
40- Keep nesting **shallow** (max 2 levels): `/users/123/orders` is fine; `/users/123/orders/456/items/789/reviews` is too deep — flatten to `/reviews?item_id=789`.
41- Use **query parameters** for filtering, sorting, pagination: `/users?role=admin&sort=-created_at&page=2`.
42- Avoid exposing internal IDs when possible — use UUIDs or slugs for public-facing resources.
43
44### HTTP methods and semantics
45
46| Method | Idempotent | Safe | Use for |
47|:--------|:-----------|:-----|:---------------------------------|
48| GET | Yes | Yes | Read resources |
49| POST | No | No | Create resources, trigger actions |
50| PUT | Yes | No | Full replace of a resource |
51| PATCH | Yes* | No | Partial update |
52| DELETE | Yes | No | Remove a resource |
53| HEAD | Yes | Yes | Like GET but headers only |
54| OPTIONS | Yes | Yes | CORS preflight, capabilities |
55
56*PATCH is idempotent when using merge-patch semantics (the common case).
57
58### HTTP status codes
59
60Use them correctly — they are part of your API contract.
61
62```
632xx — Success
64 200 OK — Standard success response
65 201 Created — Resource created (include Location header)
66 204 No Content — Success with no response body (DELETE, PUT)
67
683xx — Redirection
69 301 Moved Permanently — Resource URL changed permanently
70 304 Not Modified — Cached version is still valid
71
724xx — Client Errors
73 400 Bad Request — Malformed syntax or invalid input
74 401 Unauthorized — Missing or invalid authentication
75 403 Forbidden — Authenticated but insufficient permissions
76 404 Not Found — Resource doesn't exist
77 409 Conflict — State conflict (e.g., duplicate email)
78 422 Unprocessable Entity — Valid syntax but semantic errors
79 429 Too Many Requests — Rate limit exceeded (include Retry-After header)
80
815xx — Server Errors
82 500 Internal Server Error — Unhandled server error
83 502 Bad Gateway — Upstream service failure
84 503 Service Unavailable — Temporary overload or maintenance
85 504 Gateway Timeout — Upstream service timeout
86```
87
88### Consistent error response format
89
90Follow RFC 9457 (Problem Details for HTTP APIs):
91
92```json
93{
94 "type": "https://api.example.com/errors/validation",
95 "title": "Validation Error",
96 "status": 422,
97 "detail": "The email field is required and must be a valid email address.",
98 "instance": "/api/v1/users",
99 "errors": [
100 {
101 "field": "email",
102 "message": "This field is required.",
103 "code": "required"
104 }
105 ]
106}
107```
108
109Rules for error responses:
110- Always return a consistent structure across all endpoints.
111- Include machine-readable error codes alongside human-readable messages.
112- Include field-level validation errors so clients can fix all issues in one round-trip.
113- **Never leak internal details** (stack traces, SQL queries, file paths) in production errors.
114- Log the full error server-side; return a sanitized version to the client.
115
116---
117
118## 2. API Versioning
119
120### Strategies
121
122| Strategy | Example | Pros | Cons |
123|:--------------|:--------------------------------|:------------------------|:-------------------------|
124| URL path | `/api/v1/users` | Explicit, easy to route | URL pollution |
125| Header | `Accept: application/vnd.api+json;version=2` | Clean URLs | Hidden, harder to test |
126| Query param | `/api/users?version=2` | Simple | Inconsistent |
127
128**Recommendation**: Use **URL path versioning** (`/api/v1/`). It's the most explicit, easiest to route, cache, and document. When introducing a breaking change, create a new version (`/api/v2/`) and run both in parallel with a deprecation timeline.
129
130### Deprecation protocol
131
1321. Announce deprecation with a `Deprecation` header and `Sunset` header on old endpoints.
1332. Provide a migration guide.
1343. Give clients at least 6–12 months before removal.
1354. Monitor usage of deprecated endpoints before turning them off.
136
137---
138
139## 3. Pagination
140
141### Cursor-based pagination (recommended)
142
143Best for real-time data, large datasets, and infinite scroll. Immune to offset drift when data changes.
144
145```
146GET /api/v1/orders?limit=20&after=eyJpZCI6MTIzfQ
147
148Response:
149{
150 "data": [...],
151 "pagination": {
152 "has_next": true,
153 "next_cursor": "eyJpZCI6MTQzfQ",
154 "has_previous": true,
155 "previous_cursor": "eyJpZCI6MTI0fQ"
156 }
157}
158```
159
160### Offset-based pagination (simpler)
161
162Fine for admin dashboards, small datasets, or when users need to jump to specific pages.
163
164```
165GET /api/v1/users?page=3&per_page=25
166
167Response:
168{
169 "data": [...],
170 "pagination": {
171 "page": 3,
172 "per_page": 25,
173 "total": 482,
174 "total_pages": 20
175 }
176}
177```
178
179### Pagination rules
180
181- Always set a **maximum page size** (e.g., `per_page` capped at 100).
182- Default to a reasonable page size (20–50).
183- Include pagination metadata in every list response.
184- For very large datasets (millions of rows), cursor-based is the only performant option.
185
186---
187
188## 4. Filtering, Sorting, and Field Selection
189
190```
191# Filtering
192GET /api/v1/products?category=electronics&price_min=100&price_max=500&in_stock=true
193
194# Sorting (prefix with - for descending)
195GET /api/v1/products?sort=-price,name
196
197# Field selection (reduce payload size)
198GET /api/v1/users?fields=id,name,email
199
200# Full-text search
201GET /api/v1/products?q=wireless+headphones
202```
203
204Be consistent across all endpoints. Document every supported filter, sort field, and selection parameter.
205
206---
207
208## 5. Authentication & Authorization
209
210### Authentication methods by use case
211
212| Use case | Method |
213|:-----------------------|:------------------------|
214| User-facing apps | OAuth 2.0 / OIDC |
215| Server-to-server | API keys + HMAC signing |
216| Microservice-to-microservice | Mutual TLS (mTLS) or JWT |
217| Simple internal tools | API keys in headers |
218
219### JWT best practices
220
221```
222Authorization: Bearer eyJhbGciOiJSUzI1NiIs...
223```
224
225- Use **short-lived access tokens** (5–15 minutes) with **long-lived refresh tokens** (days/weeks).
226- Store JWTs in **httpOnly, Secure, SameSite=Strict cookies** for browser clients — never in localStorage.
227- Use **asymmetric signing** (RS256/ES256) for distributed systems so services can verify without the private key.
228- Include minimal claims: `sub` (user ID), `exp`, `iat`, `iss`, `roles/scopes`.
229- Validate `exp`, `iss`, and `aud` on every request.
230- Implement **token revocation** via a short blocklist or by keeping access tokens very short-lived.
231
232### OAuth 2.0 / OIDC
233
234- Use **Authorization Code flow with PKCE** for all client types (web, mobile, SPA).
235- Never use the Implicit flow — it's deprecated.
236- Use established providers (Auth0, Clerk, Keycloak, Supabase Auth) rather than rolling your own.
237- Always validate the `state` parameter to prevent CSRF.
238- Store client secrets server-side only.
239
240### Authorization patterns
241
242- **RBAC (Role-Based Access Control)**: Assign roles (admin, editor, viewer) with predefined permissions. Works for most applications.
243- **ABAC (Attribute-Based Access Control)**: Policies based on user attributes, resource attributes, and context. More flexible but more complex.
244- **Check authorization at every layer**: Middleware for route-level checks, service layer for business-logic checks. Never rely on a single layer.
245
246---
247
248## 6. Database Design
249
250### Schema design principles
251
252- **Normalize first, denormalize intentionally.** Start with 3NF; denormalize specific tables for read performance when you have evidence (not speculation) of a bottleneck.
253- **Use appropriate data types**: Don't store dates as strings, don't use TEXT for short fixed-length fields, use DECIMAL/NUMERIC for money (never FLOAT).
254- **Always have a primary key**: Prefer auto-incrementing integers for internal IDs; use UUIDs (v7 for sortability) for public-facing IDs.
255- **Add created_at and updated_at** to every table (with timezone).
256- **Use foreign keys** for referential integrity in OLTP databases.
257- **Index strategically**: Index columns used in WHERE, JOIN, ORDER BY. Composite indexes follow the leftmost prefix rule. Don't over-index — each index slows writes.
258
259### Query optimization
260
261```sql
262-- ✅ Use EXPLAIN ANALYZE to understand query plans
263EXPLAIN ANALYZE
264SELECT u.id, u.name, COUNT(o.id) as order_count
265FROM users u
266LEFT JOIN orders o ON o.user_id = u.id
267WHERE u.created_at > '2026-01-01'
268GROUP BY u.id, u.name
269ORDER BY order_count DESC
270LIMIT 20;
271
272-- ✅ Avoid SELECT * — fetch only needed columns
273SELECT id, name, email FROM users WHERE active = true;
274
275-- ❌ N+1 queries — the most common performance killer
276-- Fetching users, then looping to fetch each user's orders separately
277-- ✅ Fix: JOIN or batch fetch in a single query
278```
279
280### Common query pitfalls
281
282| Pitfall | Fix |
283|:--------|:----|
284| N+1 queries | Use JOINs, eager loading, or batch fetching |
285| Missing indexes on foreign keys | Always index FK columns |
286| Full table scans on large tables | Add appropriate indexes; use LIMIT |
287| Using OFFSET for deep pagination | Switch to cursor/keyset pagination |
288| Locking issues on hot tables | Use row-level locking; consider read replicas |
289| Not using connection pooling | Use PgBouncer, HikariCP, or framework-level pools |
290
291### SQL vs NoSQL decision
292
293| Choose SQL (PostgreSQL, MySQL) when | Choose NoSQL when |
294|:------------------------------------|:------------------|
295| Complex relationships and JOINs | Document-shaped data (MongoDB) |
296| ACID transactions required | High write throughput at scale (Cassandra) |
297| Structured, well-known schema | Key-value lookups / caching (Redis) |
298| Reporting and aggregation | Graph relationships (Neo4j) |
299| Most CRUD applications | Time-series data (InfluxDB, TimescaleDB) |
300
301**Default to PostgreSQL** for most applications. It handles JSON (jsonb), full-text search, geospatial data, and scales to millions of rows before you need anything exotic.
302
303---
304
305## 7. Caching
306
307### Caching layers
308
309```
310Client ──→ CDN ──→ API Gateway ──→ Application Cache ──→ Database
311 ↑ ↑ ↑
312 Static Response Redis/Memcached
313 assets caching or in-memory
314```
315
316### Caching strategies
317
318| Strategy | How it works | Best for |
319|:---------|:-------------|:---------|
320| **Cache-aside** (Lazy loading) | App checks cache → miss → query DB → write to cache | General-purpose; most common |
321| **Write-through** | App writes to cache AND DB simultaneously | Strong consistency requirements |
322| **Write-behind** (Write-back) | App writes to cache; cache async writes to DB | High write throughput |
323| **Read-through** | Cache itself fetches from DB on miss | Simpler app code |
324
325### Cache invalidation
326
327Cache invalidation is one of the two hard problems in CS. Strategies:
328
329- **TTL (Time-to-Live)**: Set expiration. Simple but can serve stale data.
330- **Event-driven invalidation**: Invalidate cache when the underlying data changes (e.g., after a write).
331- **Tag-based invalidation**: Tag cache entries; invalidate all entries with a given tag.
332- **Versioned keys**: Include a version number in the cache key; bump it on changes.
333
334### Redis patterns
335
336```
337# Cache-aside pattern (pseudocode)
338def get_user(user_id):
339 cached = redis.get(f"user:{user_id}")
340 if cached:
341 return deserialize(cached)
342
343 user = db.query("SELECT * FROM users WHERE id = %s", user_id)
344 redis.setex(f"user:{user_id}", 300, serialize(user)) # TTL: 5 min
345 return user
346
347# Distributed locking (prevent thundering herd)
348lock = redis.set(f"lock:user:{user_id}", "1", nx=True, ex=10)
349if lock:
350 # Only one process rebuilds the cache
351 ...
352```
353
354### What to cache
355
356- **Cache**: Expensive query results, computed aggregations, API responses from third parties, session data, frequently-read and rarely-changed data.
357- **Don't cache**: User-specific sensitive data without proper scoping, rapidly-changing data where staleness is unacceptable, one-time reads.
358
359---
360
361## 8. Rate Limiting
362
363Protect your API from abuse, ensure fair usage, and prevent cascading failures.
364
365### Algorithms
366
367| Algorithm | Behavior | Best for |
368|:----------|:---------|:---------|
369| **Token bucket** | Tokens refill at a steady rate; requests consume tokens | Smooth rate limiting with burst allowance |
370| **Sliding window** | Counts requests in a rolling time window | Precise rate limiting |
371| **Fixed window** | Counts requests per fixed time window | Simplest to implement |
372
373### Implementation
374
375```
376# Response headers for rate limiting
377X-RateLimit-Limit: 1000 # Max requests per window
378X-RateLimit-Remaining: 847 # Requests remaining
379X-RateLimit-Reset: 1712800000 # Unix timestamp when window resets
380Retry-After: 30 # Seconds to wait (on 429 response)
381```
382
383- Apply rate limits **per API key/user**, not just per IP (IPs are shared behind NATs/VPNs).
384- Use different limits for different tiers (free: 100/hr, pro: 10,000/hr).
385- Return `429 Too Many Requests` with a `Retry-After` header.
386- Implement at the **API gateway or reverse proxy** level for efficiency.
387- Use Redis or a distributed counter for multi-instance deployments.
388
389---
390
391## 9. Background Jobs & Async Processing
392
393### When to use background jobs
394
395- Email/notification sending
396- PDF/report generation
397- Image/video processing
398- Data import/export
399- Webhook delivery and retries
400- Scheduled tasks (cron-like)
401- Any operation that takes > 1–2 seconds
402
403### Architecture
404
405```
406API Server ──→ Message Queue ──→ Workers
407 │ (Redis, RabbitMQ, │
408 │ Kafka, SQS) │
409 │ ↓
410 └── Returns 202 Accepted Process job
411 with job status URL Update status
412 Notify on completion
413```
414
415### Job design principles
416
417- **Idempotent**: Jobs should be safe to retry. If a job runs twice, the result should be the same.
418- **Small and focused**: Each job does one thing. Chain jobs for complex workflows.
419- **Observable**: Log job start/end/failure. Track duration, success rate, and queue depth.
420- **Retry with backoff**: Use exponential backoff with jitter (e.g., 1s, 2s, 4s, 8s + random jitter).
421- **Dead letter queue**: Failed jobs after max retries go to a DLQ for manual inspection.
422- **Timeout**: Set maximum execution time to prevent zombie jobs.
423
424### Tools
425
426| Language/Ecosystem | Tool |
427|:-------------------|:-----|
428| Python | Celery + Redis/RabbitMQ, ARQ, Dramatiq |
429| Node.js | BullMQ + Redis, Temporal |
430| Java/Kotlin | Spring Batch, Quartz |
431| Language-agnostic | Kafka, RabbitMQ, AWS SQS, Temporal |
432
433---
434
435## 10. Security
436
437### OWASP Top 10 awareness (2025–2026)
438
439The most critical backend security risks. Every backend engineer should know these:
440
4411. **Broken Access Control** — Enforce authorization at every endpoint and data access layer.
4422. **Cryptographic Failures** — Use TLS everywhere. Hash passwords with bcrypt/argon2. Encrypt sensitive data at rest.
4433. **Injection** — Use parameterized queries / ORMs. Never concatenate user input into SQL/commands.
4444. **Insecure Design** — Threat model during design, not after deployment.
4455. **Security Misconfiguration** — Disable debug endpoints in production. Remove default credentials. Set security headers.
4466. **Vulnerable Components** — Keep dependencies updated. Audit with `npm audit`, `pip-audit`, Snyk, or Dependabot.
4477. **Authentication Failures** — Implement account lockout, rate limiting on login, MFA.
4488. **Data Integrity Failures** — Verify software updates and CI/CD pipelines. Use SRI for CDN assets.
4499. **Logging & Monitoring Failures** — Log security events. Set up alerts for anomalies.
45010. **SSRF** — Validate and sanitize URLs the server fetches. Allowlist permitted domains.
451
452### Security headers
453
454```
455Strict-Transport-Security: max-age=63072000; includeSubDomains; preload
456Content-Security-Policy: default-src 'self'; script-src 'self'
457X-Content-Type-Options: nosniff
458X-Frame-Options: DENY
459Referrer-Policy: strict-origin-when-cross-origin
460Permissions-Policy: camera=(), microphone=(), geolocation=()
461```
462
463### Input validation
464
465- Validate **all** input: request body, query params, path params, headers.
466- Validate on the server — client validation is for UX, not security.
467- Use schema validation libraries (Pydantic, Zod, Joi, JSON Schema).
468- Reject unexpected fields (strict mode).
469- Sanitize output when rendering user content to prevent XSS.
470
471### Password storage
472
473- **Never** store passwords in plain text or with reversible encryption.
474- Use **bcrypt** (cost factor ≥ 12) or **argon2id** (preferred for new systems).
475- Salt is built into bcrypt/argon2 — don't roll your own salting scheme.
476
477---
478
479## 11. Observability
480
481Observability = Logs + Metrics + Traces. Design it alongside your application, not as an afterthought.
482
483### The three pillars
484
485**Logs** — Structured event records.
486```json
487{
488 "timestamp": "2026-04-10T14:32:01Z",
489 "level": "error",
490 "service": "order-service",
491 "request_id": "req_abc123",
492 "user_id": "usr_456",
493 "message": "Payment processing failed",
494 "error": "stripe_timeout",
495 "duration_ms": 30042
496}
497```
498
499- Use **structured logging** (JSON) — not unstructured text.
500- Include **request IDs / correlation IDs** that propagate across services.
501- Log at appropriate levels: DEBUG (development only), INFO (significant events), WARN (recoverable issues), ERROR (failures requiring attention).
502- **Never log** passwords, tokens, PII, or credit card numbers.
503
504**Metrics** — Aggregated numerical measurements.
505
506Key metrics to track (the RED method):
507- **Rate**: Requests per second
508- **Errors**: Error rate (4xx, 5xx)
509- **Duration**: Response time (p50, p95, p99)
510
511Additional: CPU/memory usage, queue depth, active connections, cache hit ratio, DB query time.
512
513**Traces** — Request flow across services.
514
515- Use **OpenTelemetry** (the industry standard) for instrumentation.
516- Propagate trace context (`traceparent` header) across service boundaries.
517- Trace critical paths: API → service → database → external API.
518
519### Alerting rules
520
521- Alert on **symptoms** (high error rate, slow response times), not causes.
522- Set thresholds based on **SLOs** (Service Level Objectives), e.g., "99.9% of requests complete in < 500ms."
523- Use severity levels: critical (paging), warning (next business day), info (dashboard only).
524- Avoid alert fatigue — every alert should be actionable.
525
526### Tools
527
528| Category | Tools |
529|:---------|:------|
530| Logging | ELK Stack, Loki + Grafana, Datadog |
531| Metrics | Prometheus + Grafana, Datadog, CloudWatch |
532| Tracing | Jaeger, Tempo, Datadog APM, Honeycomb |
533| All-in-one | Datadog, New Relic, Grafana Cloud |
534| Instrumentation | OpenTelemetry (universal standard) |
535
536---
537
538## 12. Architecture Patterns
539
540### Monolith vs Microservices
541
542| | Modular Monolith | Microservices |
543|:--|:-----------------|:--------------|
544| **Start with** | ✅ Always start here | Migrate to when needed |
545| **Deploy** | Single artifact | Independent per service |
546| **Complexity** | Lower operational overhead | Higher (networking, observability, consistency) |
547| **Team size** | < 20–30 engineers | Large orgs with team boundaries |
548| **Data** | Single database | Database per service |
549
550**Start with a well-structured monolith.** Extract services only when you have clear domain boundaries and team/scaling needs that justify the operational complexity.
551
552### Layered architecture (for any backend)
553
554```
555┌─────────────────────────────────┐
556│ Presentation (Routes/Handlers) │ ← HTTP layer, input validation
557├─────────────────────────────────┤
558│ Service / Business Logic │ ← Domain rules, orchestration
559├─────────────────────────────────┤
560│ Repository / Data Access │ ← Database queries, external APIs
561├─────────────────────────────────┤
562│ Infrastructure │ ← DB connections, caches, queues
563└─────────────────────────────────┘
564```
565
566Rules:
567- Each layer only calls the layer directly below it.
568- Routes should be thin — delegate to services.
569- Services contain business logic — no HTTP concepts (no request/response objects).
570- Repositories abstract data access — services don't write SQL directly.
571
572### Communication patterns
573
574| Pattern | Protocol | Use when |
575|:--------|:---------|:---------|
576| Synchronous request/response | HTTP/REST, gRPC | Direct queries, CRUD |
577| Async messaging | RabbitMQ, Kafka, SQS | Decoupled processing, event-driven workflows |
578| Streaming | gRPC streams, WebSockets, SSE | Real-time updates, live feeds |
579| Webhooks | HTTP callbacks | Notifying external systems of events |
580
581### Resilience patterns
582
583- **Circuit breaker**: Stop calling a failing service; fail fast and recover gracefully.
584- **Retry with exponential backoff + jitter**: Retry transient failures without thundering herd.
585- **Timeout**: Set timeouts on all external calls. A missing timeout is a bug.
586- **Bulkhead**: Isolate failures — don't let one slow dependency bring down the entire system.
587- **Health checks**: Expose `/health` (liveness) and `/ready` (readiness) endpoints.
588
589---
590
591## 13. Webhooks
592
593### Sending webhooks
594
595- Sign payloads with HMAC-SHA256 so receivers can verify authenticity.
596- Use exponential backoff for retries (e.g., 5 attempts: 1m, 5m, 30m, 2h, 24h).
597- Consider successful delivery as 2xx response within a timeout (e.g., 30s).
598- Log all delivery attempts for debugging.
599- Provide a webhook event log / dashboard in your API.
600- Send a minimal payload with an event type and resource ID; let receivers fetch details via API.
601
602### Receiving webhooks
603
604- Verify the signature before processing.
605- Respond with `200 OK` quickly; process asynchronously.
606- Handle duplicate deliveries gracefully (idempotent processing).
607- Use a queue for processing to avoid blocking the webhook endpoint.
608
609---
610
611## 14. Testing Backend Systems
612
613### Testing pyramid
614
615```
616 ╱ E2E Tests ╲ ← Few: Full user flows
617 ╱ Integration ╲ ← Some: API endpoints with real DB
618 ╱ Unit Tests ╲ ← Many: Business logic, utilities
619 ╱────────────────────╲
620```
621
622### What to test at each level
623
624**Unit tests** (fast, isolated):
625- Business logic functions
626- Validation rules
627- Data transformations
628- Edge cases and error paths
629
630**Integration tests** (with real dependencies):
631- API endpoints end-to-end (HTTP → handler → service → DB → response)
632- Database queries against a test database
633- Cache behavior
634- Authentication/authorization flows
635
636**Contract tests** (for API consumers):
637- Verify your API responses match the documented schema.
638- Use tools like Pact or Schemathesis.
639
640**Load tests** (before production):
641- Use k6, Locust, or Gatling.
642- Test at 2–3× expected peak traffic.
643- Identify bottlenecks: slow queries, memory leaks, connection exhaustion.
644
645### Testing rules
646
647- Test **behavior**, not implementation.
648- Use **factories** (Factory Boy, Faker) to generate test data — not fixtures.
649- Isolate tests — each test should set up and tear down its own data.
650- Test **error paths** as thoroughly as happy paths.
651- Run tests in CI on every pull request.
652
653---
654
655## 15. Deployment & Infrastructure
656
657### Container best practices (Docker)
658
659```dockerfile
660# Multi-stage build for smaller images
661FROM python:3.12-slim AS builder
662WORKDIR /app
663COPY requirements.txt .
664RUN pip install --no-cache-dir -r requirements.txt
665
666FROM python:3.12-slim
667WORKDIR /app
668COPY --from=builder /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packages
669COPY . .
670# Run as non-root user
671RUN adduser --disabled-password appuser
672USER appuser
673EXPOSE 8000
674CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]
675```
676
677- Use **multi-stage builds** for smaller images.
678- Run as a **non-root user**.
679- Pin **exact versions** for base images and dependencies.
680- Use `.dockerignore` to exclude unnecessary files.
681- Scan images for vulnerabilities (Trivy, Snyk).
682
683### CI/CD pipeline
684
685```
686Push → Lint → Unit Tests → Build Image → Integration Tests → Security Scan → Deploy to Staging → E2E Tests → Deploy to Production
687```
688
689- **Every commit** triggers lint + unit tests.
690- **Every PR merge** triggers full pipeline including integration tests.
691- Use **blue-green** or **canary deployments** for zero-downtime releases.
692- **Rollback** must be a single command / automated on error spike.
693
694### Health checks
695
696```json
697// GET /health (liveness — is the process alive?)
698{ "status": "ok" }
699
700// GET /ready (readiness — can the service handle traffic?)
701{
702 "status": "ok",
703 "checks": {
704 "database": "connected",
705 "redis": "connected",
706 "queue": "connected"
707 }
708}
709```
710
711---
712
713## 16. API Documentation
714
715- Use **OpenAPI 3.1** (Swagger) for REST APIs.
716- Generate docs **from code** (not maintained separately) — most frameworks support this.
717- Include: authentication instructions, endpoint references, request/response examples, error format, rate limits, pagination details.
718- Provide **runnable examples** (curl, HTTPie, or language-specific snippets).
719- Keep docs in sync — stale docs are worse than no docs.
720- Consider publishing an `llms.txt` for AI agent consumption.
721
722---
723
724## 17. Common Pitfalls
725
726| Pitfall | Fix |
727|:--------|:----|
728| No input validation | Validate everything server-side with schema libraries |
729| Leaking internal errors to clients | Return generic errors; log details server-side |
730| N+1 database queries | Use JOINs, eager loading, or DataLoader patterns |
731| No rate limiting | Implement at API gateway level; return 429 with Retry-After |
732| OFFSET-based pagination on large datasets | Switch to cursor/keyset pagination |
733| Rolling your own auth | Use established providers or battle-tested libraries |
734| No timeouts on external calls | Set timeouts on every HTTP client, DB connection, and queue consumer |
735| Monolithic error handling | Use structured error types with consistent response format |
736| No request/correlation IDs | Generate and propagate IDs for cross-service tracing |
737| Testing only happy paths | Test error cases, edge cases, and auth failures |
738| Skipping database migrations | Use Alembic, Flyway, or Prisma Migrate — never ALTER in production manually |
739| No health check endpoints | Expose `/health` and `/ready` for orchestrator probes |
740| Storing secrets in code/env files in git | Use secret managers (Vault, AWS Secrets Manager, etc.) |
741| Ignoring CORS until frontend integration | Configure CORS from day one with explicit origins |
742
743---
744
745## 18. Production Readiness Checklist
746
747Before shipping to production, verify:
748
749**API Design:**
750- [ ] Consistent URL structure and naming
751- [ ] Proper HTTP status codes on all endpoints
752- [ ] Structured error responses (RFC 9457)
753- [ ] API versioning in place
754- [ ] Pagination on all list endpoints
755- [ ] Rate limiting configured
756
757**Security:**
758- [ ] HTTPS only (TLS 1.2+ minimum)
759- [ ] Authentication on all non-public endpoints
760- [ ] Authorization checks at route AND service layer
761- [ ] Input validation on all endpoints
762- [ ] Security headers set (HSTS, CSP, etc.)
763- [ ] Secrets in secret manager (not in code)
764- [ ] Dependencies scanned for vulnerabilities
765
766**Database:**
767- [ ] Connection pooling configured
768- [ ] Indexes on frequently-queried columns and foreign keys
769- [ ] Migrations managed with a migration tool
770- [ ] Backups automated and tested
771
772**Observability:**
773- [ ] Structured logging with correlation IDs
774- [ ] Key metrics exported (request rate, error rate, latency)
775- [ ] Alerting configured on SLO thresholds
776- [ ] Health check endpoints (`/health`, `/ready`)
777
778**Infrastructure:**
779- [ ] Containerized with non-root user
780- [ ] CI/CD pipeline with automated tests
781- [ ] Rollback procedure documented and tested
782- [ ] Load tested at 2–3× expected peak