1---2name: backend3description: Build reliable backend services with proper error handling, security, and observability.4---5
6## Error Handling
7
8- Never expose stack traces to clients—log internally, return generic message
9- Structured error responses: code, message, request ID—enables debugging without leaking
10- Fail fast on bad input—validate at entry point, not deep in business logic
11- Unexpected errors: 500 + alert—expected errors: appropriate 4xx
12
13## Input Validation
14
15- Validate everything from outside—query params, headers, body, path params
16- Whitelist valid input, don't blacklist bad—reject unknown fields
17- Validate early, before any processing—save resources, clearer errors
18- Size limits on all inputs—prevent memory exhaustion attacks
19
20## Timeouts Everywhere
21
22- Database queries: set timeout, typically 5-30s
23- External HTTP calls: connect timeout + read timeout—don't wait forever
24- Overall request timeout—gateway or middleware level
25- Background jobs: max execution time—prevent zombie processes
26
27## Retry Patterns
28
29- Exponential backoff: 1s, 2s, 4s, 8s...—prevents thundering herd
30- Add jitter: randomize delay—prevents synchronized retries
31- Idempotency keys for non-idempotent operations—safe to retry
32- Circuit breaker for failing dependencies—stop hammering, fail fast
33
34## Database Practices
35
36- Connection pooling: reuse connections—creating is expensive
37- Transactions scoped minimal—hold locks briefly
38- Read replicas for read-heavy workloads—separate read/write traffic
39- Prepared statements always—SQL injection prevention, query plan cache
40
41## Caching Strategy
42
43- Cache invalidation strategy decided upfront—TTL, event-based, or both
44- Cache at right layer: query result, computed value, HTTP response
45- Cache stampede prevention—lock or probabilistic early expiration
46- Monitor hit rate—low hit rate = wasted resources
47
48## Rate Limiting
49
50- Per-user/IP limits on expensive operations—login, signup, search
51- Different limits for different operations—read vs write
52- Return Retry-After header—tell clients when to retry
53- Rate limit early in request pipeline—save resources
54
55## Health Checks
56
57- Liveness: is process running—restart if fails
58- Readiness: can handle traffic—remove from load balancer if fails
59- Startup probe for slow-starting services—don't kill during init
60- Health checks fast and cheap—don't hit database on every probe
61
62## Graceful Shutdown
63
64- Stop accepting new requests first—drain load balancer
65- Wait for in-flight requests to complete—with timeout
66- Close database connections cleanly—prevent connection leaks
67- SIGTERM handling: graceful; SIGKILL after timeout
68
69## Logging
70
71- Structured logs (JSON)—parseable by log aggregators
72- Request ID in every log—trace request across services
73- Log level appropriate: debug for dev, info/error for prod
74- Sensitive data never logged—passwords, tokens, PII
75
76## API Design
77
78- Versioning strategy from day one—path (/v1/) or header
79- Pagination for list endpoints—cursor or offset; include total count
80- Consistent response format—same envelope everywhere
81- Meaningful status codes—201 for create, 204 for delete, 404 for not found
82
83## Security Hygiene
84
85- Secrets from environment or vault—never in code or config files
86- Dependencies updated regularly—automated with Dependabot/Renovate
87- Principle of least privilege—service accounts with minimal permissions
88- Authentication and authorization separated—who you are vs what you can do
89
90## Observability
91
92- Metrics: request count, latency percentiles, error rate—the RED method
93- Distributed tracing for microservices—follow request across services
94- Alerting on symptoms, not causes—high error rate, not CPU usage
95- Dashboards for operational visibility—know normal to spot abnormal