1---2name: backend3description: Build reliable backend services with proper error handling, security, and observability.4---56## Error Handling78- Never expose stack traces to clients—log internally, return generic message9- Structured error responses: code, message, request ID—enables debugging without leaking10- Fail fast on bad input—validate at entry point, not deep in business logic11- Unexpected errors: 500 + alert—expected errors: appropriate 4xx1213## Input Validation1415- Validate everything from outside—query params, headers, body, path params16- Whitelist valid input, don't blacklist bad—reject unknown fields17- Validate early, before any processing—save resources, clearer errors18- Size limits on all inputs—prevent memory exhaustion attacks1920## Timeouts Everywhere2122- Database queries: set timeout, typically 5-30s23- External HTTP calls: connect timeout + read timeout—don't wait forever24- Overall request timeout—gateway or middleware level25- Background jobs: max execution time—prevent zombie processes2627## Retry Patterns2829- Exponential backoff: 1s, 2s, 4s, 8s...—prevents thundering herd30- Add jitter: randomize delay—prevents synchronized retries31- Idempotency keys for non-idempotent operations—safe to retry32- Circuit breaker for failing dependencies—stop hammering, fail fast3334## Database Practices3536- Connection pooling: reuse connections—creating is expensive37- Transactions scoped minimal—hold locks briefly38- Read replicas for read-heavy workloads—separate read/write traffic39- Prepared statements always—SQL injection prevention, query plan cache4041## Caching Strategy4243- Cache invalidation strategy decided upfront—TTL, event-based, or both44- Cache at right layer: query result, computed value, HTTP response45- Cache stampede prevention—lock or probabilistic early expiration46- Monitor hit rate—low hit rate = wasted resources4748## Rate Limiting4950- Per-user/IP limits on expensive operations—login, signup, search51- Different limits for different operations—read vs write52- Return Retry-After header—tell clients when to retry53- Rate limit early in request pipeline—save resources5455## Health Checks5657- Liveness: is process running—restart if fails58- Readiness: can handle traffic—remove from load balancer if fails59- Startup probe for slow-starting services—don't kill during init60- Health checks fast and cheap—don't hit database on every probe6162## Graceful Shutdown6364- Stop accepting new requests first—drain load balancer65- Wait for in-flight requests to complete—with timeout66- Close database connections cleanly—prevent connection leaks67- SIGTERM handling: graceful; SIGKILL after timeout6869## Logging7071- Structured logs (JSON)—parseable by log aggregators72- Request ID in every log—trace request across services73- Log level appropriate: debug for dev, info/error for prod74- Sensitive data never logged—passwords, tokens, PII7576## API Design7778- Versioning strategy from day one—path (/v1/) or header79- Pagination for list endpoints—cursor or offset; include total count80- Consistent response format—same envelope everywhere81- Meaningful status codes—201 for create, 204 for delete, 404 for not found8283## Security Hygiene8485- Secrets from environment or vault—never in code or config files86- Dependencies updated regularly—automated with Dependabot/Renovate87- Principle of least privilege—service accounts with minimal permissions88- Authentication and authorization separated—who you are vs what you can do8990## Observability9192- Metrics: request count, latency percentiles, error rate—the RED method93- Distributed tracing for microservices—follow request across services94- Alerting on symptoms, not causes—high error rate, not CPU usage95- Dashboards for operational visibility—know normal to spot abnormal