Observability

Instrument a service so the questions asked during an incident are answerable from data already being collected — rate, errors, latency distribution and saturation per route, structured events carrying a correlation id that survives process and queue boundaries, and alerts on symptoms users feel. Use before a service or a new critical path goes to production, after an incident that ended in "we had no data for that", or when adding a dependency, queue or job whose failure would be silent. Not for fighting a live outage (incident-response), not for profiling a known-slow path (performance-profiling), and not for tracing agent or LLM runs (llm-observability).

nahid-sparktales Updated

File contents

nahid-sparktales/agent-dispatcher/tree/main/skills/devops/observability commit 7c5a1112ab

Frequently asked questions

npx skillmds@latest add nahid-sparktales/observability