Loki Best Practices

Operational Grafana Loki troubleshooting for already-deployed stacks — ingest rejections (rate/stream limits, out-of-order, line-too-long, invalid labels), WAL replay and chunk-flush failures, stuck compactors, ring/memberlist health, limits_config and per-tenant overrides, label cardinality, Promtail/Alloy/OTel push errors, slow LogQL queries, IRSA/object-storage breakage, and schema-period upgrades. Use whenever the user mentions Loki, LogQL, Promtail, Alloy, log ingestion or discards, `loki_discarded_samples_total`, the `/ring`, WAL replay, the Loki Helm chart, log cardinality, or any Loki incident, push failure, slow query, schema upgrade, or HA problem — even in a Grafana-stack log context that doesn't say "Loki". Do NOT use for greenfield log-tooling choices — this assumes Loki is already deployed. (Full trigger list and per-surface diagnostics are in the body's "When to use" section.)

andreab67 22b55b6 3 files · 22.9 KB Updated

File contents

andreab67/agent-skills/tree/main/loki-best-practices commit 22b55b685d

Frequently asked questions

npx skillmds@latest add andreab67/loki-best-practices