Monitoring & Observability
Comprehensive patterns for infrastructure monitoring, LLM observability, and quality drift detection. Each category has individual rule files in rules/ loaded on-demand.
Quick Reference
| Category |
Rules |
Impact |
When to Use |
| Infrastructure Monitoring |
3 |
CRITICAL |
Prometheus metrics, Grafana dashboards, alerting rules |
| LLM Observability |
3 |
HIGH |
Langfuse tracing, cost tracking, evaluation scoring |
| Drift Detection |
3 |
HIGH |
Statistical drift, quality regression, drift alerting |
| Silent Failures |
3 |
HIGH |
Tool skipping, quality degradation, loop/token spike alerting |
Total: 12 rules across 4 categories
Quick Start
# Prometheus metrics with RED method
from prometheus_client import Counter, Histogram
http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])
# Langfuse LLM tracing
from langfuse import observe, get_client
@observe()
async def analyze_content(content: str):
get_client().update_current_trace(
user_id="user_123", session_id="session_abc",
tags=["production", "orchestkit"],
)
return await llm.generate(content)
# PSI drift detection
import numpy as np
psi_score = calculate_psi(baseline_scores, current_scores)
if psi_score >= 0.25:
alert("Significant quality drift detected!")
Infrastructure Monitoring
Prometheus metrics, Grafana dashboards, and alerting for application health.
| Rule |
File |
Key Pattern |
| Prometheus Metrics |
rules/monitoring-prometheus.md |
RED method, counters, histograms, cardinality |
| Grafana Dashboards |
rules/monitoring-grafana.md |
Golden Signals, SLO/SLI, health checks |
| Alerting Rules |
rules/monitoring-alerting.md |
Severity levels, grouping, escalation, fatigue prevention |
LLM Observability
Langfuse-based tracing, cost tracking, and evaluation for LLM applications.
| Rule |
File |
Key Pattern |
| Langfuse Traces |
rules/llm-langfuse-traces.md |
@observe decorator, OTEL spans, agent graphs |
| Cost Tracking |
rules/llm-cost-tracking.md |
Token usage, spend alerts, Metrics API |
| Eval Scoring |
rules/llm-eval-scoring.md |
Custom scores, evaluator tracing, quality monitoring |
Drift Detection
Statistical and quality drift detection for production LLM systems.
| Rule |
File |
Key Pattern |
| Statistical Drift |
rules/drift-statistical.md |
PSI, KS test, KL divergence, EWMA |
| Quality Drift |
rules/drift-quality.md |
Score regression, baseline comparison, canary prompts |
| Drift Alerting |
rules/drift-alerting.md |
Dynamic thresholds, correlation, anti-patterns |
Silent Failures
Detection and alerting for silent failures in LLM agents.
| Rule |
File |
Key Pattern |
| Tool Skipping |
rules/silent-tool-skipping.md |
Expected vs actual tool calls, Langfuse traces |
| Quality Degradation |
rules/silent-degraded-quality.md |
Heuristics + LLM-as-judge, z-score baselines |
| Silent Alerting |
rules/silent-alerting.md |
Loop detection, token spikes, escalation workflow |
Key Decisions
| Decision |
Recommendation |
Rationale |
| Metric methodology |
RED method (Rate, Errors, Duration) |
Industry standard, covers essential service health |
| Log format |
Structured JSON |
Machine-parseable, supports log aggregation |
| Tracing |
OpenTelemetry |
Vendor-neutral, auto-instrumentation, broad ecosystem |
| LLM observability |
Langfuse (not LangSmith) |
Open-source, self-hosted, built-in prompt management |
| LLM tracing API |
@observe + get_client() |
OTEL-native, automatic span creation |
| Drift method |
PSI for production, KS for small samples |
PSI is stable for large datasets, KS more sensitive |
| Threshold strategy |
Dynamic (95th percentile) over static |
Reduces alert fatigue, context-aware |
| Alert severity |
4 levels (Critical, High, Medium, Low) |
Clear escalation paths, appropriate response times |
Detailed Documentation
| Resource |
Description |
| references/ |
Logging, metrics, tracing, Langfuse, drift analysis guides |
| checklists/ |
Implementation checklists for monitoring and Langfuse setup |
| examples/ |
Real-world monitoring dashboard and trace examples |
| scripts/ |
Templates: Prometheus, OpenTelemetry, health checks, Langfuse |
Related Skills
defense-in-depth - Layer 8 observability as part of security architecture
devops-deployment - Observability integration with CI/CD and Kubernetes
resilience-patterns - Monitoring circuit breakers and failure scenarios
llm-evaluation - Evaluation patterns that integrate with Langfuse scoring
caching - Caching strategies that reduce costs tracked by Langfuse
1---2name: monitoring-observability3description: Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse LLM tracing, and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.4license: MIT5---67# Monitoring & Observability89Comprehensive patterns for infrastructure monitoring, LLM observability, and quality drift detection. Each category has individual rule files in `rules/` loaded on-demand.1011## Quick Reference1213| Category | Rules | Impact | When to Use |14|----------|-------|--------|-------------|15| [Infrastructure Monitoring](#infrastructure-monitoring) | 3 | CRITICAL | Prometheus metrics, Grafana dashboards, alerting rules |16| [LLM Observability](#llm-observability) | 3 | HIGH | Langfuse tracing, cost tracking, evaluation scoring |17| [Drift Detection](#drift-detection) | 3 | HIGH | Statistical drift, quality regression, drift alerting |18| [Silent Failures](#silent-failures) | 3 | HIGH | Tool skipping, quality degradation, loop/token spike alerting |1920**Total: 12 rules across 4 categories**2122## Quick Start2324```python25# Prometheus metrics with RED method26from prometheus_client import Counter, Histogram2728http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])29http_duration = Histogram('http_request_duration_seconds', 'Request latency',30 buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])31```3233```python34# Langfuse LLM tracing35from langfuse import observe, get_client3637@observe()38async def analyze_content(content: str):39 get_client().update_current_trace(40 user_id="user_123", session_id="session_abc",41 tags=["production", "orchestkit"],42 )43 return await llm.generate(content)44```4546```python47# PSI drift detection48import numpy as np4950psi_score = calculate_psi(baseline_scores, current_scores)51if psi_score >= 0.25:52 alert("Significant quality drift detected!")53```5455## Infrastructure Monitoring5657Prometheus metrics, Grafana dashboards, and alerting for application health.5859| Rule | File | Key Pattern |60|------|------|-------------|61| Prometheus Metrics | `rules/monitoring-prometheus.md` | RED method, counters, histograms, cardinality |62| Grafana Dashboards | `rules/monitoring-grafana.md` | Golden Signals, SLO/SLI, health checks |63| Alerting Rules | `rules/monitoring-alerting.md` | Severity levels, grouping, escalation, fatigue prevention |6465## LLM Observability6667Langfuse-based tracing, cost tracking, and evaluation for LLM applications.6869| Rule | File | Key Pattern |70|------|------|-------------|71| Langfuse Traces | `rules/llm-langfuse-traces.md` | @observe decorator, OTEL spans, agent graphs |72| Cost Tracking | `rules/llm-cost-tracking.md` | Token usage, spend alerts, Metrics API |73| Eval Scoring | `rules/llm-eval-scoring.md` | Custom scores, evaluator tracing, quality monitoring |7475## Drift Detection7677Statistical and quality drift detection for production LLM systems.7879| Rule | File | Key Pattern |80|------|------|-------------|81| Statistical Drift | `rules/drift-statistical.md` | PSI, KS test, KL divergence, EWMA |82| Quality Drift | `rules/drift-quality.md` | Score regression, baseline comparison, canary prompts |83| Drift Alerting | `rules/drift-alerting.md` | Dynamic thresholds, correlation, anti-patterns |8485## Silent Failures8687Detection and alerting for silent failures in LLM agents.8889| Rule | File | Key Pattern |90|------|------|-------------|91| Tool Skipping | `rules/silent-tool-skipping.md` | Expected vs actual tool calls, Langfuse traces |92| Quality Degradation | `rules/silent-degraded-quality.md` | Heuristics + LLM-as-judge, z-score baselines |93| Silent Alerting | `rules/silent-alerting.md` | Loop detection, token spikes, escalation workflow |9495## Key Decisions9697| Decision | Recommendation | Rationale |98|----------|----------------|-----------|99| Metric methodology | RED method (Rate, Errors, Duration) | Industry standard, covers essential service health |100| Log format | Structured JSON | Machine-parseable, supports log aggregation |101| Tracing | OpenTelemetry | Vendor-neutral, auto-instrumentation, broad ecosystem |102| LLM observability | Langfuse (not LangSmith) | Open-source, self-hosted, built-in prompt management |103| LLM tracing API | `@observe` + `get_client()` | OTEL-native, automatic span creation |104| Drift method | PSI for production, KS for small samples | PSI is stable for large datasets, KS more sensitive |105| Threshold strategy | Dynamic (95th percentile) over static | Reduces alert fatigue, context-aware |106| Alert severity | 4 levels (Critical, High, Medium, Low) | Clear escalation paths, appropriate response times |107108## Detailed Documentation109110| Resource | Description |111|----------|-------------|112| [references/](references/) | Logging, metrics, tracing, Langfuse, drift analysis guides |113| [checklists/](checklists/) | Implementation checklists for monitoring and Langfuse setup |114| [examples/](examples/) | Real-world monitoring dashboard and trace examples |115| [scripts/](scripts/) | Templates: Prometheus, OpenTelemetry, health checks, Langfuse |116117## Related Skills118119- `defense-in-depth` - Layer 8 observability as part of security architecture120- `devops-deployment` - Observability integration with CI/CD and Kubernetes121- `resilience-patterns` - Monitoring circuit breakers and failure scenarios122- `llm-evaluation` - Evaluation patterns that integrate with Langfuse scoring123- `caching` - Caching strategies that reduce costs tracked by Langfuse