Monitoring and Alerts
Current guidance for monitoring LangSmith projects and configuring alerts.
Monitoring in LangSmith
LangSmith provides a Monitoring area with:
- Prebuilt dashboards (automatically generated per tracing project)
- Custom dashboards (user-defined charts/filters)
Reference: https://docs.langchain.com/langsmith/dashboards
Alerts overview
Alerts are project-scoped, so configure them per project.
LangSmith alerting is threshold-based on these metric types:
- Errored Runs
- Feedback Score
- Latency
Reference: https://docs.langchain.com/langsmith/alerts
Self-hosted requirement
For self-hosted LangSmith, alerts require Helm chart 0.10.3 or later.
Reference: https://docs.langchain.com/langsmith/alerts
Creating alert conditions
Alert conditions include:
- Aggregation method (
Average,Percentage, orCount) - Comparison operator (
>=,<=, or exceeds threshold) - Threshold value
- Aggregation window (currently
5or15minutes) - Feedback key (for Feedback Score alerts)
Reference: https://docs.langchain.com/langsmith/alerts
Webhook notifications
LangSmith supports webhook notifications for alerts.
Required webhook field:
URL
Optional webhook fields:
Headers(JSON map)Request Body Template
When webhooks fire, LangSmith appends metadata fields such as:
project_namealert_rule_idalert_rule_namealert_rule_type(currently threshold)alert_rule_attribute(error_count,feedback_score,latency)triggered_metric_valuetriggered_thresholdtimestamp
Reference: https://docs.langchain.com/langsmith/alerts-webhook
Practical setup flow
- Open target project in LangSmith.
- Go to Alerts -> Create Alert.
- Choose metric type and condition.
- Add notification channel (email/webhook).
- Save and validate by triggering a test condition.
References:
Suggested production baseline
- One Errored Runs alert for sustained failure spikes.
- One Latency alert for user-facing performance regressions.
- One Feedback Score alert for quality drift.
Tune thresholds using recent historical behavior before enabling paging notifications.
Dashboard and alert checklist
- Create a custom dashboard per production project.
- Add charts for error trend, latency trend, and feedback trend.
- Confirm alerts route to on-call channel.
- Revisit thresholds after major model/prompt changes.
- Keep alert ownership explicit (team or rotation).