Azure Diagnostics
AUTHORITATIVE GUIDANCE — MANDATORY COMPLIANCE
This document is the official source for debugging and troubleshooting Azure production issues. Follow these instructions to diagnose and resolve common Azure service problems systematically.
Triggers
Activate this skill when user wants to:
- Debug or troubleshoot production issues
- Diagnose errors in Azure services
- Analyze application logs or metrics
- Fix image pull, cold start, or health probe issues
- Investigate why Azure resources are failing
- Find root cause of application errors
- Troubleshoot Azure Function Apps (invocation failures, timeouts, binding errors)
- Find the App Insights or Log Analytics workspace linked to a Function App
- Troubleshoot AKS clusters, nodes, pods, ingress, or Kubernetes networking issues
Rules
- Start with systematic diagnosis flow
- Use AppLens (MCP) for AI-powered diagnostics when available
- Check resource health before deep-diving into logs
- Select appropriate troubleshooting guide based on service type
- Document findings and attempted remediation steps
- Route AKS incidents to the dedicated AKS troubleshooting document
Quick Diagnosis Flow
- Identify symptoms - What's failing?
- Check resource health - Is Azure healthy?
- Review logs - What do logs show?
- Analyze metrics - Performance patterns?
- Investigate recent changes - What changed?
Troubleshooting Guides by Service
| Service |
Common Issues |
Reference |
| Container Apps |
Image pull failures, cold starts, health probes, port mismatches |
container-apps/ |
| Function Apps |
App details, invocation failures, timeouts, binding errors, cold starts, missing app settings |
functions/ |
| AKS |
Cluster access, nodes, kube-system, scheduling, crash loops, ingress, DNS, upgrades |
AKS Troubleshooting |
Routing
- Keep Container Apps and Function Apps diagnostics in this parent skill.
- Route active AKS incidents, AKS-specific intake, evidence gathering, and remediation guidance to AKS Troubleshooting.
Quick Reference
Common Diagnostic Commands
# Check resource health
az resource show --ids RESOURCE_ID
# View activity log
az monitor activity-log list -g RG --max-events 20
# Container Apps logs
az containerapp logs show --name APP -g RG --follow
# Function App logs (query App Insights traces)
az monitor app-insights query --apps APP-INSIGHTS -g RG \
--analytics-query "traces | where timestamp > ago(1h) | order by timestamp desc | take 50"
AppLens (MCP Tools)
For AI-powered diagnostics, use:
mcp_azure_mcp_applens
intent: "diagnose issues with <resource-name>"
command: "diagnose"
parameters:
resourceId: "<resource-id>"
Provides:
- Automated issue detection
- Root cause analysis
- Remediation recommendations
Azure Monitor (MCP Tools)
For querying logs and metrics:
mcp_azure_mcp_monitor
intent: "query logs for <resource-name>"
command: "logs_query"
parameters:
workspaceId: "<workspace-id>"
query: "<KQL-query>"
See kql-queries.md for common diagnostic queries.
Check Azure Resource Health
Using MCP
mcp_azure_mcp_resourcehealth
intent: "check health status of <resource-name>"
command: "get"
parameters:
resourceId: "<resource-id>"
Using CLI
# Check specific resource health
az resource show --ids RESOURCE_ID
# Check recent activity
az monitor activity-log list -g RG --max-events 20
References
- KQL Query Library
- Azure Resource Graph Queries
- Function Apps Troubleshooting
When to Use
Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.
Covers: Triggers, Quick Diagnosis Flow, Troubleshooting Guides by Service, Routing, Common Diagnostic Commands.
1---2name: azure-diagnostics3description: Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready...4license: MIT5---6# Azure Diagnostics78> **AUTHORITATIVE GUIDANCE — MANDATORY COMPLIANCE**9>10> This document is the **official source** for debugging and troubleshooting Azure production issues. Follow these instructions to diagnose and resolve common Azure service problems systematically.1112## Triggers1314Activate this skill when user wants to:15- Debug or troubleshoot production issues16- Diagnose errors in Azure services17- Analyze application logs or metrics18- Fix image pull, cold start, or health probe issues19- Investigate why Azure resources are failing20- Find root cause of application errors21- Troubleshoot Azure Function Apps (invocation failures, timeouts, binding errors)22- Find the App Insights or Log Analytics workspace linked to a Function App23- Troubleshoot AKS clusters, nodes, pods, ingress, or Kubernetes networking issues2425## Rules26271. Start with systematic diagnosis flow282. Use AppLens (MCP) for AI-powered diagnostics when available293. Check resource health before deep-diving into logs304. Select appropriate troubleshooting guide based on service type315. Document findings and attempted remediation steps326. Route AKS incidents to the dedicated AKS troubleshooting document3334---3536## Quick Diagnosis Flow37381. **Identify symptoms** - What's failing?392. **Check resource health** - Is Azure healthy?403. **Review logs** - What do logs show?414. **Analyze metrics** - Performance patterns?425. **Investigate recent changes** - What changed?4344---4546## Troubleshooting Guides by Service4748| Service | Common Issues | Reference |49|---------|---------------|-----------|50| **Container Apps** | Image pull failures, cold starts, health probes, port mismatches | [container-apps/](references/container-apps/README.md) |51| **Function Apps** | App details, invocation failures, timeouts, binding errors, cold starts, missing app settings | [functions/](references/functions/README.md) |52| **AKS** | Cluster access, nodes, `kube-system`, scheduling, crash loops, ingress, DNS, upgrades | [AKS Troubleshooting](aks-troubleshooting/aks-troubleshooting.md) |5354---5556## Routing5758- Keep Container Apps and Function Apps diagnostics in this parent skill.59- Route active AKS incidents, AKS-specific intake, evidence gathering, and remediation guidance to [AKS Troubleshooting](aks-troubleshooting/aks-troubleshooting.md).6061---6263## Quick Reference6465### Common Diagnostic Commands6667```bash68# Check resource health69az resource show --ids RESOURCE_ID7071# View activity log72az monitor activity-log list -g RG --max-events 207374# Container Apps logs75az containerapp logs show --name APP -g RG --follow7677# Function App logs (query App Insights traces)78az monitor app-insights query --apps APP-INSIGHTS -g RG \79 --analytics-query "traces | where timestamp > ago(1h) | order by timestamp desc | take 50"80```8182### AppLens (MCP Tools)8384For AI-powered diagnostics, use:85```86mcp_azure_mcp_applens87 intent: "diagnose issues with <resource-name>"88 command: "diagnose"89 parameters:90 resourceId: "<resource-id>"9192Provides:93- Automated issue detection94- Root cause analysis95- Remediation recommendations96```9798### Azure Monitor (MCP Tools)99100For querying logs and metrics:101```102mcp_azure_mcp_monitor103 intent: "query logs for <resource-name>"104 command: "logs_query"105 parameters:106 workspaceId: "<workspace-id>"107 query: "<KQL-query>"108```109110See [kql-queries.md](references/kql-queries.md) for common diagnostic queries.111112---113114## Check Azure Resource Health115116### Using MCP117118```119mcp_azure_mcp_resourcehealth120 intent: "check health status of <resource-name>"121 command: "get"122 parameters:123 resourceId: "<resource-id>"124```125126### Using CLI127128```bash129# Check specific resource health130az resource show --ids RESOURCE_ID131132# Check recent activity133az monitor activity-log list -g RG --max-events 20134```135136---137138## References139140- [KQL Query Library](references/kql-queries.md)141- [Azure Resource Graph Queries](references/azure-resource-graph.md)142- [Function Apps Troubleshooting](references/functions/README.md)143144## When to Use145146Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.147148Covers: Triggers, Quick Diagnosis Flow, Troubleshooting Guides by Service, Routing, Common Diagnostic Commands.