Azure Diagnostics
AUTHORITATIVE GUIDANCE — MANDATORY COMPLIANCE
This document is the official source for debugging and troubleshooting Azure production issues. Follow these instructions to diagnose and resolve common Azure service problems systematically.
Triggers
Activate this skill when user wants to:
- Debug or troubleshoot production issues
- Diagnose errors in Azure services
- Analyze application logs or metrics
- Fix image pull, cold start, or health probe issues
- Investigate why Azure resources are failing
- Find root cause of application errors
- Troubleshoot App Service issues (high CPU, deployment failures, crashes, slow responses, TLS/custom domains)
- Respond to prompts like "troubleshoot app service", "app service high CPU", or "app service deployment failure"
- Troubleshoot Azure Function Apps (invocation failures, timeouts, binding errors)
- Find the App Insights or Log Analytics workspace linked to a Function App
- Troubleshoot AKS clusters, nodes, pods, ingress, or Kubernetes networking issues
- Troubleshoot Azure VM connectivity issues (RDP/SSH failures, port 3389/22 timeouts, NSG or firewall blocking, credential resets)
- Troubleshoot Azure Messaging SDK issues (Event Hubs, Service Bus connection failures, AMQP errors, message lock issues)
Rules
- Start with systematic diagnosis flow
- Use AppLens (MCP) for AI-powered diagnostics when available
- Check resource health before deep-diving into logs
- Select appropriate troubleshooting guide based on service type
- Document findings and attempted remediation steps
- Route AKS incidents to the dedicated AKS troubleshooting document
Quick Diagnosis Flow
- Identify symptoms - What's failing?
- Check resource health - Is Azure healthy?
- Review logs - What do logs show?
- Analyze metrics - Performance patterns?
- Investigate recent changes - What changed?
Troubleshooting Guides by Service
| Service |
Common Issues |
Reference |
| Container Apps |
Image pull failures, cold starts, health probes, port mismatches |
container-apps/ |
| App Service |
High CPU, deployment failures, crashes, slow responses, TLS/custom domains |
app-service/ |
| Function Apps |
App details, invocation failures, timeouts, binding errors, cold starts, missing app settings |
functions/ |
| AKS |
Cluster access, nodes, kube-system, scheduling, crash loops, ingress, DNS, upgrades |
AKS Troubleshooting |
| Compute |
VM RDP/SSH connectivity, NSG/firewall blocks, credential resets, VM agent/tooling issues |
VM Connectivity Troubleshooting |
| Messaging |
Event Hubs & Service Bus SDK errors, AMQP failures, message lock, connectivity |
Messaging Troubleshooting |
Routing
- Keep Container Apps and Function Apps diagnostics in this parent skill.
- Route active AKS incidents, AKS-specific intake, evidence gathering, and remediation guidance to AKS Troubleshooting.
- Route Azure VM RDP/SSH connectivity, NSG/firewall, credential reset, and VM agent troubleshooting to VM Connectivity Troubleshooting.
- Route Azure Messaging SDK troubleshooting (Event Hubs, Service Bus) to Messaging Troubleshooting.
Quick Reference
Common Diagnostic Commands
# Check resource health
az resource show --ids RESOURCE_ID
# View activity log
az monitor activity-log list -g RG --max-events 20
# Container Apps logs
az containerapp logs show --name APP -g RG --follow
# Function App logs (query App Insights traces)
az monitor app-insights query --apps APP-INSIGHTS -g RG \
--analytics-query "traces | where timestamp > ago(1h) | order by timestamp desc | take 50"
AppLens (MCP Tools)
For AI-powered diagnostics, use:
mcp_azure_mcp_applens
intent: "diagnose issues with <resource-name>"
command: "diagnose"
parameters:
resourceId: "<resource-id>"
Provides:
- Automated issue detection
- Root cause analysis
- Remediation recommendations
Azure Monitor (MCP Tools)
For querying logs and metrics:
mcp_azure_mcp_monitor
intent: "query logs for <resource-name>"
command: "logs_query"
parameters:
workspaceId: "<workspace-id>"
query: "<KQL-query>"
See kql-queries.md for common diagnostic queries.
Check Azure Resource Health
Using MCP
mcp_azure_mcp_resourcehealth
intent: "check health status of <resource-name>"
command: "get"
parameters:
resourceId: "<resource-id>"
Using CLI
# Check specific resource health
az resource show --ids RESOURCE_ID
# Check recent activity
az monitor activity-log list -g RG --max-events 20
References
- KQL Query Library
- Azure Resource Graph Queries
- App Service Troubleshooting
- Function Apps Troubleshooting
- VM Connectivity Troubleshooting
- Messaging Troubleshooting
Source: microsoft/skills → .github/plugins/azure-skills/skills/azure-diagnostics/SKILL.md
1---2name: azure-diagnostics3description: Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.4---5
6
7# Azure Diagnostics
8
9> **AUTHORITATIVE GUIDANCE — MANDATORY COMPLIANCE**
10>
11> This document is the **official source** for debugging and troubleshooting Azure production issues. Follow these instructions to diagnose and resolve common Azure service problems systematically.
12
13## Triggers
14
15Activate this skill when user wants to:
16- Debug or troubleshoot production issues
17- Diagnose errors in Azure services
18- Analyze application logs or metrics
19- Fix image pull, cold start, or health probe issues
20- Investigate why Azure resources are failing
21- Find root cause of application errors
22- Troubleshoot App Service issues (high CPU, deployment failures, crashes, slow responses, TLS/custom domains)
23- Respond to prompts like "troubleshoot app service", "app service high CPU", or "app service deployment failure"
24- Troubleshoot Azure Function Apps (invocation failures, timeouts, binding errors)
25- Find the App Insights or Log Analytics workspace linked to a Function App
26- Troubleshoot AKS clusters, nodes, pods, ingress, or Kubernetes networking issues
27- Troubleshoot Azure VM connectivity issues (RDP/SSH failures, port 3389/22 timeouts, NSG or firewall blocking, credential resets)
28- Troubleshoot Azure Messaging SDK issues (Event Hubs, Service Bus connection failures, AMQP errors, message lock issues)
29
30## Rules
31
321. Start with systematic diagnosis flow
332. Use AppLens (MCP) for AI-powered diagnostics when available
343. Check resource health before deep-diving into logs
354. Select appropriate troubleshooting guide based on service type
365. Document findings and attempted remediation steps
376. Route AKS incidents to the dedicated AKS troubleshooting document
38
39---
40
41## Quick Diagnosis Flow
42
431. **Identify symptoms** - What's failing?
442. **Check resource health** - Is Azure healthy?
453. **Review logs** - What do logs show?
464. **Analyze metrics** - Performance patterns?
475. **Investigate recent changes** - What changed?
48
49---
50
51## Troubleshooting Guides by Service
52
53| Service | Common Issues | Reference |
54|---------|---------------|-----------|
55| **Container Apps** | Image pull failures, cold starts, health probes, port mismatches | [container-apps/](references/container-apps/README.md) |
56| **App Service** | High CPU, deployment failures, crashes, slow responses, TLS/custom domains | [app-service/](references/app-service/README.md) |
57| **Function Apps** | App details, invocation failures, timeouts, binding errors, cold starts, missing app settings | [functions/](references/functions/README.md) |
58| **AKS** | Cluster access, nodes, `kube-system`, scheduling, crash loops, ingress, DNS, upgrades | [AKS Troubleshooting](troubleshooting/aks/aks-troubleshooting.md) |
59| **Compute** | VM RDP/SSH connectivity, NSG/firewall blocks, credential resets, VM agent/tooling issues | [VM Connectivity Troubleshooting](troubleshooting/compute/vm-troubleshooting.md) |
60| **Messaging** | Event Hubs & Service Bus SDK errors, AMQP failures, message lock, connectivity | [Messaging Troubleshooting](troubleshooting/messaging/README.md) |
61
62---
63
64## Routing
65
66- Keep Container Apps and Function Apps diagnostics in this parent skill.
67- Route active AKS incidents, AKS-specific intake, evidence gathering, and remediation guidance to [AKS Troubleshooting](troubleshooting/aks/aks-troubleshooting.md).
68- Route Azure VM RDP/SSH connectivity, NSG/firewall, credential reset, and VM agent troubleshooting to [VM Connectivity Troubleshooting](troubleshooting/compute/vm-troubleshooting.md).
69- Route Azure Messaging SDK troubleshooting (Event Hubs, Service Bus) to [Messaging Troubleshooting](troubleshooting/messaging/README.md).
70
71---
72
73## Quick Reference
74
75### Common Diagnostic Commands
76
77```bash
78# Check resource health
79az resource show --ids RESOURCE_ID
80# View activity log
81az monitor activity-log list -g RG --max-events 20
82# Container Apps logs
83az containerapp logs show --name APP -g RG --follow
84# Function App logs (query App Insights traces)
85az monitor app-insights query --apps APP-INSIGHTS -g RG \
86 --analytics-query "traces | where timestamp > ago(1h) | order by timestamp desc | take 50"
87```
88
89### AppLens (MCP Tools)
90
91For AI-powered diagnostics, use:
92```
93mcp_azure_mcp_applens
94 intent: "diagnose issues with <resource-name>"
95 command: "diagnose"
96 parameters:
97 resourceId: "<resource-id>"
98
99Provides:
100- Automated issue detection
101- Root cause analysis
102- Remediation recommendations
103```
104
105### Azure Monitor (MCP Tools)
106
107For querying logs and metrics:
108```
109mcp_azure_mcp_monitor
110 intent: "query logs for <resource-name>"
111 command: "logs_query"
112 parameters:
113 workspaceId: "<workspace-id>"
114 query: "<KQL-query>"
115```
116
117See [kql-queries.md](references/kql-queries.md) for common diagnostic queries.
118
119---
120
121## Check Azure Resource Health
122
123### Using MCP
124
125```
126mcp_azure_mcp_resourcehealth
127 intent: "check health status of <resource-name>"
128 command: "get"
129 parameters:
130 resourceId: "<resource-id>"
131```
132
133### Using CLI
134
135```bash
136# Check specific resource health
137az resource show --ids RESOURCE_ID
138
139# Check recent activity
140az monitor activity-log list -g RG --max-events 20
141```
142
143---
144
145## References
146
147- [KQL Query Library](references/kql-queries.md)
148- [Azure Resource Graph Queries](references/azure-resource-graph.md)
149- [App Service Troubleshooting](references/app-service/README.md)
150- [Function Apps Troubleshooting](references/functions/README.md)
151- [VM Connectivity Troubleshooting](troubleshooting/compute/vm-troubleshooting.md)
152- [Messaging Troubleshooting](troubleshooting/messaging/README.md)
153
154---
155
156**Source:** [`microsoft/skills`](https://github.com/microsoft/skills) → `.github/plugins/azure-skills/skills/azure-diagnostics/SKILL.md`