Sibling skills (local only)
Sibling CloudBase skills ship beside this skill. Use local relative paths such as ../auth-tool-cloudbase/SKILL.md.
If a referenced sibling skill file is missing from this environment, ask the user to install the full CloudBase plugin (or the missing skill). Do not HTTP-fetch remote skill or protocol markdown into the agent context.
Activation Contract
Use this first when
- The user wants to check the health or status of CloudBase resources (cloud functions, CloudRun, databases, storage, etc.).
- The user reports errors, failures, or abnormal behavior and wants a quick diagnosis.
- The user asks for an "inspection", "health check", "巡检", "诊断", or "troubleshooting" of their CloudBase environment.
- The user wants to review recent error logs across services.
Read before writing code if
- The inspection reveals code-level issues in cloud functions or CloudRun services — then read the relevant implementation skill before suggesting fixes.
- The user wants to fix a problem found during inspection rather than just diagnose it.
Then also read
- Cloud function issues ->
../cloud-functions/SKILL.md
- CloudRun issues ->
../cloudrun-development/SKILL.md
- Database issues ->
../postgresql-development-cloudbase/SKILL.md for CloudBase PG / PostgreSQL, ../relational-database-mcp-cloudbase/SKILL.md for MySQL, or ../cloudbase-document-database-web-sdk/SKILL.md for NoSQL
- Platform overview ->
../cloudbase-platform/SKILL.md
Do NOT use for
- Deploying new resources or writing application code. This skill is read-only and diagnostic.
- Replacing proper monitoring/alerting infrastructure. It provides point-in-time inspection, not continuous monitoring.
- Directly fixing problems — it diagnoses and recommends; actual fixes should use the appropriate implementation skill.
Common mistakes / gotchas
- Running a full inspection without first confirming the environment is bound (
auth tool must show logged-in and env-bound state).
- Ignoring CLS log service status — if CLS is not enabled,
queryLogs will fail; always check first with queryLogs(action="checkLogService").
- Searching logs without a time range — this can return excessive or irrelevant results. Always scope searches to a relevant time window.
- Treating a single error log as the root cause without correlating across resources. A function error may stem from a database or config issue.
Minimal checklist
How to use this skill (for a coding agent)
Inspection Modes
The skill supports two modes based on user intent:
| Mode |
When to use |
Scope |
| Full inspection |
User asks for a general health check / 巡检 / 全面检查 |
All resource types in the environment |
| Targeted inspection |
User reports a specific error or asks about a specific resource |
One resource type or a specific resource |
Full Inspection Workflow
Follow these steps in order for a comprehensive environment health check:
Step 1 — Environment Check
envQuery(action="info")
Confirm the environment is accessible. Record the envId for console link generation.
Step 2 — Log Service Status
queryLogs(action="checkLogService")
If CLS is not enabled, note this as a warning — log-based diagnosis will be unavailable. Recommend enabling CLS in the console: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log
Step 3 — Cloud Functions Inspection
queryFunctions(action="listFunctions")
For each function, check:
- Status: Is the function in an active/deployed state?
- Recent errors:
queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<recent>")
- Common issues:
- Timeout errors (execution exceeded limit)
- Memory limit exceeded
- Runtime errors (unhandled exceptions)
- Cold start frequency
Step 4 — CloudRun Services Inspection
queryCloudRun(action="list")
For each service, check:
- Status: Is the service running?
- Detail:
queryCloudRun(action="detail", detailServerName="<name>")
- Common issues:
- Service not running (scaled to zero or crashed)
- Image pull failures
- OOMKilled events
- Health check failures
Step 5 — Error Log Aggregation (if CLS is enabled)
queryLogs(action="searchLogs", queryString="ERROR", service="tcb", startTime="<24h-ago>", limit=50)
queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", startTime="<24h-ago>", limit=50)
Look for patterns:
- Repeated error messages (same error many times)
- Cascading failures (errors in multiple services around the same time)
- Timeout patterns
Step 6 — Summary Report
Generate a structured report:
# CloudBase Resource Inspection Report
**Environment**: ${envId}
**Inspection Time**: ${timestamp}
## Overall Health: ✅ Healthy / ⚠️ Warnings Found / ❌ Issues Found
### Cloud Functions
| Function | Status | Recent Errors | Severity |
|----------|--------|---------------|----------|
| ... | ... | ... | ... |
### CloudRun Services
| Service | Status | Issues | Severity |
|---------|--------|--------|----------|
| ... | ... | ... | ... |
### Error Log Summary
- Total errors in last 24h: N
- Top error patterns: ...
## Recommendations
1. ...
2. ...
## Console Links
- Cloud Functions: https://tcb.cloud.tencent.com/dev?envId=${envId}#/scf
- CloudRun: https://tcb.cloud.tencent.com/dev?envId=${envId}#/platform-run
- Logs: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log
Targeted Inspection Workflow
When the user specifies a resource type or a specific resource:
- Cloud function errors:
queryFunctions(action="listFunctionLogs", functionName="<name>") then queryLogs(action="searchLogs", queryString="* AND functionName:<name> AND level:ERROR", ...)
- CloudRun errors:
queryCloudRun(action="detail", detailServerName="<name>") then queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", ...)
- If logs show DB / Redis connection failures (
ECONNREFUSED, timeout, "could not connect"): check whether VpcConf is set and matches the database VPC. See cloudrun-development/references/vpc-and-database.md.
- Database issues: Check
queryPgDatabase(action="context"|"metadata"|"objects") for CloudBase PG, queryMysqlDatabase for MySQL, or readNoSqlDatabaseStructure for NoSQL depending on type
- General error search:
queryLogs(action="searchLogs", queryString="<error-keyword>", ...)
AIOps Methodology
This skill follows AIOps principles for intelligent inspection:
- Data Collection: Gather logs and resource states via MCP tools
- Pattern Recognition: Identify recurring errors, anomaly patterns, and correlations across services
- Root Cause Hypothesis: Based on error patterns, suggest likely root causes (e.g., a function timeout may be caused by a database query bottleneck)
- Actionable Recommendations: Provide specific, prioritized remediation steps with links to relevant skills and console pages
Severity Levels
| Level |
Icon |
Meaning |
| Critical |
❌ |
Service is down or data is at risk; requires immediate action |
| Warning |
⚠️ |
Errors detected but service is still partially functional; investigate soon |
| Info |
ℹ️ |
No errors found; informational status only |
| Healthy |
✅ |
Resource is operating normally |
Preferred Tool Map
| Operation |
MCP Tool Call |
| Check environment |
envQuery(action="info") |
| Check CLS status |
queryLogs(action="checkLogService") |
| List cloud functions |
queryFunctions(action="listFunctions") |
| Get function detail |
queryFunctions(action="getFunctionDetail", functionName="<name>") |
| Get function logs |
queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<time>", endTime="<time>") |
| Get function log detail |
queryFunctions(action="getFunctionLogDetail", requestId="<id>") |
| List CloudRun services |
queryCloudRun(action="list") |
| Get CloudRun detail |
queryCloudRun(action="detail", detailServerName="<name>") |
| Search CLS logs |
queryLogs(action="searchLogs", queryString="<query>", service="tcb|tcbr", startTime="<time>", endTime="<time>") |
| Check NoSQL structure |
readNoSqlDatabaseStructure(action="listCollections") |
| Check PostgreSQL context |
queryPgDatabase(action="context") |
| Check PostgreSQL metadata |
queryPgDatabase(action="metadata", limit=20) |
| Check MySQL status |
queryMysqlDatabase(action="getContext") |
Common CLS Query Patterns
| Scenario |
queryString |
| All errors |
ERROR |
| Function timeout |
timeout OR 超时 |
| Function OOM |
OOM OR out of memory OR 内存超限 |
| CloudRun crash |
crash OR OOMKilled OR Error |
| Specific function errors |
functionName:<name> AND level:ERROR |
| 5xx HTTP errors |
statusCode:>499 |
| Cold start issues |
coldStart OR 冷启动 |
Time Range Guidance
- Quick check: Last 1 hour (
startTime = 1 hour ago)
- Standard inspection: Last 24 hours
- Trend analysis: Last 7 days
- Specific incident: Narrow to the reported time window
Always use ISO 8601 format for startTime/endTime, e.g., "2025-01-15 00:00:00".
Related Skills
cloud-functions — Cloud function development, deployment, and debugging
cloudrun-development — CloudRun backend deployment and management
cloudbase-platform — General platform knowledge and console navigation
postgresql-development-cloudbase — CloudBase PostgreSQL / PG diagnostics and schema/RLS checks
relational-database-mcp-cloudbase — MySQL database management and diagnostics
1---2name: ops-inspector3description: AIOps-style one-click inspection skill for CloudBase resources. Use this skill when users need to diagnose errors, check resource health, inspect logs, or run a comprehensive health check across cloud functions, CloudRun services, databases, and other CloudBase resources.4---5
6## Sibling skills (local only)
7
8Sibling CloudBase skills ship beside this skill. Use local relative paths such as `../auth-tool-cloudbase/SKILL.md`.
9
10If a referenced sibling skill file is missing from this environment, ask the user to install the full CloudBase plugin (or the missing skill). Do **not** HTTP-fetch remote skill or protocol markdown into the agent context.
11
12## Activation Contract
13
14### Use this first when
15
16- The user wants to check the health or status of CloudBase resources (cloud functions, CloudRun, databases, storage, etc.).
17- The user reports errors, failures, or abnormal behavior and wants a quick diagnosis.
18- The user asks for an "inspection", "health check", "巡检", "诊断", or "troubleshooting" of their CloudBase environment.
19- The user wants to review recent error logs across services.
20
21### Read before writing code if
22
23- The inspection reveals code-level issues in cloud functions or CloudRun services — then read the relevant implementation skill before suggesting fixes.
24- The user wants to fix a problem found during inspection rather than just diagnose it.
25
26### Then also read
27
28- Cloud function issues -> `../cloud-functions/SKILL.md`
29- CloudRun issues -> `../cloudrun-development/SKILL.md`
30- Database issues -> `../postgresql-development-cloudbase/SKILL.md` for CloudBase PG / PostgreSQL, `../relational-database-mcp-cloudbase/SKILL.md` for MySQL, or `../cloudbase-document-database-web-sdk/SKILL.md` for NoSQL
31- Platform overview -> `../cloudbase-platform/SKILL.md`
32
33### Do NOT use for
34
35- Deploying new resources or writing application code. This skill is read-only and diagnostic.
36- Replacing proper monitoring/alerting infrastructure. It provides point-in-time inspection, not continuous monitoring.
37- Directly fixing problems — it diagnoses and recommends; actual fixes should use the appropriate implementation skill.
38
39### Common mistakes / gotchas
40
41- Running a full inspection without first confirming the environment is bound (`auth` tool must show logged-in and env-bound state).
42- Ignoring CLS log service status — if CLS is not enabled, `queryLogs` will fail; always check first with `queryLogs(action="checkLogService")`.
43- Searching logs without a time range — this can return excessive or irrelevant results. Always scope searches to a relevant time window.
44- Treating a single error log as the root cause without correlating across resources. A function error may stem from a database or config issue.
45
46### Minimal checklist
47
48- [ ] Environment is bound and accessible (`envQuery(action="info")`)
49- [ ] CLS log service is enabled (`queryLogs(action="checkLogService")`)
50- [ ] All target resources are listed before diving into details
51- [ ] Time range is specified for any log searches
52- [ ] Findings are summarized with severity levels and actionable recommendations
53
54---
55
56## How to use this skill (for a coding agent)
57
58### Inspection Modes
59
60The skill supports two modes based on user intent:
61
62| Mode | When to use | Scope |
63|------|-------------|-------|
64| **Full inspection** | User asks for a general health check / 巡检 / 全面检查 | All resource types in the environment |
65| **Targeted inspection** | User reports a specific error or asks about a specific resource | One resource type or a specific resource |
66
67### Full Inspection Workflow
68
69Follow these steps in order for a comprehensive environment health check:
70
71**Step 1 — Environment Check**
72
73```
74envQuery(action="info")
75```
76
77Confirm the environment is accessible. Record the `envId` for console link generation.
78
79**Step 2 — Log Service Status**
80
81```
82queryLogs(action="checkLogService")
83```
84
85If CLS is not enabled, note this as a **warning** — log-based diagnosis will be unavailable. Recommend enabling CLS in the console: `https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log`
86
87**Step 3 — Cloud Functions Inspection**
88
89```
90queryFunctions(action="listFunctions")
91```
92
93For each function, check:
94- **Status**: Is the function in an active/deployed state?
95- **Recent errors**: `queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<recent>")`
96- **Common issues**:
97 - Timeout errors (execution exceeded limit)
98 - Memory limit exceeded
99 - Runtime errors (unhandled exceptions)
100 - Cold start frequency
101
102**Step 4 — CloudRun Services Inspection**
103
104```
105queryCloudRun(action="list")
106```
107
108For each service, check:
109- **Status**: Is the service running?
110- **Detail**: `queryCloudRun(action="detail", detailServerName="<name>")`
111- **Common issues**:
112 - Service not running (scaled to zero or crashed)
113 - Image pull failures
114 - OOMKilled events
115 - Health check failures
116
117**Step 5 — Error Log Aggregation** (if CLS is enabled)
118
119```
120queryLogs(action="searchLogs", queryString="ERROR", service="tcb", startTime="<24h-ago>", limit=50)
121queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", startTime="<24h-ago>", limit=50)
122```
123
124Look for patterns:
125- Repeated error messages (same error many times)
126- Cascading failures (errors in multiple services around the same time)
127- Timeout patterns
128
129**Step 6 — Summary Report**
130
131Generate a structured report:
132
133```markdown
134# CloudBase Resource Inspection Report
135
136**Environment**: ${envId}
137**Inspection Time**: ${timestamp}
138
139## Overall Health: ✅ Healthy / ⚠️ Warnings Found / ❌ Issues Found
140
141### Cloud Functions
142| Function | Status | Recent Errors | Severity |
143|----------|--------|---------------|----------|
144| ... | ... | ... | ... |
145
146### CloudRun Services
147| Service | Status | Issues | Severity |
148|---------|--------|--------|----------|
149| ... | ... | ... | ... |
150
151### Error Log Summary
152- Total errors in last 24h: N
153- Top error patterns: ...
154
155## Recommendations
1561. ...
1572. ...
158
159## Console Links
160- Cloud Functions: https://tcb.cloud.tencent.com/dev?envId=${envId}#/scf
161- CloudRun: https://tcb.cloud.tencent.com/dev?envId=${envId}#/platform-run
162- Logs: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log
163```
164
165### Targeted Inspection Workflow
166
167When the user specifies a resource type or a specific resource:
168
1691. **Cloud function errors**: `queryFunctions(action="listFunctionLogs", functionName="<name>")` then `queryLogs(action="searchLogs", queryString="* AND functionName:<name> AND level:ERROR", ...)`
1702. **CloudRun errors**: `queryCloudRun(action="detail", detailServerName="<name>")` then `queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", ...)`
171 - If logs show DB / Redis connection failures (`ECONNREFUSED`, timeout, "could not connect"): check whether `VpcConf` is set and matches the database VPC. See `cloudrun-development/references/vpc-and-database.md`.
1723. **Database issues**: Check `queryPgDatabase(action="context"|"metadata"|"objects")` for CloudBase PG, `queryMysqlDatabase` for MySQL, or `readNoSqlDatabaseStructure` for NoSQL depending on type
1734. **General error search**: `queryLogs(action="searchLogs", queryString="<error-keyword>", ...)`
174
175### AIOps Methodology
176
177This skill follows AIOps principles for intelligent inspection:
178
1791. **Data Collection**: Gather logs and resource states via MCP tools
1802. **Pattern Recognition**: Identify recurring errors, anomaly patterns, and correlations across services
1813. **Root Cause Hypothesis**: Based on error patterns, suggest likely root causes (e.g., a function timeout may be caused by a database query bottleneck)
1824. **Actionable Recommendations**: Provide specific, prioritized remediation steps with links to relevant skills and console pages
183
184### Severity Levels
185
186| Level | Icon | Meaning |
187|-------|------|---------|
188| Critical | ❌ | Service is down or data is at risk; requires immediate action |
189| Warning | ⚠️ | Errors detected but service is still partially functional; investigate soon |
190| Info | ℹ️ | No errors found; informational status only |
191| Healthy | ✅ | Resource is operating normally |
192
193### Preferred Tool Map
194
195| Operation | MCP Tool Call |
196|-----------|---------------|
197| Check environment | `envQuery(action="info")` |
198| Check CLS status | `queryLogs(action="checkLogService")` |
199| List cloud functions | `queryFunctions(action="listFunctions")` |
200| Get function detail | `queryFunctions(action="getFunctionDetail", functionName="<name>")` |
201| Get function logs | `queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<time>", endTime="<time>")` |
202| Get function log detail | `queryFunctions(action="getFunctionLogDetail", requestId="<id>")` |
203| List CloudRun services | `queryCloudRun(action="list")` |
204| Get CloudRun detail | `queryCloudRun(action="detail", detailServerName="<name>")` |
205| Search CLS logs | `queryLogs(action="searchLogs", queryString="<query>", service="tcb\|tcbr", startTime="<time>", endTime="<time>")` |
206| Check NoSQL structure | `readNoSqlDatabaseStructure(action="listCollections")` |
207| Check PostgreSQL context | `queryPgDatabase(action="context")` |
208| Check PostgreSQL metadata | `queryPgDatabase(action="metadata", limit=20)` |
209| Check MySQL status | `queryMysqlDatabase(action="getContext")` |
210
211### Common CLS Query Patterns
212
213| Scenario | queryString |
214|----------|-------------|
215| All errors | `ERROR` |
216| Function timeout | `timeout OR 超时` |
217| Function OOM | `OOM OR out of memory OR 内存超限` |
218| CloudRun crash | `crash OR OOMKilled OR Error` |
219| Specific function errors | `functionName:<name> AND level:ERROR` |
220| 5xx HTTP errors | `statusCode:>499` |
221| Cold start issues | `coldStart OR 冷启动` |
222
223### Time Range Guidance
224
225- **Quick check**: Last 1 hour (`startTime` = 1 hour ago)
226- **Standard inspection**: Last 24 hours
227- **Trend analysis**: Last 7 days
228- **Specific incident**: Narrow to the reported time window
229
230Always use ISO 8601 format for `startTime`/`endTime`, e.g., `"2025-01-15 00:00:00"`.
231
232## Related Skills
233
234- `cloud-functions` — Cloud function development, deployment, and debugging
235- `cloudrun-development` — CloudRun backend deployment and management
236- `cloudbase-platform` — General platform knowledge and console navigation
237- `postgresql-development-cloudbase` — CloudBase PostgreSQL / PG diagnostics and schema/RLS checks
238- `relational-database-mcp-cloudbase` — MySQL database management and diagnostics