gh-logs
CI log analyst. Fetches GitHub Actions logs via gh, reasons about them, classifies the failure, and suggests a fix with a verify step. The terminal is faster than the web UI and won't crash on 50MB of test output.
This skill runs in a forked context so that megabytes of raw log never land in the main conversation — only your finished diagnosis returns. That also means you have no conversation history: everything you need is below, plus $ARGUMENTS.
Request: $ARGUMENTS — if that is blank or still shows the literal placeholder, no arguments were passed: run the default diagnose flow against the most recent failure.
Preflight
- Repo: !
gh repo view --json nameWithOwner --jq .nameWithOwner - Branch: !
git branch --show-current - Auth: !
gh auth status - Recent failures (all branches): !
gh run list --status failure --limit 5 --json databaseId,displayTitle,workflowName,headBranch,createdAt
If auth failed above, tell the user to run gh auth login and stop. If the repo lookup failed, ask which repo to target. Otherwise prefer a failure on the current branch; fall back to the list above.
Reference files live in ${CLAUDE_SKILL_DIR}/references/.
Modes
| Mode | Triggered by | Primary reference |
|---|---|---|
| Diagnose (default) | "CI is broken / red / failing", "why did this run fail?", a bare run ID or URL | references/failure-patterns.md |
--flaky |
"is this test flaky?", "it passes locally" | references/failure-patterns.md (test signatures) |
--slow |
"why is the build so slow?" | references/gh-commands.md (timing jq) |
--history [n] |
"has this been failing for a while?" (default 10) | references/analysis-templates.md |
--watch |
"watch this run and tell me when it's done" | references/gh-commands.md (gh run watch flags) |
Every mode reports through a template in references/analysis-templates.md. A <workflow-name> argument narrows any mode to one workflow. Full gh invocation cookbook: references/gh-commands.md.
When NOT to use
- Viewing passing-run logs for debugging successful runs — use
gh run view <id> --logdirectly. - Editing workflow YAML — this skill reads runs, it doesn't author workflows.
- Running or re-running workflows — that's
gh workflow run/gh run rerun.
Default workflow (diagnose)
1. Find the run
The Preflight block already resolved repo, branch, and recent failures. If a run ID was passed, use it directly. Otherwise narrow to the current branch:
BRANCH=$(git branch --show-current)
gh run list --branch "$BRANCH" --status failure --limit 1 \
--json databaseId,displayTitle,conclusion,event,headBranch,workflowName,createdAt
If that is empty, fall back to the all-branches list from Preflight. With a workflow name — add --workflow <name>.
2. Get the overview
gh run view <run-id> --json jobs \
--jq '.jobs[] | {name, conclusion, steps: [.steps[] | select(.conclusion == "failure") | {name, conclusion}]}'
This tells you which jobs failed and which steps inside them.
3. Fetch the failing logs
gh run view <run-id> --log-failed > /tmp/run-<run-id>.log
${CLAUDE_SKILL_DIR}/scripts/classify-log.sh /tmp/run-<run-id>.log
Run the classifier before reading the log by hand: failed-step output is routinely 5000+ lines, and the script greps it against one signature regex per category, printing the hit count and the first five matching lines (with line numbers) for each, then primary: <category> — the most upstream category that matched. That narrows a wall of text to the lines that matter and tells you where to sed -n next.
If the log is still unwieldy, narrow to one job: gh run view <run-id> --job <job-id> --log-failed. If even that is too large, gh api the raw log and grep for error/FAIL/fatal.
4. Classify
Start from the classifier's primary: line, then confirm against the surrounding log and references/failure-patterns.md — the script only sees signatures, not causality. Pick one primary category:
| Category | Strongest signals |
|---|---|
| test | FAIL, AssertionError, --- FAIL:, snapshot mismatch |
| build | error TS, Build failed, Rollup failed to resolve, undefined: |
| deps | ERESOLVE, 404 Not Found, ETARGET, ECONNREFUSED to registry |
| lint | X errors found, biome/eslint/prettier output |
| auth | 403, 401, Permission denied (publickey), missing secret |
| infra | Killed (137), No space left, heap out of memory, runner shutdown |
| timeout | exceeded maximum execution time, stuck for 10+ min |
When multiple categories match (common — e.g., an OOM during tests looks like both infra and test): pick the most upstream cause, because that's what needs to be fixed. Priority: auth > deps > build > infra > lint > test > timeout. A test failing because deps didn't install is a deps bug, not a test bug. When it's genuinely ambiguous, surface both and ask the user which feels right — a wrong classification leads to a wrong fix.
5. Report
Use the diagnosis template in references/analysis-templates.md. Always include:
- Category and the failed step / job name
- Root cause — 1–2 sentences, specific
- Log excerpt — the lines that proved it, truncated if long
- Suggested fix — actionable, with commands or code
- Verify command — how to re-run and confirm the fix worked
Other modes
Flaky test detection — --flaky
gh run list --branch "$BRANCH" --limit 20 --json databaseId,conclusion
For each failed run, extract failed test names from the log. Tests that fail in some runs but pass in others are flaky. Report with pass/fail ratio and the suspected mechanism (race condition, timing dependency, shared port, external service, test ordering). See references/analysis-templates.md for the output shape.
Slow step profiling — --slow
gh run view --json jobs returns startedAt / completedAt on every step (gh 2.60+, camelCase):
gh run view <run-id> --json jobs --jq '
[.jobs[].steps[] | select(.completedAt != null and .startedAt != null) |
{name, duration: ((.completedAt | fromdateiso8601) - (.startedAt | fromdateiso8601))}] |
sort_by(-.duration) | .[] | "\(.duration)s\t\(.name)"'
Durations come back sorted descending — identify bottlenecks and suggest cache, parallelism, or dropping the step. On gh older than 2.60 the steps carry no timestamps; use the REST fallback in references/gh-commands.md (same filter, snake_case fields).
History — --history [n]
gh run list --status failure --limit <n> \
--json databaseId,displayTitle,conclusion,event,headBranch,workflowName,createdAt
For each, pull failed job/step names. Surface recurring patterns: same step, specific branch, time-of-day correlation.
Watch — --watch
gh run watch --exit-status
When it finishes, if it failed, drop into diagnose mode on the resulting run ID.
Example — auto-diagnose
User: /gh-logs
Claude: Checking CI for branch feat/auth-flow...
Found failed run #4521 (CI / test) from 3 minutes ago.
## Diagnosis
Category: test
Failed step: Run tests (job: test-ubuntu)
Root cause: Snapshot mismatch in LoginForm — expected output changed after adding the "Remember me" checkbox.
Log excerpt:
FAIL src/components/LoginForm.test.tsx
- renders login form (2ms)
Expected: "<form>..."
Received: "<form>...<label>Remember me</label>..."
1 snapshot failed.
Fix:
<test-runner> -u src/components/LoginForm.test.tsx # e.g. vitest -u / jest -u, via the project's package manager
Verify:
Commit the updated snapshot (hand off to /commit), push, then `gh run watch`.
More session shapes (flaky / slow / history) are in references/analysis-templates.md.
Edge cases
| Situation | Handling |
|---|---|
| No failures found | Report "no failed runs on " and suggest widening (different branch, include success, workflow filter). |
gh rate limit (403) |
Back off, tell the user which call hit the limit. |
| Logs >5000 lines | Narrow to failing job, then grep for error/FAIL/fatal if still too large. |
No gh installed |
brew install gh or https://cli.github.com. |
| Not authenticated | gh auth login. |
| Private repo / no access | gh returns 404; explain required permissions. |
| Multiple failed jobs | Diagnose each; lead the report with the most upstream cause. |
| Cancelled runs | Infra category, unless the log carries a timeout signature (exceeded the maximum execution time, The operation was canceled after a long silent step) — then it's timeout. Check whether cancellation was manual, concurrency, or timeout. |
Reference index
references/failure-patterns.md— log signature database by language / categoryreferences/analysis-templates.md— output templates for each modereferences/gh-commands.md— completeghCLI command cookbook with jq filters