Paths: File paths (shared/, references/, ../ln-*) are relative to skills repo root. If not found at CWD, locate this SKILL.md directory and go up one level for repo root.
Test Log Analyzer
Two-layer analysis of application logs. Python script handles collection and quantitative analysis; AI handles classification, quality assessment, and fix recommendations.
Inputs
No required inputs. Runs in current project directory, auto-detects log sources.
Optional args — caller instructions (natural language): time window, expected errors, test context. Example: "review logs for last 30min, auth 401 errors expected from negative tests".
Purpose & Scope
- Analyze application logs (after test runs, during development, or on demand)
- Classify errors into 4 categories: Real Bug, Test Artifact, Expected Behavior, Operational Warning
- Assess log quality: noisiness, completeness, level correctness, format, structured logging
- Map stack traces to source files; provide fix recommendations
- Report findings for quality verdict (only Real Bugs block)
- No status changes or task creation — report only
When to Use
- Analyze application logs in any project (default: last 1h)
- After test runs to classify errors and assess log quality
- Can be invoked with context instructions:
Skill(skill: "ln-514-test-log-analyzer", args: "review last 30min, 401 errors expected")
Workflow
Phase 0: Parse Instructions
If args provided — extract: time window (default: 1h), expected errors list, test context.
If no args — use defaults (last 1h, no expected errors).
Phase 1: Log Source Detection and Script Execution
MANDATORY READ: Load docs/project/infrastructure.md, docs/project/runbook.md
- Check if
scripts/analyze_test_logs.py exists in target project. If missing, copy from references/analyze_test_logs.py.
- Detect log source mode (auto-detection priority: docker → file → loki):
| Mode |
Detection |
Source |
docker |
docker compose ps returns running containers |
docker compose logs --since {window} |
file |
.log files exist, or tests/manual/results/ has output |
File paths from infrastructure.md or *.log glob |
loki |
LOKI_URL env var or tools_config.md observability section |
Loki HTTP query_range API |
- Run script:
python scripts/analyze_test_logs.py --mode {detected} [options]
- If no log sources found → return
NO_LOG_SOURCES status, skip to Phase 5.
Phase 2: 4-Category Error Classification
Classify each error group from script JSON output:
| Category |
Action |
Criteria |
| Real Bug |
Fix |
Unexpected crash, data loss, broken pipeline |
| Test Artifact |
Skip |
From test scripts, deliberate error-path validation |
| Expected Behavior |
Skip |
Rate limiting, input validation, auth failures from invalid tokens |
| Operational Warning |
Monitor |
Clock drift, resource pressure, temporary unavailability |
Test artifact detection heuristics:
- Test name contains:
invalid, error, fail, reject, unauthorized, forbidden, not_found, bad_request, timeout
- Test asserts non-2xx status codes (4xx, 5xx)
- Test uses
pytest.raises, expect(...).rejects, assertThrows, should.throw
- Errors correlate with test execution timestamps from regression test output
- Patterns matching
tests/manual/ scripts
Error taxonomy per references/error_taxonomy.md (9 categories: CRASH, TIMEOUT, AUTH, DB, NETWORK, VALIDATION, CONFIG, RESOURCE, DEPRECATION).
Phase 3: Log Quality Assessment
MANDATORY READ: Load references/error_taxonomy.md (per-level criteria table + level correctness reference)
Step 1: Detect configured log level. Check in order:
LOG_LEVEL / LOGLEVEL env var (.env, docker-compose.yml, infrastructure.md)
- Framework config: Python
logging.conf / Django LOGGING / Node LOG_LEVEL
- Default: assume
INFO if not detected
Configured level determines WHICH levels appear in logs, but each level has its own noise threshold regardless.
Step 2: Assess 6 quality dimensions:
| Dimension |
What to Check |
Signal |
| Noisiness |
Per-level noise thresholds from error_taxonomy.md section 4: TRACE (zero in prod), DEBUG (>50% monopoly), INFO (>30%), WARNING (>1% of total), ERROR (>0.1% of total) |
NOISY: {level} template "{msg}" at {ratio}% |
| Completeness & Traceability |
Critical operations missing log entries + traceability gaps (see table below) |
MISSING: No log for {operation} / TRACEABILITY_GAP: {type} in {file}:{line} |
| Level correctness |
Per-level criteria from error_taxonomy.md section 4: content, anti-patterns, library rule |
WRONG_LEVEL: should be {level} |
| Structured logging |
Missing trace_id/request_id/user context; unstructured plaintext |
UNSTRUCTURED: lacks {field} |
| Sensitivity |
PII/secrets/tokens/passwords in log messages |
SENSITIVE: {type} exposure |
| Context richness |
Errors without actionable context (order_id, user_id, operation) |
LOW_CONTEXT: lacks context |
Traceability gap detection — scan source code for operations without INFO-level logging:
| Operation Type |
Expected Log |
Where to Add |
| Incoming request handling |
Request received + response status |
Entry/exit of route handler |
| External API call |
Request sent + response status + duration |
Before/after HTTP client call |
| DB write (INSERT/UPDATE/DELETE) |
Operation + affected entity + count |
Before/after ORM/query call |
| Auth decision |
Result (allow/deny) + reason |
After auth check |
| State transition |
Old state → new state + trigger |
At transition point |
| Background job |
Start + complete/fail + duration |
Entry/exit of job handler |
| File/resource operation |
Open/close + path + size |
At I/O operation |
Log Format Quality (10-criterion checklist per references/log_analysis_output_format.md):
| # |
Criterion |
Check |
| 1 |
Dual format |
JSON in prod, readable in dev |
| 2 |
Timestamp |
Consistent, timezone-aware |
| 3 |
Level field |
Present, uppercase |
| 4 |
Trace/Correlation ID |
Present in every entry, async-safe |
| 5 |
Service name |
Identifies source service |
| 6 |
Source location |
module:line + function |
| 7 |
Extra context |
Structured fields, not string interpolation |
| 8 |
PII redaction |
Passwords, API keys, emails handled |
| 9 |
Noise suppression |
Duplicate filters, third-party suppressed |
| 10 |
Parseability |
Dev: pipe-delimited; prod: valid JSON per line |
Score: passed criteria / 10.
Phase 4: Stack Trace Mapping + Fix Recommendations
For each Real Bug:
- Extract stack trace frames; identify origin frame (first frame in project code, not in node_modules/site-packages)
- Map to source file:line
- Generate fix recommendation: what to change, where, effort estimate (S/M/L)
Prioritize using Sentry-inspired dimensions:
- High-volume (occurrence count), Post-test regression (new errors), High-impact path (auth/payment/DB), Correlated traces (trace_id across services)
Phase 5: Generate Report
MANDATORY READ: Load references/log_analysis_output_format.md
Output report to chat with header ## Test Log Analysis. Include:
- Signals table (Real Bugs count, Test Artifacts filtered, Log Noise status, Log Format score, Log Quality score)
- Real Bugs table (priority, category, error, source, fix recommendation)
- Filtered table (category, count, examples)
- Log Quality Issues table (dimension, service, issue, recommendation)
- Noise Report table (count, ratio, service, level, template, action)
- Machine-readable block
<!-- LOG-ANALYSIS-DATA ... --> for programmatic consumption
Phase 6: Meta-Analysis
MANDATORY READ: Load shared/references/meta_analysis_protocol.md
Skill type: execution-worker. Run after all phases complete.
Verdict Contribution
Quality coordinator normalization matrix component:
| Status |
Maps To |
Penalty |
| CLEAN |
-- |
0 |
| WARNINGS_ONLY |
-- |
0 |
| REAL_BUGS_FOUND |
FAIL |
-20 |
| SKIPPED / NO_LOG_SOURCES |
ignored |
0 |
Log quality/format issues are INFORMATIONAL — do not affect quality verdict. Only Real Bugs block.
Critical Rules
- No status changes or task creation; report only.
- Test Artifacts and Expected Behavior are ALWAYS filtered — never count as bugs.
- Log quality issues are advisory — inform, don't block.
- Script must handle gracefully: no Docker, no log files, no Loki →
NO_LOG_SOURCES.
- Language preservation in comments (EN/RU).
Definition of Done
- Script deployed to target project
scripts/ (or already exists).
- Log source detected and script executed (or NO_LOG_SOURCES returned).
- Errors classified into 4 categories; Real Bugs identified.
- Log quality assessed (6 dimensions + 10-criterion format checklist).
- Stack traces mapped to source files for Real Bugs.
- Report output to chat with signals table + machine-readable block.
Reference Files
- Error taxonomy:
references/error_taxonomy.md
- Output format:
references/log_analysis_output_format.md
- Analysis script:
references/analyze_test_logs.py
Version: 1.0.0
Last Updated: 2026-03-13
1---2name: ln-514-test-log-analyzer3description: Analyzes application logs: classifies errors, checks log quality/format, maps stack traces to source, recommends fixes.4license: MIT5---67> **Paths:** File paths (`shared/`, `references/`, `../ln-*`) are relative to skills repo root. If not found at CWD, locate this SKILL.md directory and go up one level for repo root.89# Test Log Analyzer1011Two-layer analysis of application logs. Python script handles collection and quantitative analysis; AI handles classification, quality assessment, and fix recommendations.1213## Inputs1415No required inputs. Runs in current project directory, auto-detects log sources.1617Optional `args` — caller instructions (natural language): time window, expected errors, test context. Example: `"review logs for last 30min, auth 401 errors expected from negative tests"`.1819## Purpose & Scope20- Analyze application logs (after test runs, during development, or on demand)21- Classify errors into 4 categories: Real Bug, Test Artifact, Expected Behavior, Operational Warning22- Assess log quality: noisiness, completeness, level correctness, format, structured logging23- Map stack traces to source files; provide fix recommendations24- Report findings for quality verdict (only Real Bugs block)25- **No status changes or task creation** — report only2627## When to Use28- Analyze application logs in any project (default: last 1h)29- After test runs to classify errors and assess log quality30- Can be invoked with context instructions: `Skill(skill: "ln-514-test-log-analyzer", args: "review last 30min, 401 errors expected")`3132## Workflow3334### Phase 0: Parse Instructions3536If `args` provided — extract: time window (default: 1h), expected errors list, test context.37If no `args` — use defaults (last 1h, no expected errors).3839### Phase 1: Log Source Detection and Script Execution4041**MANDATORY READ:** Load `docs/project/infrastructure.md`, `docs/project/runbook.md`42431) Check if `scripts/analyze_test_logs.py` exists in target project. If missing, copy from `references/analyze_test_logs.py`.442) Detect log source mode (auto-detection priority: docker → file → loki):4546| Mode | Detection | Source |47|------|-----------|--------|48| `docker` | `docker compose ps` returns running containers | `docker compose logs --since {window}` |49| `file` | `.log` files exist, or `tests/manual/results/` has output | File paths from infrastructure.md or `*.log` glob |50| `loki` | `LOKI_URL` env var or `tools_config.md` observability section | Loki HTTP query_range API |51523) Run script: `python scripts/analyze_test_logs.py --mode {detected} [options]`534) If no log sources found → return `NO_LOG_SOURCES` status, skip to Phase 5.5455### Phase 2: 4-Category Error Classification5657Classify each error group from script JSON output:5859| Category | Action | Criteria |60|----------|--------|----------|61| **Real Bug** | Fix | Unexpected crash, data loss, broken pipeline |62| **Test Artifact** | Skip | From test scripts, deliberate error-path validation |63| **Expected Behavior** | Skip | Rate limiting, input validation, auth failures from invalid tokens |64| **Operational Warning** | Monitor | Clock drift, resource pressure, temporary unavailability |6566**Test artifact detection heuristics:**67- Test name contains: `invalid`, `error`, `fail`, `reject`, `unauthorized`, `forbidden`, `not_found`, `bad_request`, `timeout`68- Test asserts non-2xx status codes (4xx, 5xx)69- Test uses `pytest.raises`, `expect(...).rejects`, `assertThrows`, `should.throw`70- Errors correlate with test execution timestamps from regression test output71- Patterns matching `tests/manual/` scripts7273**Error taxonomy per** `references/error_taxonomy.md` **(9 categories: CRASH, TIMEOUT, AUTH, DB, NETWORK, VALIDATION, CONFIG, RESOURCE, DEPRECATION).**7475### Phase 3: Log Quality Assessment7677**MANDATORY READ:** Load `references/error_taxonomy.md` (per-level criteria table + level correctness reference)7879**Step 1: Detect configured log level.** Check in order:801. `LOG_LEVEL` / `LOGLEVEL` env var (`.env`, `docker-compose.yml`, `infrastructure.md`)812. Framework config: Python `logging.conf` / Django `LOGGING` / Node `LOG_LEVEL`823. Default: assume `INFO` if not detected8384Configured level determines WHICH levels appear in logs, but each level has its own noise threshold regardless.8586**Step 2: Assess 6 quality dimensions:**8788| Dimension | What to Check | Signal |89|-----------|---------------|--------|90| **Noisiness** | Per-level noise thresholds from `error_taxonomy.md` section 4: TRACE (zero in prod), DEBUG (>50% monopoly), INFO (>30%), WARNING (>1% of total), ERROR (>0.1% of total) | `NOISY: {level} template "{msg}" at {ratio}%` |91| **Completeness & Traceability** | Critical operations missing log entries + traceability gaps (see table below) | `MISSING: No log for {operation}` / `TRACEABILITY_GAP: {type} in {file}:{line}` |92| **Level correctness** | Per-level criteria from `error_taxonomy.md` section 4: content, anti-patterns, library rule | `WRONG_LEVEL: should be {level}` |93| **Structured logging** | Missing trace_id/request_id/user context; unstructured plaintext | `UNSTRUCTURED: lacks {field}` |94| **Sensitivity** | PII/secrets/tokens/passwords in log messages | `SENSITIVE: {type} exposure` |95| **Context richness** | Errors without actionable context (order_id, user_id, operation) | `LOW_CONTEXT: lacks context` |9697**Traceability gap detection** — scan source code for operations without INFO-level logging:9899| Operation Type | Expected Log | Where to Add |100|---------------|-------------|--------------|101| Incoming request handling | Request received + response status | Entry/exit of route handler |102| External API call | Request sent + response status + duration | Before/after HTTP client call |103| DB write (INSERT/UPDATE/DELETE) | Operation + affected entity + count | Before/after ORM/query call |104| Auth decision | Result (allow/deny) + reason | After auth check |105| State transition | Old state → new state + trigger | At transition point |106| Background job | Start + complete/fail + duration | Entry/exit of job handler |107| File/resource operation | Open/close + path + size | At I/O operation |108109**Log Format Quality** (10-criterion checklist per `references/log_analysis_output_format.md`):110111| # | Criterion | Check |112|---|-----------|-------|113| 1 | Dual format | JSON in prod, readable in dev |114| 2 | Timestamp | Consistent, timezone-aware |115| 3 | Level field | Present, uppercase |116| 4 | Trace/Correlation ID | Present in every entry, async-safe |117| 5 | Service name | Identifies source service |118| 6 | Source location | module:line + function |119| 7 | Extra context | Structured fields, not string interpolation |120| 8 | PII redaction | Passwords, API keys, emails handled |121| 9 | Noise suppression | Duplicate filters, third-party suppressed |122| 10 | Parseability | Dev: pipe-delimited; prod: valid JSON per line |123124Score: passed criteria / 10.125126### Phase 4: Stack Trace Mapping + Fix Recommendations127128For each Real Bug:1291) Extract stack trace frames; identify origin frame (first frame in project code, not in node_modules/site-packages)1302) Map to source file:line1313) Generate fix recommendation: what to change, where, effort estimate (S/M/L)132133**Prioritize using Sentry-inspired dimensions:**134- High-volume (occurrence count), Post-test regression (new errors), High-impact path (auth/payment/DB), Correlated traces (trace_id across services)135136### Phase 5: Generate Report137138**MANDATORY READ:** Load `references/log_analysis_output_format.md`139140Output report to chat with header `## Test Log Analysis`. Include:141- Signals table (Real Bugs count, Test Artifacts filtered, Log Noise status, Log Format score, Log Quality score)142- Real Bugs table (priority, category, error, source, fix recommendation)143- Filtered table (category, count, examples)144- Log Quality Issues table (dimension, service, issue, recommendation)145- Noise Report table (count, ratio, service, level, template, action)146- Machine-readable block `<!-- LOG-ANALYSIS-DATA ... -->` for programmatic consumption147148### Phase 6: Meta-Analysis149150**MANDATORY READ:** Load `shared/references/meta_analysis_protocol.md`151152Skill type: `execution-worker`. Run after all phases complete.153154## Verdict Contribution155156Quality coordinator normalization matrix component:157158| Status | Maps To | Penalty |159|--------|---------|---------|160| CLEAN | -- | 0 |161| WARNINGS_ONLY | -- | 0 |162| REAL_BUGS_FOUND | FAIL | -20 |163| SKIPPED / NO_LOG_SOURCES | ignored | 0 |164165Log quality/format issues are INFORMATIONAL — do not affect quality verdict. Only Real Bugs block.166167## Critical Rules168- No status changes or task creation; report only.169- Test Artifacts and Expected Behavior are ALWAYS filtered — never count as bugs.170- Log quality issues are advisory — inform, don't block.171- Script must handle gracefully: no Docker, no log files, no Loki → `NO_LOG_SOURCES`.172- Language preservation in comments (EN/RU).173174## Definition of Done175- Script deployed to target project `scripts/` (or already exists).176- Log source detected and script executed (or NO_LOG_SOURCES returned).177- Errors classified into 4 categories; Real Bugs identified.178- Log quality assessed (6 dimensions + 10-criterion format checklist).179- Stack traces mapped to source files for Real Bugs.180- Report output to chat with signals table + machine-readable block.181182## Reference Files183- **Error taxonomy:** `references/error_taxonomy.md`184- **Output format:** `references/log_analysis_output_format.md`185- **Analysis script:** `references/analyze_test_logs.py`186187---188**Version:** 1.0.0189**Last Updated:** 2026-03-13