oncall-validation-debugger
Execution Mode — FULLY AUTONOMOUS
NEVER ask the user for confirmation. Fetch logs, interpret, fix immediately. No user-facing handoffs (rollout triggering belongs to oncall-iteration-loop). If you can't map a pattern with confidence, escalate to horizon-agentic-reviewer.
Trigger
Oracle ≠ 1.0/passed; No-Op score > 0; grader exception; feedback: null (setup crash); or before editing setup/solution/grader after a validation failure. NOT for rollout failures — use oncall-score-tuner.
Task Identity
| Field |
Value |
| UUID |
<TASK_ID> |
| Slug / path |
oncall-schedule-coverage-collapse / tasks/oncall-schedule-coverage-collapse/ |
| Subscores |
coverage_restored (0.5) + monitoring_governance (0.5) |
| Venv |
source /Users/mac/Documents/tasks/horizon_env/bin/activate |
Root-only files the grader reads: /opt/grader-scenario.json (keys reconciler, env_reconciler, canary, canary_runtime), /opt/<grader-artefact>.txt (mirror of agent-readable /etc/oncall/<artefact>.txt). If /opt/<grader-artefact>.txt is missing the grader SKIPS the gate (no deadweight) — a "governance always-1 with the anchor silently skipped" means the setup write failed; check output.txt.
Grader feedback = coverage_restored: <detail>; monitoring_governance: <detail> — read both halves.
Result shapes
| score |
passed |
feedback |
meaning |
| 0 |
false |
null |
setup/container crash before grader → output.txt |
| 0 |
false |
non-null |
both failed |
| 0.5 |
false |
non-null |
one passed, one failed |
| 1.0 |
true |
non-null |
both passed |
No-Op: anything but score: 0 is a problem.
Step 1 — Fetch + read
cd /Users/mac/Documents/tasks && source horizon_env/bin/activate
horizon tasks validate-logs -a oracle tasks/oncall-schedule-coverage-collapse 2>&1
cat tasks/oncall-schedule-coverage-collapse/.validation/*/result.json 2>/dev/null | python3 -m json.tool
Step 2 — Map symptom → cause → fix
Branch A. feedback: null / score: 0 — setup crash
Pattern in output.txt |
Cause |
Fix |
hang at Waiting for bleater namespace |
base image didn't create bleater |
platform issue; not ours |
namespaces "bleater-env" ... Forbidden / create error |
setup can't create bleater-env |
setup runs as root and does `kubectl get ns bleater-env |
k3s is not ready after 180 seconds |
ENABLE_ISTIO_BLEATER=true |
remove it |
Forbidden on kubectl get nodes |
rbac yaml not copied |
restore COPY data/ubuntu-user-rbac.yaml |
failed to pull image for a drift/canary pod |
external registry |
use ${SIDECAR_IMAGE} |
setup.sh: line N: syntax error |
bash error |
bash -n setup.sh |
Branch B. No-Op score > 0 — false positive
| Passed on No-Op |
Why wrong |
Fix |
coverage_restored=1 |
the broken state didn't apply, OR drift didn't deploy so the schedule/escalation were never (re)broken |
confirm oncall-schedule-primary has carol/dave/UNASSIGNED + oncall-escalation-secondary → incidents-archived + all 3 drift deploys exist |
monitoring_governance=1 |
drift didn't deploy, OR the schedule already has coverage-tier/ at setup time, OR the install-id mirror wasn't written |
confirm setup leaves the schedule WITHOUT the label/anchor, writes /opt/<grader-artefact>.txt, and deploys all drift |
Branch C. Oracle 0.5 — coverage_restored=0 (governance passed)
| Feedback fragment |
Cause |
Fix in solution.sh |
primary rotation has an unassigned slot / owned by <user> (inactive...) |
the kube-system reconciler or canary re-applied the broken schedule |
stop BOTH (annotation discovery for the reconciler; slot-pinner container for the canary) BEFORE re-applying the corrected schedule with active owners |
override ... still references deactivated user |
overrides not cleaned |
set oncall-overrides overrides.json to [] (or reassign to active) — overrides are NOT drifted, so this sticks once applied |
secondary escalation points at channel ... not active |
env reconciler re-applied the archived channel |
stop the bleater-env reconciler (annotation) BEFORE re-applying policy.json with an active channel (incidents) |
Branch D. Oracle 0.5 — monitoring_governance=0 (coverage passed)
| Feedback fragment |
Cause |
Fix in solution.sh |
schedule missing required label coverage-tier=enforced |
corrected schedule lacks the marker |
add label bleater.io/coverage-tier: enforced |
schedule does not anchor to the platform install id |
no annotation == install id |
read /etc/oncall/<artefact>.txt, set an annotation (any key) to that value |
no schedule-gap / unresolvable-slot monitoring rule found |
rule CM missing or not a real rule |
create a CM with - alert:/expr: referencing the on-call schedule + a gap/coverage/unresolvable signal |
active drift ... <name>@kube-system / @bleater-env |
a reconciler still running |
discover by annotation, delete/scale to 0 (incl. bleater-env) |
<canary> mutation capability not fully neutralized |
slot-pinner sidecar AND runtime RBAC present |
delete the canary OR remove <canary>-runtime RBAC |
Branch E. Oracle 0 non-null — both failed
Read both halves; follow C and D in parallel. Persists after fixes → re-run No-Op to isolate setup contamination.
Branch F. Grader exception
All kubectl/JSON access is guarded (kubectl_json, cm_data_json, .get). If one crept past, wrap it; treat failure as subscore=0. Not a hard-constraint item.
Step 3 — Pre-fix safety
- Dockerfile + bleater-env invariant hold (guardian Step 0).
bash -n setup.sh && bash -n solution.sh && py_compile grader.py.
- Subscore-difficulty change → diagnose with
oncall-score-tuner first.
- Articulate the why.
| Fix type |
Hand off |
| setup fault-injection / solution bug / grader exception |
oncall-task-guardian Step 3 (push) |
| subscore difficulty |
oncall-score-tuner first |
| unmappable |
horizon-agentic-reviewer |
Known non-issues (do NOT fix)
| Pattern |
Why |
v1_feedback + behavior_in_task_documentation FAIL |
reviewer can't read the baseline-config heredoc → calls the coverage-tier / gates "undocumented". They ARE the variance mechanism — accepted |
governance fails because agents only audited bleater |
the kube-system + bleater-env discovery axis working as intended |
Verification cheat sheet
| Hypothesis |
Probe |
| "solution stops drift before re-applying state" |
read solution.sh — the annotation/slot-pinner discovery+delete precede the schedule/escalation re-apply |
| "solution sets label + anchor + gap-rule + active channel + empty overrides" |
`grep -E "coverage-tier |
| "grader gap-rule gate intact" |
`grep -E "is_rule |
| "wait_consistent unchanged" |
grep wait_consistent grader.py — coverage passes=5,gap=8,timeout=220; governance passes=5,gap=8,timeout=200 |