# Operational Resilience Tester

> Describe what this skill helps an agent do.

- Skill: `domehahn/operational-resilience-tester-2` (Agent Skill, multi-file: 10 files)
- Install (CLI): `npx skillmds@latest add domehahn/operational-resilience-tester-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/domehahn/operational-resilience-tester-2/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: domehahn (https://skillmd.com/u/domehahn)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/domehahn/operational-resilience-tester-2

---

# Operational Resilience Tester

## Purpose

Review backup and restore, failover, disaster recovery, restart procedures, crisis exercises, scenario tests, and lessons learned. Treat regulatory, security, and operational references as review and evidence guidance, not legal advice.

## When to use

- operational resilience testing decisions, controls, or operating practices need independent review.
- A change affects operational resilience testing artifacts such as backup policy, restore test log, failover runbook, DR plan, crisis exercise report, lessons-learned register.
- The user needs evidence-oriented findings for risks such as unproven restore, manual failover bottleneck, RTO mismatch, unrehearsed crisis role, scenario gap, unclosed lesson.
- Audit, security, operations, or platform stakeholders need a concise readiness position.
- Existing documentation, tickets, tests, or logs must be turned into actionable remediation items.

## Operating model

1. Identify the relevant operational resilience testing artifacts, owners, systems, environments, and review boundary.
2. Compare the available artifacts against expected signals such as RPO evidence, RTO measurement, failover transcript, exercise participant record, defect ticket, retest proof.
3. Separate confirmed gaps from assumptions, missing evidence, and advisory improvement opportunities.
4. Rate findings by operational, security, compliance, customer, and auditability impact.
5. Recommend minimal remediation steps, validation evidence, owners, and review cadence.

## Spec-Driven Change Context

- Treat repository specs, ADRs, runbooks, change proposals, design notes, and task files as durable context that outlives a chat session.
- For non-trivial changes, prefer a checked-in change artifact or equivalent proposal/design/tasks record before implementation begins.
- Capture requirement deltas explicitly: added, modified, removed, deprecated, or unchanged behavior.
- Keep implementation tasks traceable to acceptance criteria, affected specs, validation commands, and owners.
- During verification, compare the implementation against the proposal, design decisions, task checklist, and spec deltas.
- After completion, sync or archive completed change artifacts so the repository's source of truth reflects the final behavior.
- If the repository has no spec workflow yet, report the missing artifact and provide a minimal proposal/spec/tasks outline instead of relying on chat-only intent.

## Skill-Specific Review Scope

- Primary artifacts: backup policy, restore test log, failover runbook, DR plan, crisis exercise report, lessons-learned register.
- Risk themes: unproven restore, manual failover bottleneck, RTO mismatch, unrehearsed crisis role, scenario gap, unclosed lesson.
- Evidence signals: RPO evidence, RTO measurement, failover transcript, exercise participant record, defect ticket, retest proof.
- Ownership, approvals, review cadence, exception handling, and residual-risk decisions.
- Traceability from requirement or control intent to implementation, validation, and retained evidence.

## Skill-Specific Checklist

- [ ] Confirm the review boundary covers the right operational resilience testing systems, teams, and environments.
- [ ] Inventory and inspect the current backup policy.
- [ ] Check whether restore test log is current, approved, versioned, and owned.
- [ ] Verify that failover runbook has test, ticket, log, or approval support.
- [ ] Look for unproven restore and record concrete repository or process evidence.
- [ ] Look for manual failover bottleneck and identify affected assets, services, or stakeholders.
- [ ] Look for RTO mismatch and classify the operational or audit impact.
- [ ] Use RPO evidence to validate that the control or practice is operating.
- [ ] Use RTO measurement to confirm ownership, timing, and reproducibility.
- [ ] Check exception, risk-acceptance, and expiry handling for operational resilience testing.
- [ ] Confirm remediation items have owners, due dates, validation steps, and evidence expectations.
- [ ] Identify missing artifacts separately from weak artifacts so the next action is unambiguous.
- [ ] Review whether logging, reporting, or retained evidence exposes sensitive data unnecessarily.

## Decision Rules

- If backup policy is missing for a critical service, raise at least a high-severity readiness gap.
- If RTO measurement cannot be tied to an owner and approval, treat the outcome as unauditable until corrected.
- If DR plan is present but expired or untested, require validation before accepting residual risk.
- If the only support is verbal or chat-only context, request durable ticket, document, log, or test evidence.
- If remediation would require a process or architecture decision, assign a decision owner instead of prescribing legal conclusions.
- If compensating measures reduce likelihood but not impact, keep the residual-risk statement explicit.

## Finding Categories

- Missing or stale operational resilience testing artifact.
- Unclear ownership, approval, review cadence, or accountability.
- Insufficient validation, test proof, logs, ticket trail, or retained audit material.
- Unreviewed exception, residual risk, expiry, or compensating measure.
- Policy, architecture, operational, or platform implementation drift.
- Sensitive-data exposure in logs, reports, prompts, artifacts, or evidence packages.

## Severity Guidance

- Critical: a gap in operational resilience testing creates immediate outage, data-loss, privilege, regulatory-reporting, or irreversible business risk.
- High: backup policy is missing, unowned, untested, or unauditable for a critical service or material change.
- Medium: restore test log exists but is stale, incomplete, inconsistently enforced, or weakly evidenced.
- Low: wording, metadata, formatting, link freshness, or minor traceability improvements are needed.

## DevSecOps Guardrails

- Do not read secrets, `.env` files, private keys, production credentials, masked CI/CD variables, database dumps, or sensitive logs unless explicitly required.
- Do not push, deploy, publish, merge, or create releases unless explicitly asked.
- Prefer merge requests, reviewable diffs, and auditable validation evidence.
- Prefer least privilege, minimal changes, and explicit rollback notes.
- Do not fabricate test results, repository state, commands, security findings, or validation outcomes.
- Report assumptions, uncertainty, residual risk, and validation gaps clearly.

## Output Requirements

- Findings ordered by severity with affected operational resilience testing artifacts and evidence references.
- Coverage note for reviewed artifacts: backup policy, restore test log, failover runbook, DR plan, crisis exercise report, lessons-learned register.
- Risk note covering relevant themes: unproven restore, manual failover bottleneck, RTO mismatch, unrehearsed crisis role, scenario gap, unclosed lesson.
- Evidence request list using expected signals: RPO evidence, RTO measurement, failover transcript, exercise participant record, defect ticket, retest proof.
- Deliverables or updates needed: resilience test findings, RPO/RTO evidence table, scenario coverage map, recovery gap backlog, lessons-learned closure plan.
- Residual-risk, assumptions, missing-context, and validation-gap summary.

## Acceptance Criteria

- Relevant operational resilience testing artifacts are identified, current, owned, and versioned where applicable.
- Each high-impact finding includes evidence, impact, likelihood, owner, and remediation guidance.
- Missing evidence is separated from failed controls or weak implementation.
- Exceptions and risk acceptances include owner, rationale, expiry, and compensating measures.
- Recommendations are review-oriented and avoid presenting regulatory interpretation as legal advice.
- Final output states pass, conditional pass, or blocked readiness with validation gaps.

## Anti-Patterns

- Treating a policy title or control name as proof that the practice operates effectively.
- Collapsing missing evidence and failed implementation into one vague finding.
- Accepting open-ended exceptions without owner, expiry, impact, likelihood, and compensating measures.
- Making legal, regulatory, or audit conclusions beyond the available evidence and review scope.
- Recommending broad process rewrites when a targeted owner, test, ticket, or evidence fix is enough.
- Copying sensitive production data into examples, evidence packages, prompts, or reports.

## Changelog

### 0.1.0 - 2026-07-28

- Initial generated production-ready SDLC / DevSecOps skill.

