AEM Backup & Disaster Recovery
Purpose
Design and implement backup strategies, disaster recovery plans, and business continuity procedures for AEM environments including repository backup, content recovery, and failover architecture.
When to Use (Triggers)
- User mentions "backup," "disaster recovery," "DR," "failover," or "business continuity"
- References to repository backup, restore procedures, or data protection
- Questions about RTO/RPO requirements, backup scheduling, or recovery testing
- Requests involving high availability, geographic redundancy, or failover configuration
- Discussion of data loss prevention, backup validation, or restore procedures
Core Capabilities
- Design backup strategies for all AEM data (repository, datastore, configurations)
- Implement automated backup scheduling with retention policies
- Configure high availability and failover architectures
- Build and document disaster recovery runbooks
- Test and validate recovery procedures with measurable RTO/RPO
Domain Knowledge Required
Technical Foundation
- Backup types: full, incremental, differential, snapshot-based
- Recovery objectives: RTO (Recovery Time Objective), RPO (Recovery Point Objective)
- High availability patterns: active-active, active-passive, cold standby
- Data consistency during backup (quiescing, point-in-time snapshots)
AEM-Specific Context
- Oak TarMK online backup (FileDataStore + SegmentNodeStore backup)
- AEM Cold Standby (TarMK standby sync)
- Document store (MongoDB) backup patterns for DocumentMK
- Backup of external datastore (S3, Azure Blob, shared filesystem)
- AEM Cloud Service managed backup vs. on-premise manual backup
- Restore procedures and data consistency verification
Implementation Approach
Step 1: Requirements Analysis
Define recovery objectives and scope.
- Establish RTO and RPO for each environment (author, publish)
- Identify critical data: repository content, configurations, datastore binaries
- Determine compliance requirements (retention periods, geographic restrictions)
- Assess acceptable data loss and downtime per tier
Step 2: Backup Architecture Design
Plan the backup approach for each data component.
- Configure Oak online backup for repository segments
- Set up datastore backup (filesystem copy, S3 versioning, or snapshots)
- Plan configuration backup (OSGi configs, dispatcher, infrastructure-as-code)
- Design backup storage location (offsite, different region, different provider)
Step 3: Automation & Scheduling
Implement automated backup execution.
- Schedule Oak online backup via cron/maintenance window
- Automate datastore backup aligned with repository backup timing
- Implement backup rotation and retention (daily, weekly, monthly)
- Configure backup monitoring and failure alerting
Step 4: High Availability Configuration
Set up failover capabilities.
- Configure TarMK Cold Standby for author instance failover
- Set up publish farm with load balancer health checks
- Implement dispatcher-level failover for publish tier
- Design cross-region failover for geographic disaster scenarios
Step 5: Recovery Testing
Validate recovery procedures regularly.
- Schedule quarterly DR drills with documented runbooks
- Test full restore to isolated environment
- Measure actual RTO/RPO against targets
- Document and remediate gaps found during testing
Quality Checklist
Related Skills
- aem-monitoring-alerting (backup monitoring)
- aem-versioning-content-rollback (content-level recovery)
- aem-upgrade-patch-management (safe upgrade with backup)
Example Use Cases
- Enterprise DR Strategy: Design a multi-region disaster recovery architecture for a critical AEM author instance with 4-hour RTO, 1-hour RPO, automated failover, and documented recovery runbooks tested quarterly.
- Cloud-to-On-Premise Backup: Implement backup strategy for AEM Cloud Service content with nightly exports to on-premise storage for compliance requirements, including content package extraction and verification.
- Ransomware Recovery Plan: Create an air-gapped backup architecture with immutable backups, integrity verification, and a tested recovery procedure that can restore a clean AEM instance within 8 hours.
Notes
- AEM Cloud Service manages backups automatically — focus DR planning on content recovery, not infrastructure
- Oak online backup creates a consistent snapshot — do NOT use filesystem copy while AEM is running
- Cold Standby provides near-real-time replication but requires dedicated standby instance
- Test restores are critical — a backup that hasn't been tested is not a backup
1---2name: aem-backup-disaster-recovery3description: AEM Backup & Disaster Recovery4---5# AEM Backup & Disaster Recovery67## Purpose8Design and implement backup strategies, disaster recovery plans, and business continuity procedures for AEM environments including repository backup, content recovery, and failover architecture.910## When to Use (Triggers)11- User mentions "backup," "disaster recovery," "DR," "failover," or "business continuity"12- References to repository backup, restore procedures, or data protection13- Questions about RTO/RPO requirements, backup scheduling, or recovery testing14- Requests involving high availability, geographic redundancy, or failover configuration15- Discussion of data loss prevention, backup validation, or restore procedures1617## Core Capabilities18- Design backup strategies for all AEM data (repository, datastore, configurations)19- Implement automated backup scheduling with retention policies20- Configure high availability and failover architectures21- Build and document disaster recovery runbooks22- Test and validate recovery procedures with measurable RTO/RPO2324## Domain Knowledge Required25### Technical Foundation26- Backup types: full, incremental, differential, snapshot-based27- Recovery objectives: RTO (Recovery Time Objective), RPO (Recovery Point Objective)28- High availability patterns: active-active, active-passive, cold standby29- Data consistency during backup (quiescing, point-in-time snapshots)3031### AEM-Specific Context32- Oak TarMK online backup (FileDataStore + SegmentNodeStore backup)33- AEM Cold Standby (TarMK standby sync)34- Document store (MongoDB) backup patterns for DocumentMK35- Backup of external datastore (S3, Azure Blob, shared filesystem)36- AEM Cloud Service managed backup vs. on-premise manual backup37- Restore procedures and data consistency verification3839## Implementation Approach40### Step 1: Requirements Analysis41Define recovery objectives and scope.42- Establish RTO and RPO for each environment (author, publish)43- Identify critical data: repository content, configurations, datastore binaries44- Determine compliance requirements (retention periods, geographic restrictions)45- Assess acceptable data loss and downtime per tier4647### Step 2: Backup Architecture Design48Plan the backup approach for each data component.49- Configure Oak online backup for repository segments50- Set up datastore backup (filesystem copy, S3 versioning, or snapshots)51- Plan configuration backup (OSGi configs, dispatcher, infrastructure-as-code)52- Design backup storage location (offsite, different region, different provider)5354### Step 3: Automation & Scheduling55Implement automated backup execution.56- Schedule Oak online backup via cron/maintenance window57- Automate datastore backup aligned with repository backup timing58- Implement backup rotation and retention (daily, weekly, monthly)59- Configure backup monitoring and failure alerting6061### Step 4: High Availability Configuration62Set up failover capabilities.63- Configure TarMK Cold Standby for author instance failover64- Set up publish farm with load balancer health checks65- Implement dispatcher-level failover for publish tier66- Design cross-region failover for geographic disaster scenarios6768### Step 5: Recovery Testing69Validate recovery procedures regularly.70- Schedule quarterly DR drills with documented runbooks71- Test full restore to isolated environment72- Measure actual RTO/RPO against targets73- Document and remediate gaps found during testing7475## Quality Checklist76- [ ] Backup covers all data components (repository, datastore, config)77- [ ] Backup consistency verified (no corrupt or partial backups)78- [ ] RTO/RPO measured and within business requirements79- [ ] Backup retention meets compliance requirements80- [ ] Automated alerting on backup failures81- [ ] Recovery procedures documented in runbook format82- [ ] DR tested quarterly with documented results83- [ ] Backup storage in separate failure domain from production8485## Related Skills86- aem-monitoring-alerting (backup monitoring)87- aem-versioning-content-rollback (content-level recovery)88- aem-upgrade-patch-management (safe upgrade with backup)8990## Example Use Cases911. **Enterprise DR Strategy:** Design a multi-region disaster recovery architecture for a critical AEM author instance with 4-hour RTO, 1-hour RPO, automated failover, and documented recovery runbooks tested quarterly.922. **Cloud-to-On-Premise Backup:** Implement backup strategy for AEM Cloud Service content with nightly exports to on-premise storage for compliance requirements, including content package extraction and verification.933. **Ransomware Recovery Plan:** Create an air-gapped backup architecture with immutable backups, integrity verification, and a tested recovery procedure that can restore a clean AEM instance within 8 hours.9495## Notes96- AEM Cloud Service manages backups automatically — focus DR planning on content recovery, not infrastructure97- Oak online backup creates a consistent snapshot — do NOT use filesystem copy while AEM is running98- Cold Standby provides near-real-time replication but requires dedicated standby instance99- Test restores are critical — a backup that hasn't been tested is not a backup