CISA · Domain 4
Information Systems Operations and Business Resilience
About 26% of the exam
Incident vs problem management
- Incident
- restore service fast, minimal impact
- Problem
- find and remove the root cause
- Known error
- root cause found, workaround documented
- Service desk
- single point of contact
- SLA
- agreed service levels with customer
- OLA
- internal agreement supporting the SLA
- CSI
- continual improvement, often missing
Incident asks how fast can we restore; problem asks why it keeps happening
Change and configuration
- CAB evaluates, prioritizes, authorizes changes
- Standard, normal, emergency change types
- Emergency changes reviewed after deployment
- CMDB records assets and their relationships
- Impact analysis depends on an accurate CMDB
- Job schedule changes need approval too
- Canary or blue green with fast rollback
- VM cloning and migration under approval
Capacity, performance, patching
- Capacity planning meets current and future demand
- Still needed in auto-scaling cloud: cost, limits
- Automated monitoring with threshold alerts
- Logs collected, reviewed, retained long enough
- Seven day log retention hides slow attacks
- Patch within policy window, track exceptions
- End-of-life systems are a high-risk finding
- Observability data needs access control too
Backups
- Full
- everything, slow backup, fast restore
- Incremental
- changes since last backup, any type
- Differential
- changes since last full
- Restore incremental
- last full plus every incremental
- Restore differential
- last full plus latest differential
- Grandfather father son
- monthly, weekly, daily rotation
- Immutable offline copy
- survives ransomware dwell time
- Offsite
- far enough to escape the disaster
Backups nobody has restored are hope, not a control; test the restore
Business continuity lifecycle
- Policy and scope
- BIA
- Risk assessment
- Strategy
- Plan
- Test
- Maintain
- Policy
- board mandate, business driven, not IT
- BIA
- critical processes, impact over time
- Strategy
- recovery options priced against RTO and RPO
- Plan
- roles, procedures, contacts, vital records
- Test
- prove RTO and RPO are met
- Maintain
- annually and on significant change
BCP owned by IT alone is the most fundamental flaw; the business decides what is critical
BIA outputs
- RTO
- maximum time to restore the service
- RPO
- maximum data loss measured in time
- MTD
- outage that threatens survival
- WRT
- verify data, catch up backlog
- SDO
- reduced service level accepted during recovery
- MTBF, MTTR
- reliability and repair time
- Criticality ranking
- recovery sequence and dependencies
RTO plus WRT must fit inside MTD; RPO drives backup frequency and replication choice
Recovery site options
- Hot site
- equipped, current data, hours
- Warm site
- equipped, data loaded on arrival, days
- Cold site
- space and power only, weeks
- Mirrored
- active-active, near zero RTO
- Mobile site
- trailer delivered to location
- Reciprocal agreement
- rarely enough spare capacity
- DRaaS
- cloud recovery, regional outage still risk
Same city for primary and recovery sites fails the regional disaster test
DR test types
- Checklist
- Tabletop
- Simulation
- Parallel
- Full interruption
- Checklist: desk review, least effective
- Tabletop: talk through roles and gaps
- Simulation: act the scenario, no failover
- Parallel: recovery site processes alongside production
- Full interruption: real failover, highest risk
- Measure actual RTO and data loss achieved
- Post-test review fixes root causes, retest
Replication and architecture
- Synchronous
- zero data loss, distance limited
- Asynchronous
- lag equals potential data loss
- Active-active
- both sites live, automated failover
- Active-passive
- standby waits, manual or scripted
- Near-zero RPO
- synchronous replication required
- Dependency map
- recover services in the right order
- Cloud DR
- customer still owns the plan
Thirty minute asynchronous lag means up to thirty minutes lost, whatever management believes
What the plan must contain
- Crisis communication to staff, customers, regulators
- Named roles with trained successors
- Vital records program for essential documents
- Work area recovery for people, not just servers
- Pandemic and remote workforce scenarios
- Third-party dependencies inside the BIA
- Regulatory DR expectations validated
- Contact trees, call procedures, escalation
Operational access and data
- Terminated accounts disabled the same day
- No shared root or administrator accounts
- PAM for individual accountability and rotation
- DBA cannot also control audit logs
- Secrets vault, never hardcoded credentials
- Pull printing for sensitive reports
- Decommissioned drives sanitized with chain of custody
- Multi-cloud needs one governance and monitoring view
Cyber resilience
- DRP must cover ransomware, not just fire
- Backups may already be compromised
- Rebuild from verified clean images
- Close the infection vector before restoring
- Retention longer than attacker dwell time
- Immutable, air-gapped backup copies
- Insurance compensates, it does not recover
- Single DR coordinator is a single point of failure
Key formulas
- MTD
- RTO + WRT, never exceeded
- RPO
- sets backup or replication interval
- Availability
- MTBF divided by MTBF plus MTTR
- Incremental restore
- full plus each incremental in order
- Differential restore
- full plus the latest differential
- Recovery cost
- falls as RTO lengthens
Where the cost of downtime curve meets the cost of recovery curve is your target RTO
Know the order
- BIA
- Strategy
- Plan
- Test
- Maintain
- Tests: checklist, tabletop, simulation, parallel, full
- Sites by readiness: mirrored, hot, warm, cold
- Incident: detect, log, categorize, restore, close
- Problem: identify, root cause, known error, fix
- Change: request, assess, approve, implement, review
Reference strip: ITSM, backup, BIA metrics, sites, tests
ITSM processes
- Incident, problem, change, release
- Configuration and CMDB
- Capacity, availability, continuity
- Service level, service desk, CSI
- ITIL as the common reference
Backup and media
- Full, incremental, differential
- GFS rotation, tower of Hanoi
- Onsite, offsite, cloud, immutable
- Restore testing, media retention, sanitization
BIA metrics
- RTO, RPO, MTD, WRT, SDO
- MTBF, MTTR, availability
- Criticality tiers and dependencies
- Impacts: financial, operational, legal, reputational
Site and replication types
- Mirrored, hot, warm, cold, mobile
- Reciprocal, DRaaS, multi-region cloud
- Synchronous vs asynchronous replication
- Active-active vs active-passive
Test rigor ladder
- Checklist review, desk based
- Tabletop walkthrough with key roles
- Simulation without failover
- Parallel processing at recovery site
- Full interruption, production cut over
Quick exam traps
- Trap: Problem management restores service as quickly as possible
- Trap: A four hour RPO means systems come back within four hours
- Trap: The cloud provider handles disaster recovery once you migrate
- Trap: Auto-scaling makes capacity planning unnecessary
- Trap: Restoring the most recent backup is a complete ransomware recovery plan
- Trap: A warm site has current data ready to run
- Trap: Full interruption is the right first test of an untested plan
- Trap: The IT department owns the business continuity plan
cybercertprep.com · original revision sheet written from the public body of knowledge