CC · Domain 2
Business Continuity, Disaster Recovery and Incident Response
About 10% of the exam
Three plans, three jobs
- Incident response plan
- handle the event as it happens
- Business continuity plan
- keep critical functions running
- Disaster recovery plan
- restore the technology afterwards
- Order of use
- respond, continue, then recover
- Scope
- whole business versus IT systems
- Approval
- senior management signs each plan
One event can trigger all three: respond to the incident, continue the business, then recover the systems
Business impact analysis
- RTO
- target time to restore service
- RPO
- acceptable data loss in time
- MTD
- longest outage the business survives
- Work recovery time
- checking and catching up afterwards
- MTBF
- average uptime between failures
- MTTR
- average time to repair
- Critical function
- the work that must not stop
- SLA
- the promised service level
Back up more often than the RPO and restore faster than the RTO, and never exceed the maximum tolerable downtime
Backup types
- Full
- everything, slowest to take
- Incremental
- changes since any last backup
- Differential
- changes since the last full
- Snapshot
- a point in time image
- Immutable
- cannot be altered or encrypted
- Offsite copy
- survives loss of the site
- 3-2-1 rule
- three copies, two media, one offsite
- Restore test
- an untested backup is guesswork
Recovery sites
- Hot site
- live, current data, minutes
- Warm site
- equipment ready, data restored
- Cold site
- space, power, nothing else
- Mobile site
- a trailer brought to you
- Cloud recovery
- pay for capacity when needed
- Reciprocal
- another organization hosts you
The incident response lifecycle
Prepare and detect
- Plan, roles and contacts ready
- Train staff to report quickly
- An event is any observable occurrence
- An incident harms or could harm
- Users are a real detection source
Contain and eradicate
- Short term: isolate the host
- Long term: rebuild and harden
- Preserve evidence while containing
- Remove malware, backdoors and rogue accounts
- Establish how they got in
Recover and learn
- Restore from known good backups
- Monitor closely after the return
- Confirm normal operations have resumed
- Lessons learned with owners and dates
- Update the plan and playbooks
Containment comes before eradication: stop the bleeding, remove the cause, and only then restore service
Evidence handling
- Chain of custody
- who held it, when and why
- Forensic image
- a bit for bit copy
- Hashing
- proves the copy is unchanged
- Volatile data
- memory vanishes when powered off
- Write blocker
- prevents changes to the original
- Legal hold
- stop deleting relevant records
Testing, least to most disruptive
- Read-through
- Tabletop
- Walkthrough
- Simulation
- Parallel
- Full interruption
- Start with the least disruptive
- Parallel runs the alternate site too
- Full interruption stops the primary site
- Findings have to change the plan
Redundancy
- RAID 1 mirrors two disks
- RAID 5 stripes with distributed parity
- RAID 6 survives two disk failures
- RAID 10 mirrors and stripes together
- RAID gives uptime, not backup
- UPS bridges the gap to generator
- Clustering fails over between nodes
- Remove single points of failure
How incidents get noticed
- SIEM
- correlates logs into alerts
- IDS
- detects and alerts only
- IPS
- detects and blocks inline
- EDR
- records endpoint activity continuously
- Antivirus
- matches known malicious files
- User report
- often the very first signal
- Impossible travel
- two logins too far apart
People and communication
- Incident lead
- runs the response and decides
- Escalation path
- who is told, and when
- Out-of-band channel
- assume normal email is watched
- Legal and privacy
- decide notification obligations
- Communications team
- handles customers and press
- Regulator clock
- 72 hours under GDPR
- Law enforcement
- involved per policy and severity
Know the order
- Plans
- incident, continuity, recovery
- Response
- prepare, detect, contain, eradicate, recover
- After recovery
- lessons learned, then plan update
- Tests
- read-through up to full interruption
- Volatility
- memory before disk before backups
- Restore
- full first, then incrementals in order
Reference strip: plans, metrics, backups, response
Plans
- Incident response, continuity, disaster recovery
- Evacuation and emergency procedures
- Playbooks for common scenarios
- Contact lists kept current
Metrics
- RTO, RPO, MTD, work recovery time
- MTBF and MTTR
- Service level agreement targets
- Downtime cost per hour
Backups and sites
- Full, incremental, differential, snapshot
- Immutable and offsite copies
- Hot, warm, cold, mobile, cloud
- Three copies, two media, one offsite
Response
- Event versus incident versus breach
- Short term and long term containment
- Eradication then recovery
- Chain of custody and legal hold
Resilience
- RAID levels 1, 5, 6 and 10
- Clustering and load balancing
- UPS, generator, dual feeds
- Geographic separation of sites
Quick exam traps
- Trap: A backup nobody has restored is still proven
- Trap: RAID protects you from accidental file deletion
- Trap: RTO and RPO measure the same interval
- Trap: The disaster recovery plan covers the whole business
- Trap: Rebuilding at once beats preserving the evidence
- Trap: Offsite backups make a recovery site unnecessary
- Trap: Only IT staff need to attend the tabletop exercise
cybercertprep.com · original revision sheet written from the public body of knowledge