CISM · Domain 4
Incident Management
About 30% of the exam
The incident lifecycle
- Preparation
- Identification
- Containment
- Eradication
- Recovery
- Lessons learned
- Preparation
- plan, team, tools, contacts, training
- Identification
- detect, triage, declare, classify
- Containment
- limit spread, preserve evidence
- Eradication
- remove root cause and artifacts
- Recovery
- restore, validate, monitor closely
- Lessons learned
- improve plan, controls, training
Contain before you eradicate, eradicate before you recover, and always close with lessons learned
Event, incident, triage
- Event
- any observable occurrence
- Incident
- event that harms or violates policy
- Triage
- is it real, how bad, who
- Declaration
- formal decision to activate the plan
- Severity
- business impact plus urgency
- Precursor
- sign an incident may occur
- Indicator
- sign an incident has occurred
Classify by business impact, not by how technical the attack looks
Preparation essentials
- Written plan approved by management
- Contact lists and communication channels
- Playbooks for common incident types
- Detection tooling and logging in place
- Retainers: forensics, legal, PR
- Regular exercises, then plan updates
- Authority to act defined in advance
Team and roles
- Incident manager
- decides, coordinates, tracks
- Technical lead
- analysis, containment, eradication
- Legal counsel
- evidence, notification, liability
- Communications
- internal and external messaging
- HR
- insider and staff matters
- Business owner
- impact, recovery priorities
- Executive sponsor
- authority, funding, escalation
Cross-functional by design; a purely technical team misses legal and business impact
Classification and escalation
- Severity scale defined before the incident
- Criteria: impact, scope, data type, urgency
- Escalation paths named per severity level
- Who is notified, when, through what channel
- Regulated data raises severity automatically
- Legal advises on notification triggers
- Reclassify as facts change
Containment choices
- Short-term
- isolate, block, disable accounts
- Long-term
- harden while rebuilding
- Ransomware first
- isolate affected systems, activate plan
- Evidence
- image before you wipe
- Business trade-off
- downtime versus continued exposure
- Watch and learn
- only with counsel and control
Decide containment with the business owner: stopping a system is a business decision
Eradication and recovery
- Find and remove the root cause
- Remove all artifacts, not the first one
- Close the vulnerability that let it in
- Rebuild from known good, then patch
- Restore in order of business criticality
- Validate integrity before reconnecting
- Heightened monitoring after return to service
- Activate DRP when normal response cannot recover
Evidence and forensics
Chain of custody
- Who collected, when, where, how
- Every handoff signed and timed
- Hash images, work on copies
- Secure storage, limited access
- Gap in the chain, evidence questioned
Order of volatility
- Registers and cache first
- Memory, then network state
- Running processes, then disk
- Logs and remote data
- Archive and backup media last
Manager's decisions
- Engage counsel before collecting for court
- Balance investigation against recovery
- Preserve, then restore, when law is involved
- Forensic depth matches legal exposure
- Document every decision and its time
Admissibility rests on integrity and custody; a fast recovery that destroys evidence can cost more later
Communications
- Timely, accurate, audience appropriate
- One trained spokesperson for media
- Legal reviews external statements first
- Regulators and customers per notification law
- Internal updates on a set cadence
- Out-of-band channel if email is compromised
- Never speculate on cause or scope
Lessons learned
- Held soon after closure, blameless
- What happened, what worked, what failed
- Root cause, not just the symptom
- Actions with owners and due dates
- Update plan, playbooks, controls, training
- Track actions to completion
- Repeat incidents mean actions never landed
The output is tracked change, not a report that sits in a folder
BIA and recovery metrics
Business impact analysis
- Identify critical processes and dependencies
- Quantify impact over time of outage
- Owners set the numbers, security facilitates
- Drives recovery priority and strategy
- Revisit after major business change
The metrics
- MTD
- maximum tolerable downtime, the outer limit
- RTO
- time to restore the service
- RPO
- acceptable data loss, in time
- WRT
- work recovery after systems return
- SDO
- service delivery objective, degraded level
- AIW
- acceptable interruption window, business view
RTO plus WRT must fit inside MTD; RPO sets backup frequency; shorter targets cost more
Plans and recovery sites
Which plan
- IRP
- handle the security incident
- BCP
- keep the business running
- DRP
- restore IT after disaster
- Crisis management
- people, safety, communications
- Trigger
- IRP escalates to DRP when recovery fails
Site options
- Hot
- ready in hours, highest cost
- Warm
- equipment present, data to load
- Cold
- space and power, days to weeks
- Mobile
- trailer-based, flexible location
- Reciprocal
- agreement with a peer, rarely tested
- Cloud
- elastic, mind the shared responsibility
Match the site to the RTO the business paid for, not to the cheapest option
Testing and exercises
- Checklist
- Tabletop
- Walkthrough
- Simulation
- Parallel
- Full interruption
What each proves
- Checklist: plan is complete and current
- Tabletop: roles understood, decisions rehearsed
- Simulation: response under realistic pressure
- Parallel: alternate site works alongside
- Full interruption: real failover, real risk
Manager's rules
- Test regularly and after significant change
- Start low rigor, earn the right to go further
- Every test ends in plan updates
- Recovery slower than RTO is a finding
- Full interruption needs executive approval
Know the order
- Detect
- Declare
- Contain
- Preserve evidence
- Eradicate
- Recover
- Review
- Isolate before you investigate deeply
- Image before you rebuild
- Restore critical services first
- Notify per law, on the legal clock
Rapid recall: the FIRST move
- Activate the plan, then act within it
- Contain to limit impact, preserve evidence
- PII involved: check notification obligations
- Legal exposure: engage counsel early
- No matching playbook: apply principles, document
- Recovery order follows business criticality
- Repeat incidents: fix the action tracking
Reference strip: phases, metrics, plans, evidence
Lifecycle phases
- Preparation: plan, team, tools
- Identification: detect, triage, declare
- Containment: limit spread, keep evidence
- Eradication: remove cause and artifacts
- Recovery: restore, validate, monitor
- Lessons learned: tracked improvements
Recovery metrics
- MTD: the outer survival limit
- RTO: restore time target
- RPO: data loss tolerated
- WRT: catch-up work after restore
- SDO and AIW: degraded service, window
Plans
- IRP: security incidents
- BCP: business keeps operating
- DRP: IT restored after disaster
- Crisis plan: people and messaging
- All aligned on objectives and communications
Evidence words
- Chain of custody
- Order of volatility
- Forensic image and hash
- Admissibility and integrity
- Legal hold and preservation
Test ladder
- Checklist review of the plan
- Tabletop and walkthrough discussion
- Simulation with scenario injects
- Parallel run at the alternate site
- Full interruption, highest risk
Quick exam traps
- Trap: Every security event is an incident
- Trap: Eradicate the malware first, then worry about containment
- Trap: Wiping and rebuilding quickly is always the right call
- Trap: The incident response team should be technical staff only
- Trap: A full interruption test is the place to start testing
- Trap: RTO can be longer than MTD if the business agrees
- Trap: The DRP is activated at the start of every incident
- Trap: Lessons learned is complete once the report is written
cybercertprep.com · original revision sheet written from the public body of knowledge