Denial of Service (DoS / DDoS)
Also known as: DoS, DDoS, Distributed denial of service, Resource exhaustion, Layer 7 flood
An attack that exhausts the bandwidth, connection capacity or compute of a service so legitimate users cannot use it, either by sheer traffic volume or by repeatedly hitting expensive application operations.
How it works
A denial of service makes a system unavailable to the people who should be able to use it by exhausting something finite: network bandwidth, connection table slots, CPU, memory, database capacity or a downstream dependency. When the traffic comes from many sources at once it is a distributed denial of service (DDoS), which is harder to filter by a single address.
Floods fall into broad families. Volumetric floods try to fill the network link itself. Protocol floods consume state in devices such as firewalls and load balancers. Application-layer floods send requests that each look legitimate but are expensive to serve, such as search, report generation or login, so a modest request rate can exhaust a back end.
From the analyst's seat the question is always the same: is the load abnormal, where does it land, and who is it from? Volumetric events show up as link saturation and packet-per-second spikes at the edge. Application-layer events show up as a rising latency and error rate concentrated on a few endpoints, with a broad spread of source addresses that individually look quiet.
Defence is layered because no single control covers every family. Capacity and upstream scrubbing absorb volume, CDN and WAF rules filter at the edge, rate limits and quotas protect expensive operations, caching and load shedding keep core functions alive, and timeouts and circuit breakers stop one overloaded dependency from taking everything down.
Walk through it
- 1Read the dashboard
- 2Check the source distribution
- 3Review the protective rule
- Escalate for upstream scrubbing
- Contain, scale safely and recover
Pager at 20:02: checkout is timing out. Before blaming a bug, read the metrics. Bandwidth is only mildly up, but one endpoint stands out for latency and errors. Type the path that is under stress.
Last 15 minutes, by endpoint
Link utilisation: 38% (normal peak 30%)
/ p95 120 ms errors 0.2% 1,900 req/s
/events p95 210 ms errors 0.4% 2,300 req/s
/search?q= p95 9,800 ms errors 41% 6,400 req/s
/checkout p95 5,200 ms errors 18% 210 req/s
Spot it
- Sudden rise in request rate, packets per second or connection count well above the normal daily curve.
- Latency and error rate climbing on a small set of endpoints while the rest of the site is healthy.
- A large increase in distinct source addresses, often with identical headers, user agents or request shapes.
- Cache hit ratio collapsing because requests carry unique parameters.
- Connection table, thread pool or database connection pool near exhaustion on the origin or a load balancer.
Edge / CDN log
ts=2026-10-11T20:02:11Z src=203.0.113.44 method=GET path=/search?q=a81f2 status=200 ms=9120 cache=MISS ua="Mozilla/5.0 (X11)"
ts=2026-10-11T20:02:11Z src=198.51.100.9 method=GET path=/search?q=77c0e status=504 ms=10000 cache=MISS ua="Mozilla/5.0 (X11)"Load balancer
ts=2026-10-11T20:02:30Z backend=app-pool healthy=3/12 active_conn=48210 5xx_rate=0.41 queue_depth=9800splSplunk, endpoints with abnormal volume and error rate
index=edge sourcetype=cdn
| bin _time span=1m
| stats count AS reqs, dc(src) AS sources, avg(ms) AS avg_ms, count(eval(status>=500)) AS errors by _time, path
| where reqs > 3*avg(reqs) AND errors/reqs > 0.1Compare against a rolling baseline for the same hour of day to avoid flagging a planned sale.
kqlKQL, distinct sources per path in the last 10 minutes
EdgeLogs
| where TimeGenerated > ago(10m)
| summarize Requests=count(), Sources=dcount(ClientIp), P95=percentile(DurationMs,95) by Path
| where Sources > 2000 and P95 > 2000
| order by Requests descStop it
Absorb volume upstream and filter at the edge
Put the service behind a CDN and a DDoS scrubbing provider with more capacity than any single flood, and use anycast so load spreads across sites. Volumetric traffic must be dropped before it reaches your link.
Protect expensive operations
Apply rate limits, quotas and challenges per client and per endpoint, cache aggressively, and cap the cost of a single request with timeouts, pagination and query limits.
Degrade gracefully
Shed low-priority work, reserve capacity for critical functions, and use circuit breakers and timeouts so an overloaded dependency fails fast instead of dragging everything down. Autoscaling helps but can become an expensive way to feed a flood.
Edge rate limit with challenge and caching
Hardened
rules:
, name: protect-search
match: { path: /search }
rate_limit: { key: client_ip, requests: 20, per: 60s }
on_exceed: challenge
cache: { ttl: 30s, key: normalised_query }Syntax is illustrative; map it to your CDN or WAF rule language.
Timeout and bounded work on an expensive handler
Vulnerable
const rows = await db.query(buildSearch(req.query.q));Hardened
const q = String(req.query.q).slice(0, 64);
const rows = await withTimeout(db.query(buildSearch(q), { limit: 25 }), 1500);
// breaker.open() after repeated timeouts so the handler fails fastwithTimeout and breaker stand for your platform's timeout and circuit-breaker helpers.
- Front public services with a CDN and a contracted DDoS scrubbing service, and keep origin addresses private.
- Rate limit and challenge per client on expensive endpoints such as search, login and report generation.
- Cache responses and normalise query keys so unique parameters do not defeat the cache.
- Set timeouts, connection limits and circuit breakers on every service and dependency.
- Reserve capacity for critical flows and define what to shed first under load.
- Rehearse a DDoS runbook with your provider, including who can authorise upstream mitigation.
If it already happened
Engage the CDN or scrubbing provider, enable stricter edge rules on the affected endpoints, add rate limits and challenges, and shed low-priority traffic to protect critical functions.
Block or challenge the abusive patterns (shared headers, source networks, request shapes) with narrowly scoped rules, and confirm the origin is only reachable through the protective layer.
Relax emergency rules gradually while watching latency and error rates, restore normal capacity, and verify checkout and other critical flows against baseline.
Record the timeline and signals, tune alert thresholds, add the endpoint-level limits permanently, update the runbook and check whether an extortion demand or a distraction for another intrusion accompanied the event.
Check yourself
1. Latency and errors are climbing on one search endpoint while link utilisation is only slightly above normal and thousands of sources each send a few requests. What is the most likely situation?
2. Why does blocking individual source IP addresses work poorly against a large DDoS?
3. Which combination best protects a service from both volumetric and application-layer floods?