XML Bombs (Entity Expansion DoS)
Also known as: XML bomb, billion laughs, XML entity expansion attack, XML denial of service
A denial-of-service flaw in which an XML parser with no expansion limits lets a tiny document consume huge amounts of CPU and memory; it is prevented by disabling DTDs, enabling secure processing and capping entity expansion, document size and depth.
How it works
An XML bomb is a denial-of-service flaw in how an application configures its XML parser. XML lets a document define entities, named substitutions that the parser expands while reading. Because entities can be defined in terms of other entities, a document can be tiny on the wire yet describe an expansion that grows exponentially in memory once the parser follows every reference.
The root cause is a parser with no limits on expansion, not a bug in business logic. Parsers that process document type definitions and have no cap on entity expansion, total size or nesting depth will happily spend CPU and memory on whatever the document asks. The exposure appears wherever untrusted XML arrives: uploads, SOAP and other web service requests, SVG and office documents, single sign-on messages and configuration imports.
The impact is availability. A handful of small requests can pin processors, exhaust the heap, trigger out-of-memory kills and take a shared service down for every user. Nothing is read or stolen, which is why the flaw is easy to miss in functional testing: ordinary documents parse correctly.
The fix is to remove the capability and bound what remains. Disable document type definitions wherever they are not needed, enable the parser's secure-processing mode so the library enforces its own expansion limits, and cap document size, nesting and parse time. Where the data does not need XML, accept JSON. Request size limits and resource quotas on the parsing service contain the damage if a parser is missed.
Walk through it
- 1Read the parser setup
- 2Name the root cause
- 3Read the signals, pick the fix
- Find every XML entry point
- Harden, bound and verify
A pre-release review flags the partner gateway that accepts XML order files. Read the parsing code as a reviewer and note how the parser is configured and what bounds it applies before it touches the document.
1public Document read(InputStream body) throws Exception {2 DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();3 // no feature flags set: library defaults apply4 dbf.setExpandEntityReferences(true);5 DocumentBuilder builder = dbf.newDocumentBuilder();6 return builder.parse(body); // untrusted partner XML, no size cap7}Spot it
- Parser CPU or memory spikes on small request bodies to an upload or web service endpoint that normally handles larger documents cheaply.
- Parse timeouts and out-of-memory kills on the parsing service, clustered from one source.
- WAF matches for XML entity or document type declarations in request bodies that normally carry plain data.
- Repeated container restarts or heap exhaustion that correlate with requests to a single XML endpoint.
- Latency for all users degrading while request volume stays flat, because a few requests consume disproportionate resources.
WAF
ts=2026-10-11T10:02:11Z rule=xml-entity-declaration action=log src=203.0.113.88 host=gateway.contoso-orders.example path=/orders/import
ts=2026-10-11T10:02:19Z rule=xml-entity-declaration action=log src=203.0.113.88 host=gateway.contoso-orders.example path=/orders/importApplication and host telemetry
ts=2026-10-11T10:02:12Z svc=gateway level=error msg="XML parse timeout" body_bytes=812 client=203.0.113.88
ts=2026-10-11T10:02:40Z proc=gateway event=oom_kill note=heap_limit_reached
ts=2026-10-11T10:02:41Z proc=gateway event=cpu_sustained pct=98 note=unexpected_for_request_sizesplSplunk: small bodies that cause parse timeouts
index=app sourcetype=gateway level=error "XML parse timeout" body_bytes<10000
| stats count by client
| where count > 5Tune the threshold to the endpoint's baseline and correlate with WAF matches and host memory or restart events before escalating.
kqlKQL: WAF hits for XML entity rules
AzureDiagnostics
| where Category == "ApplicationGatewayFirewallLog" and ruleGroup_s has "XML"
| summarize hits=count() by clientIp_s, requestUri_s
| where hits > 10Stop it
Disable DTDs and entity expansion in every XML parser
Turn off document type definitions entirely where they are not needed, and enable the parser's secure-processing mode so the library enforces its own expansion limits. Set this explicitly per parser instead of trusting library defaults, and enforce it in review and static analysis.
Set entity-expansion, size and depth limits
Where some expansion must remain, cap the number of entity expansions, total expanded size, element nesting and parse time. Apply the same limits to every component that parses XML, including SVG, office-document and SAML handling.
Prefer simpler formats and enforce request size limits
Accept JSON where the data does not need XML features. Enforce a maximum request body size at the edge and in the application, and give the parsing service memory and CPU quotas so a missed flaw cannot take down neighbours.
Disable DTDs, enable secure processing and cap size
Vulnerable
DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
DocumentBuilder builder = dbf.newDocumentBuilder();
return builder.parse(body);Hardened
DocumentBuilderFactory dbf = DocumentBuilderFactory.newInstance();
dbf.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
dbf.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
dbf.setXIncludeAware(false);
dbf.setExpandEntityReferences(false);
DocumentBuilder builder = dbf.newDocumentBuilder();
return builder.parse(new LimitedInputStream(body, MAX_ORDER_BYTES));With document type declarations disallowed there is nothing to expand. Secure processing and the size cap bound anything else. Apply equivalent settings to every parser type in use.
Use a parser that refuses entity processing
Vulnerable
from lxml import etree
tree = etree.parse(body)Hardened
from lxml import etree
parser = etree.XMLParser(resolve_entities=False, load_dtd=False, no_network=True, huge_tree=False)
tree = etree.parse(body, parser)Or use the defusedxml package, which wraps the standard parsers with safe defaults and rejects entity declarations.
- Disable DTD processing explicitly on every XML parser, factory and reader in the codebase.
- Turn on the parser's secure-processing mode and keep XML libraries patched; re-check after upgrades.
- Cap request body size at the edge and again in the application before parsing.
- Limit element nesting depth and parse time, and fail closed on timeout.
- Give the parsing service memory and CPU limits so one bad request cannot starve other workloads.
- Prefer JSON or another simpler format for new interfaces that do not need XML features.
- Add a static-analysis rule that flags parsers created without the hardening flags.
If it already happened
Block the offending sources, enforce a tight request size limit and rate limit on the XML endpoint, and restart or scale the affected service while the fix ships.
Disable DTDs and enable secure processing with expansion limits on every affected parser, and review host and application logs to scope which services were degraded.
Redeploy the hardened service, run regression tests that confirm entity declarations are rejected and oversized documents fail fast, and communicate any availability impact to affected parties.
Add a static-analysis rule and a review checklist item for parser configuration, inventory every place XML is parsed, and keep the resource-signal detections above permanently.
Check yourself
1. What is the root cause of an XML bomb denial of service?
2. Which control is the primary fix for XML entity-expansion denial of service?
3. Which signal best suggests runaway entity expansion on an XML endpoint?