Buffer Overflows
Also known as: Buffer overrun, Out-of-bounds write, Stack overflow (memory), Memory corruption
A memory-safety flaw in which code writes more data into a fixed-size buffer than it can hold, corrupting adjacent memory; it is prevented by bounds-checked code, memory-safe languages and compiler and operating-system mitigations.
How it works
A buffer overflow is a flaw in which a program writes more data into a fixed-size block of memory than the block can hold. In languages such as C and C++ the runtime does not check the size of a write for you, so the extra data lands in whatever memory sits next to the buffer. The weakness is catalogued as CWE-120 (copy without checking size) and, more broadly, as CWE-787 (out-of-bounds write).
Conceptually, the memory next to a buffer often holds other variables and control data that the program relies on to keep running correctly. When an unchecked write spills into it, the program may crash, behave unpredictably, or, in the worst case, be steered into running something its author never intended. This is why the flaw matters well beyond a simple crash, and why defenders treat it as a code-execution risk rather than a stability bug.
The root cause is a missing length check between the amount of data arriving and the space reserved for it. It shows up wherever input meets a fixed buffer: network protocol parsers, file format readers, command-line and environment handling, and legacy string-copy helpers such as strcpy, strcat, sprintf and gets. Code review looks for exactly this pairing of an unbounded copy and a fixed-size destination.
Defence is layered. The strongest control is to avoid the problem class with a memory-safe language. Where C or C++ must be used, prefer bounds-checked APIs and safer string types, validate input length before copying, and build with stack canaries, ASLR, DEP/NX and fortified libraries so that a missed bug is much harder to turn into compromise. Fuzz testing and static analysis in CI find the remaining defects before release.
Walk through it
- 1Read the vulnerable copy
- 2Name the unsafe call
- 3Read the signals, pick the fix
- Scope every unbounded copy
- Fix, harden the build and fuzz
A pre-release review flags the gateway's message parser. It copies a device name from a network message into a local buffer. Read the function as a reviewer and note whether anything limits how much data is copied.
1/* Handle a HELLO message from a badge reader. */2void handle_hello(const char *msg) {3 char device_name[32];4 strcpy(device_name, msg); /* copy the name as received */5 register_device(device_name);6}Spot it
- Repeated segmentation faults or access-violation crashes in one service, with respawn counts above its baseline.
- Stack-protector aborts such as 'stack smashing detected' and heap-corruption aborts in system or application logs.
- EDR or OS memory-protection alerts: blocked execution from non-executable memory, or control-flow integrity violations.
- Crash dumps whose faulting address or call stack sits inside a parser or string-handling function.
- Network messages to one service that are far longer than its protocol normally allows, seen in length-aware sensors or logs.
Linux host log
2026-10-11T10:02:11Z gateway-01 kernel: gateway[4127]: segfault at 0 ip 00007f3a in libgateway.so
2026-10-11T10:02:12Z gateway-01 gateway[4140]: *** stack smashing detected ***: terminatedEDR
ts=2026-10-11T10:02:12Z host=gateway-01 process=gateway event=memory_protection action=blocked reason=exec_from_nonexec_region
ts=2026-10-11T10:02:40Z host=gateway-01 process=gateway event=crash count_1h=9 baseline_1h=0splSplunk: services with repeated crashes or stack-protector aborts
index=os (sourcetype=syslog OR sourcetype=linux_secure) ("segfault" OR "stack smashing detected")
| stats count by host, process
| where count > 5Compare with the service's normal crash rate and correlate with EDR memory-protection alerts before escalating.
kqlKQL: repeated process crashes on one host
DeviceEvents
| where ActionType in ("ExploitGuardNonMicrosoftSignedBlocked", "AsrExecutableOfficeAppsBlocked", "ProcessCrash")
| summarize crashes=count() by DeviceName, InitiatingProcessFileName
| where crashes > 5Action type names vary by sensor; adapt to the crash and memory-protection events your EDR emits.
Stop it
Avoid the problem class: prefer memory-safe languages and safe string types
Write new parsing and network-facing code in a memory-safe language, or use safe containers such as std::string and std::vector in C++. Where C must remain, ban unbounded functions such as gets, strcpy and sprintf in coding standards and enforce it in static analysis.
Bound every copy and validate input length first
Use APIs that take the destination size, such as snprintf and bounded copy helpers, check the input length against the buffer before copying, and reject or truncate deliberately. Treat every length that comes from outside the process as untrusted.
Build with mitigations and test with fuzzing
Enable stack canaries, ASLR, DEP/NX and FORTIFY_SOURCE so a missed bug is far harder to abuse, then fuzz parsers and run sanitizers in CI to find defects before release. Mitigations reduce impact; they never replace the fix.
Bound the copy to the destination size
Vulnerable
char device_name[32];
strcpy(device_name, msg);Hardened
char device_name[32];
if (strlen(msg) >= sizeof(device_name)) {
return ERR_NAME_TOO_LONG; /* reject over-long input */
}
snprintf(device_name, sizeof(device_name), "%s", msg);The size check rejects input that does not fit, and the bounded call can never write past the destination.
Compiler and linker hardening flags (GCC/Clang)
Hardened
cc -O2 -D_FORTIFY_SOURCE=2 -fstack-protector-strong -fPIE -pie \
-Wl,-z,relro,-z,now -Wformat -Wformat-security -o gateway src/*.cCanaries, position-independent code, fortified libc calls and full RELRO make surviving defects harder to turn into compromise.
- Ban gets, strcpy, strcat and sprintf in coding standards and enforce the ban with static analysis in CI.
- Prefer memory-safe languages or safe string and container types for new network-facing and file-parsing code.
- Compile with stack canaries, PIE/ASLR, FORTIFY_SOURCE and full RELRO; keep DEP/NX enabled on every host.
- Run address and undefined-behaviour sanitizers in test builds and fuzz every parser that handles external input.
- Run services as an unprivileged account with sandboxing or seccomp so a missed flaw has a small blast radius.
- Alert on crash-rate anomalies and EDR memory-protection events, and patch third-party libraries promptly.
If it already happened
Isolate or rate-limit the affected service, block the offending sources at the edge, and if exposure is likely take the feature offline while the code fix ships.
Replace every unbounded copy on the affected path with bounds-checked code, patch or rebuild with hardening flags, and review host and EDR telemetry to scope whether any process was compromised.
Redeploy the fixed build, rotate secrets the service held if compromise is possible, and verify the crash rate returns to baseline.
Add a static-analysis rule and review checklist item against unbounded copies, add the parser to the fuzzing corpus, and keep the crash and memory-protection detections permanently.
Check yourself
1. What is the root cause of a classic buffer overflow?
2. Which change is the best code-level fix for an unbounded string copy in C?
3. What is the role of stack canaries, ASLR and DEP/NX?