Reading Crash Reports: Signals, Backtraces and Core Dumps
Triage native crashes like a defender: what SIGSEGV and SIGABRT mean, glibc abort messages, gdb and core dumps, ASan reports, WinDbg, and how to prioritise memory-safety crashes.
When a native program crashes, the first minutes of analysis decide whether the crash is filed as noise or recognised as a memory-safety bug. This guide is a practical triage workflow for defenders, developers and incident responders: what the signal tells you, how to read glibc and sanitizer messages, how to pull a useful backtrace from a core dump, and how to prioritise crashes by severity. It assumes the memory model from how a process lays out memory.
Step 1: read the signal
On Linux, a crash is almost always the delivery of a fatal signal. The signal narrows the cause immediately.
| Signal | Usual meaning | Memory-safety relevance |
|---|---|---|
SIGSEGV | Access to an unmapped page, or a permission violation (write to read-only, execute non-executable), or a shadow-stack violation under CET | High: invalid pointer, overflow, use-after-free |
SIGBUS | Misaligned access on strict architectures, or access to a truncated memory-mapped file | Medium: often a file or alignment issue |
SIGABRT | The program called abort(): failed assert, glibc heap check, stack canary, FORTIFY check, sanitizer | High when raised by glibc or the stack protector |
SIGILL | Illegal instruction: ud2 from a compiler trap (__builtin_trap, UBSan trap mode, CFI), or execution of garbage | High if unexpected: possible control-flow corruption |
SIGFPE | Integer division by zero or overflow in division | Low to medium: usually an input-validation bug |
SIGTRAP | Breakpoint or trap instruction | Debugging, or a deliberate trap in hardened code |
The kernel log records segfaults with the faulting address and instruction pointer:
server[4182]: segfault at 0 ip 000055d0c3a1b2f4 sp 00007ffd2c1e8a90 error 4 in server[55d0c3a1a000+5000]
segfault at 0 is a NULL dereference. On x86, the error field is the page-fault error code, printed in hexadecimal: 0x1 means the page was present (a permission violation rather than a missing page), 0x2 a write, 0x4 user mode, and 0x10 an instruction fetch. So error 4 is a user-mode read of an unmapped page, error 6 a user-mode write to an unmapped page, error 7 a user-mode write to a read-only page, and error 15 (0x15) a user-mode attempt to execute a non-executable page.
Step 2: read the abort message
When the C library or the compiler's instrumentation detects corruption, it prints a message to stderr before aborting. These messages are precise indicators:
| Message | What detected it | What it usually means |
|---|---|---|
*** stack smashing detected ***: terminated | Stack protector (__stack_chk_fail) | A stack buffer overflow overwrote the canary |
*** buffer overflow detected ***: terminated | FORTIFY_SOURCE | A checked memcpy, strcpy, sprintf, etc. exceeded the destination size |
free(): double free detected in tcache 2 | glibc tcache | The same pointer was freed twice |
free(): invalid pointer | glibc | free was given a pointer not returned by malloc, or metadata was corrupted |
malloc(): corrupted top size | glibc | A heap overflow corrupted the top chunk's size |
malloc(): unaligned tcache chunk detected | glibc safe-linking | A free-list pointer was corrupted |
corrupted size vs. prev_size | glibc | Chunk headers are inconsistent, typically from an overflow |
*** %n in writable segment detected *** | FORTIFY_SOURCE | A format-string bug with a user-controlled format |
Heap messages indicate where corruption was detected, which is often long after it happened. Do not debug the free call that aborted; reproduce the crash under ASan, which stops at the original bad write. See heap corruption and use-after-free and stack buffer overflows for the underlying bugs.
Step 3: get a core dump
A core dump is a snapshot of the process's memory and registers at the moment of the crash. Make sure you will have one:
# Allow cores for this shell and its children
ulimit -c unlimited
# Where do cores go?
cat /proc/sys/kernel/core_pattern
# |/usr/lib/systemd/systemd-coredump ... -> managed by systemd-coredump
# On systemd systems
coredumpctl list server
coredumpctl info server # metadata, signal, backtrace summary
coredumpctl debug server # opens gdb on the most recent core
Core dumps contain process memory, which may include secrets and personal data. Handle them like any other sensitive incident artefact: restrict access, set retention, and avoid copying them to shared tickets.
Step 4: triage in gdb
With the core and the matching binary and debug information:
gdb ./server core.4182
(gdb) bt # backtrace of the crashing thread
(gdb) info registers rip rsp rbp
(gdb) x/6i $rip # instructions at the crash site
(gdb) frame 2 # select a frame
(gdb) info locals # locals in that frame (needs debug info)
(gdb) thread apply all bt # every thread, for concurrency bugs
(gdb) info proc mappings # memory map (live process); 'info files' for a core
What to look for:
- Where is
rip? Inside your code or a known library is normal. An address that is not inside any mapping, or inside a data region, suggests the instruction pointer itself was corrupted. - What was accessed?
x/i $ripshows the faulting instruction; the register it dereferences tells you which pointer was bad. - Does the backtrace make sense? A backtrace that ends in
??frames or repeats nonsense addresses suggests stack corruption, missing debug info or missing unwind information.
The debugger extensions GEF and pwndbg add a friendlier context view (registers, stack, code) that many analysts prefer for this kind of inspection.
Step 5: reproduce under sanitizers
Once you have the input or a way to trigger the crash, rebuild with -fsanitize=address,undefined -g -O1 and reproduce. ASan reports the first invalid access with allocation and free stacks, which turns an ambiguous crash into a diagnosis. The sanitizers guide explains every part of the report. If the crash came from a fuzzer, the input is already on disk; minimise it first as described in fuzzing with libFuzzer and AFL++.
Windows crashes, briefly
On Windows the equivalents are access violations (0xC0000005), heap corruption (0xC0000374), stack buffer overrun / fail-fast (0xC0000409, raised by /GS checks and other fast-fail conditions) and CFG violations. Crash dumps come from Windows Error Reporting or procdump. In WinDbg:
!analyze -v automated analysis: exception, faulting frame, bucket ID
k stack backtrace
r registers
!heap -p -a <addr> page heap details (enable with gflags /p /enable app.exe /full)
Page heap (gflags) plays a role similar to ASan's redzones, placing allocations next to guard pages so overflows fault at the exact instruction.
Step 6: prioritise
A defender triaging hundreds of crashes needs a quick, evidence-based severity estimate. The questions below are the same ones tools such as the exploitable gdb plugin or ClusterFuzz's security severity heuristics ask. They inform patch priority; they do not replace a proper root-cause analysis.
| Evidence | Suggested priority | Reasoning |
|---|---|---|
| NULL or near-NULL read | Low to medium | Usually denial of service only |
| Out-of-bounds read, heap or stack | Medium to high | Information disclosure, can undermine ASLR |
| Out-of-bounds write, heap or stack | High | Memory corruption with integrity impact |
| Use-after-free or double free | High | Type confusion, frequently exploitable |
| Stack canary, FORTIFY or glibc heap abort | High | Proves corruption occurred, even though the mitigation stopped it |
Faulting rip outside code, CET/CFI violation | Critical | Control flow was already corrupted |
| Crash reachable from untrusted network input | Raise one level | Attack surface matters as much as bug type |
Keep the rule simple: writes beat reads, attacker-reachable beats local, and anything that corrupted control data gets fixed first.
Deduplication and bisection
Large campaigns produce many crashes with the same root cause. Group them by bug type plus the top few frames of the stack in your own code (ignoring allocator and libc frames). For regressions, git bisect run with a script that builds and replays the crashing input identifies the commit that introduced the bug, which usually points straight at the faulty change.
Checklist
- Record signal, faulting address, instruction pointer and the kernel's
errorcode for every crash. - Treat glibc heap, stack-protector and FORTIFY aborts as confirmed memory corruption.
- Keep debug information for every release so cores from stripped binaries can be symbolised.
- Reproduce with ASan and UBSan before guessing at the root cause.
- Prioritise by access type, reachability and whether control data was corrupted; deduplicate by top frames.