Skip to content

Reading Crash Reports: Signals, Backtraces and Core Dumps

Triage native crashes like a defender: what SIGSEGV and SIGABRT mean, glibc abort messages, gdb and core dumps, ASan reports, WinDbg, and how to prioritise memory-safety crashes.

Published on 6 min read

When a native program crashes, the first minutes of analysis decide whether the crash is filed as noise or recognised as a memory-safety bug. This guide is a practical triage workflow for defenders, developers and incident responders: what the signal tells you, how to read glibc and sanitizer messages, how to pull a useful backtrace from a core dump, and how to prioritise crashes by severity. It assumes the memory model from how a process lays out memory.

Step 1: read the signal

On Linux, a crash is almost always the delivery of a fatal signal. The signal narrows the cause immediately.

SignalUsual meaningMemory-safety relevance
SIGSEGVAccess to an unmapped page, or a permission violation (write to read-only, execute non-executable), or a shadow-stack violation under CETHigh: invalid pointer, overflow, use-after-free
SIGBUSMisaligned access on strict architectures, or access to a truncated memory-mapped fileMedium: often a file or alignment issue
SIGABRTThe program called abort(): failed assert, glibc heap check, stack canary, FORTIFY check, sanitizerHigh when raised by glibc or the stack protector
SIGILLIllegal instruction: ud2 from a compiler trap (__builtin_trap, UBSan trap mode, CFI), or execution of garbageHigh if unexpected: possible control-flow corruption
SIGFPEInteger division by zero or overflow in divisionLow to medium: usually an input-validation bug
SIGTRAPBreakpoint or trap instructionDebugging, or a deliberate trap in hardened code

The kernel log records segfaults with the faulting address and instruction pointer:

server[4182]: segfault at 0 ip 000055d0c3a1b2f4 sp 00007ffd2c1e8a90 error 4 in server[55d0c3a1a000+5000]

segfault at 0 is a NULL dereference. On x86, the error field is the page-fault error code, printed in hexadecimal: 0x1 means the page was present (a permission violation rather than a missing page), 0x2 a write, 0x4 user mode, and 0x10 an instruction fetch. So error 4 is a user-mode read of an unmapped page, error 6 a user-mode write to an unmapped page, error 7 a user-mode write to a read-only page, and error 15 (0x15) a user-mode attempt to execute a non-executable page.

Step 2: read the abort message

When the C library or the compiler's instrumentation detects corruption, it prints a message to stderr before aborting. These messages are precise indicators:

MessageWhat detected itWhat it usually means
*** stack smashing detected ***: terminatedStack protector (__stack_chk_fail)A stack buffer overflow overwrote the canary
*** buffer overflow detected ***: terminatedFORTIFY_SOURCEA checked memcpy, strcpy, sprintf, etc. exceeded the destination size
free(): double free detected in tcache 2glibc tcacheThe same pointer was freed twice
free(): invalid pointerglibcfree was given a pointer not returned by malloc, or metadata was corrupted
malloc(): corrupted top sizeglibcA heap overflow corrupted the top chunk's size
malloc(): unaligned tcache chunk detectedglibc safe-linkingA free-list pointer was corrupted
corrupted size vs. prev_sizeglibcChunk headers are inconsistent, typically from an overflow
*** %n in writable segment detected ***FORTIFY_SOURCEA format-string bug with a user-controlled format

Heap messages indicate where corruption was detected, which is often long after it happened. Do not debug the free call that aborted; reproduce the crash under ASan, which stops at the original bad write. See heap corruption and use-after-free and stack buffer overflows for the underlying bugs.

Step 3: get a core dump

A core dump is a snapshot of the process's memory and registers at the moment of the crash. Make sure you will have one:

# Allow cores for this shell and its children
ulimit -c unlimited

# Where do cores go?
cat /proc/sys/kernel/core_pattern
# |/usr/lib/systemd/systemd-coredump ...   -> managed by systemd-coredump

# On systemd systems
coredumpctl list server
coredumpctl info server          # metadata, signal, backtrace summary
coredumpctl debug server         # opens gdb on the most recent core

Core dumps contain process memory, which may include secrets and personal data. Handle them like any other sensitive incident artefact: restrict access, set retention, and avoid copying them to shared tickets.

Step 4: triage in gdb

With the core and the matching binary and debug information:

gdb ./server core.4182
(gdb) bt                        # backtrace of the crashing thread
(gdb) info registers rip rsp rbp
(gdb) x/6i $rip                 # instructions at the crash site
(gdb) frame 2                   # select a frame
(gdb) info locals               # locals in that frame (needs debug info)
(gdb) thread apply all bt       # every thread, for concurrency bugs
(gdb) info proc mappings        # memory map (live process); 'info files' for a core

What to look for:

  • Where is rip? Inside your code or a known library is normal. An address that is not inside any mapping, or inside a data region, suggests the instruction pointer itself was corrupted.
  • What was accessed? x/i $rip shows the faulting instruction; the register it dereferences tells you which pointer was bad.
  • Does the backtrace make sense? A backtrace that ends in ?? frames or repeats nonsense addresses suggests stack corruption, missing debug info or missing unwind information.

The debugger extensions GEF and pwndbg add a friendlier context view (registers, stack, code) that many analysts prefer for this kind of inspection.

Step 5: reproduce under sanitizers

Once you have the input or a way to trigger the crash, rebuild with -fsanitize=address,undefined -g -O1 and reproduce. ASan reports the first invalid access with allocation and free stacks, which turns an ambiguous crash into a diagnosis. The sanitizers guide explains every part of the report. If the crash came from a fuzzer, the input is already on disk; minimise it first as described in fuzzing with libFuzzer and AFL++.

Windows crashes, briefly

On Windows the equivalents are access violations (0xC0000005), heap corruption (0xC0000374), stack buffer overrun / fail-fast (0xC0000409, raised by /GS checks and other fast-fail conditions) and CFG violations. Crash dumps come from Windows Error Reporting or procdump. In WinDbg:

!analyze -v        automated analysis: exception, faulting frame, bucket ID
k                  stack backtrace
r                  registers
!heap -p -a <addr> page heap details (enable with gflags /p /enable app.exe /full)

Page heap (gflags) plays a role similar to ASan's redzones, placing allocations next to guard pages so overflows fault at the exact instruction.

Step 6: prioritise

A defender triaging hundreds of crashes needs a quick, evidence-based severity estimate. The questions below are the same ones tools such as the exploitable gdb plugin or ClusterFuzz's security severity heuristics ask. They inform patch priority; they do not replace a proper root-cause analysis.

EvidenceSuggested priorityReasoning
NULL or near-NULL readLow to mediumUsually denial of service only
Out-of-bounds read, heap or stackMedium to highInformation disclosure, can undermine ASLR
Out-of-bounds write, heap or stackHighMemory corruption with integrity impact
Use-after-free or double freeHighType confusion, frequently exploitable
Stack canary, FORTIFY or glibc heap abortHighProves corruption occurred, even though the mitigation stopped it
Faulting rip outside code, CET/CFI violationCriticalControl flow was already corrupted
Crash reachable from untrusted network inputRaise one levelAttack surface matters as much as bug type

Keep the rule simple: writes beat reads, attacker-reachable beats local, and anything that corrupted control data gets fixed first.

Deduplication and bisection

Large campaigns produce many crashes with the same root cause. Group them by bug type plus the top few frames of the stack in your own code (ignoring allocator and libc frames). For regressions, git bisect run with a script that builds and replays the crashing input identifies the commit that introduced the bug, which usually points straight at the faulty change.

Checklist

  • Record signal, faulting address, instruction pointer and the kernel's error code for every crash.
  • Treat glibc heap, stack-protector and FORTIFY aborts as confirmed memory corruption.
  • Keep debug information for every release so cores from stripped binaries can be symbolised.
  • Reproduce with ASan and UBSan before guessing at the root cause.
  • Prioritise by access type, reachability and whether control data was corrupted; deduplicate by top frames.