Coverage-Guided Fuzzing with libFuzzer and AFL++
Write a fuzz harness, build it with sanitizers, run libFuzzer and AFL++, manage corpora and dictionaries, and run continuous fuzzing in CI with OSS-Fuzz or ClusterFuzzLite.
Fuzzing feeds a program large numbers of automatically generated inputs and watches for crashes. Coverage-guided fuzzers make this dramatically more effective: they instrument the target, keep any input that reaches new code, and mutate those inputs further. Combined with sanitizers, which turn silent memory corruption into immediate crashes, coverage-guided fuzzing is the most productive way to find memory-safety bugs in parsers, decoders and protocol handlers. This guide shows how to write a harness, run it under libFuzzer and AFL++, and keep it running in CI.
How coverage guidance works
+-----------+ mutate +-----------+
| corpus | ------------> | new input |
+-----------+ +-----+-----+
^ |
| keep if new v run instrumented target
| coverage +---------------+
+------------------ | coverage map | --> crash? save it
+---------------+
The compiler inserts lightweight counters on edges between basic blocks. After each execution, the fuzzer compares the coverage map with everything it has seen. Inputs that reach new edges join the corpus and become parents for future mutations. Over time the corpus climbs deeper into the program's logic, far beyond what random inputs could reach.
Writing a harness
A harness is a small function that takes a byte buffer and passes it to the code under test. The libFuzzer convention, also understood by AFL++, honggfuzz and OSS-Fuzz, is:
#include <stddef.h>
#include <stdint.h>
#include "config_parser.h"
int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {
struct config cfg;
if (config_parse(&cfg, data, size) == 0) {
config_free(&cfg);
}
return 0; /* non-zero return values are reserved */
}
Good harnesses share a few properties:
| Property | Why |
|---|---|
| Deterministic | The same input must produce the same behaviour, or crashes will not reproduce |
| Fast | Aim for thousands of executions per second; avoid disk, network and sleeps |
| No global state leaks | Free what you allocate, reset globals, or leaks and state carry-over mask bugs |
| Narrow entry point | Fuzz the parser directly, not the whole application around it |
| Exercises the API fully | If there is a serializer, round-trip parse and serialize to reach more code |
For structured APIs that take several arguments, FuzzedDataProvider (a header shipped with Clang) splits the input into typed values:
#include <fuzzer/FuzzedDataProvider.h>
extern "C" int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size) {
FuzzedDataProvider fdp(data, size);
auto mode = fdp.ConsumeIntegralInRange<int>(0, 3);
auto key = fdp.ConsumeRandomLengthString(64);
auto value = fdp.ConsumeRemainingBytes<uint8_t>();
kv_store_put(mode, key.c_str(), value.data(), value.size());
return 0;
}
Running libFuzzer
libFuzzer is built into Clang:
clang -g -O1 -fsanitize=fuzzer,address,undefined \
config_parser.c fuzz_config.c -o fuzz_config
mkdir -p corpus
cp tests/fixtures/*.conf corpus/ # seed with real, valid examples
./fuzz_config corpus/ -max_len=4096 -jobs=4 -workers=4
Useful options:
| Option | Purpose |
|---|---|
corpus/ | Directory of seed inputs; new interesting inputs are written back here |
-max_len=N | Upper bound on input size |
-dict=file.dict | Tokens to splice into inputs (keywords, magic numbers) |
-jobs=N -workers=N | Parallel runs |
-max_total_time=S | Stop after S seconds, useful in CI |
-runs=0 corpus/ | Run the corpus once without fuzzing, to check for regressions |
When a crash occurs, libFuzzer prints the sanitizer report and writes the input as crash-<sha1>. Reproduce it by passing the file as an argument: ./fuzz_config crash-<sha1>.
Seeds and dictionaries
A good seed corpus is the single biggest accelerator. Use real, valid files from your test suite; the fuzzer then mutates from meaningful structure instead of discovering it from scratch. For text or token-based formats, a dictionary helps the fuzzer get past keyword comparisons:
# config.dict
kw_section="[section]"
kw_include="include"
kw_true="true"
magic="\x7fCFG"
Running AFL++
AFL++ is a community-maintained successor to AFL with many improvements. It can build the same LLVMFuzzerTestOneInput harness, or fuzz a program that reads from a file or stdin.
# Build with AFL++'s LLVM instrumentation and ASan
export AFL_USE_ASAN=1
afl-clang-fast -g -O1 -fsanitize=fuzzer config_parser.c fuzz_config.c -o fuzz_config_afl
# Or instrument a whole program that reads a file
CC=afl-clang-lto ./configure && make
afl-fuzz -i corpus/ -o findings/ -x config.dict -- ./fuzz_config_afl
A few AFL++ features worth knowing:
- LTO instrumentation (
afl-clang-lto) assigns collision-free edge IDs at link time, which improves coverage accuracy. - Persistent mode runs many inputs in one process, the same idea libFuzzer uses; building a libFuzzer-style harness with
-fsanitize=fuzzerunderafl-clang-fastgives you this automatically. - CmpLog (
-cwith a separately built CmpLog binary) captures comparison operands so the fuzzer can solve magic values and checksums that pure mutation rarely hits. - Parallel fuzzing: run one main instance (
-M) and several secondaries (-S) sharing the same output directory.
Triage and minimisation
A campaign quickly accumulates crashing inputs, many of them duplicates. Shrink and deduplicate before filing bugs:
# libFuzzer: minimise a crashing input while keeping the crash
./fuzz_config -minimize_crash=1 -runs=10000 crash-<sha1>
# libFuzzer: minimise a corpus to the smallest set with the same coverage
./fuzz_config -merge=1 corpus_min/ corpus/
# AFL++: minimise one input, and the corpus
afl-tmin -i findings/default/crashes/id:000000* -o crash.min -- ./fuzz_config_afl
afl-cmin -i findings/default/queue -o corpus_min -- ./fuzz_config_afl
Group crashes by the sanitizer report's top frames and bug type, then fix the root cause, add the minimised input to your regression tests, and re-run the corpus. The crash triage guide covers deduplication and severity assessment in more depth.
Continuous fuzzing
Fuzzing pays off most when it runs continuously and new code is fuzzed as it lands.
| Option | What it is | Fits |
|---|---|---|
| OSS-Fuzz | Google's free continuous fuzzing service for critical open-source projects, built on ClusterFuzz | Open-source projects that are accepted into the program |
| ClusterFuzzLite | A lightweight version that runs in your own CI (GitHub Actions, GitLab and others) | Any project, including private code |
| CI smoke fuzzing | Run each harness for a few minutes on every pull request with -max_total_time | Catching regressions early |
| Self-hosted ClusterFuzz | The full platform on your own infrastructure | Large organisations with many targets |
A minimal CI step can be as simple as building the fuzzers with sanitizers, running the stored corpus with -runs=0 to catch regressions, then fuzzing for a fixed time budget and uploading any crash-* files as artifacts.
Beyond C and C++
The same approach works in memory-safe languages, where it finds panics, infinite loops, excessive allocation and logic bugs:
- Rust:
cargo fuzz(libFuzzer-based), plusarbitraryfor structured inputs. - Go: native fuzzing with
func FuzzXxx(f *testing.F)since Go 1.18. - Python: Atheris. Java/JVM: Jazzer.
Checklist
- Write one harness per parser or decoder, targeting the narrowest useful entry point.
- Build fuzzers with ASan and UBSan; seed them with real, valid inputs and add a dictionary for token-heavy formats.
- Track coverage; when it plateaus, improve the harness or seeds rather than only adding CPU.
- Minimise and deduplicate crashes, fix the root cause, and keep every crash input as a regression test.
- Run fuzzers continuously with OSS-Fuzz or ClusterFuzzLite and briefly on every pull request.