Fuzz Testing#
A fuzz test states a property that must hold for every input, then lets Google FuzzTest search the input domain for a counterexample. Where a unit test pins one input to one expected output, a fuzz test pins the whole domain to an invariant:
// Unit test: one scenario
EXPECT_EQ(algorithm.update(knownInput), expectedOutput);
// Fuzz test: a property of every valid input
void fuzzAlgorithm(int32_t x, int32_t y) {
auto result = algorithm.update(x, y);
EXPECT_GE(result.count, 0);
EXPECT_LE(result.count, x + y);
}
FUZZ_TEST(MySuite, fuzzAlgorithm)
.WithDomains(fuzztest::InRange(0, 1'000), fuzztest::InRange(0, 1'000));
The same FUZZ_TEST serves two roles. In a build without fuzzing
instrumentation it runs as an ordinary parameterized GTest over a small set of
generated inputs, which is quick enough for a normal ctest run. In a build
with FUZZTEST_FUZZING_MODE, it becomes a coverage-guided fuzzer that mutates
inputs to reach new branches.
This page covers running the fuzz tests, writing one for your module, and reproducing a finding.
There are two modules with fuzz tests today. Read them before writing your own:
src/fswAlgorithms/imageProcessing/regionsOfInterest/tests/test_regionsOfInterest_fuzz.cpp– the plain idiom, with each property function followed by itsFUZZ_TEST.src/fswAlgorithms/imageProcessing/centerOfBrightness/tests/– the variant that keeps the property functions and a fake image reader incenterOfBrightnessTestHelpers.hpp, so the_fuzz.cppfile holds only the registrations. It also shows the reference oracle idiom.
Running the Fuzz Tests#
Fuzz targets are built only when XMERA_ENABLE_FUZZTESTS is on. Two presets in
src/CMakePresets.json turn it on:
fuzz-smoke-test– setsXMERA_ENABLE_FUZZTESTS=ON. EachFUZZ_TESTruns as a normal GTest.fuzz-test– inherits the above and addsFUZZTEST_FUZZING_MODE=ON, which builds the coverage instrumentation and the sanitizer.
The first configure with either preset fetches Google FuzzTest through
FetchContent (see src/cmake/XmeraGoogleTest.cmake), thus it takes several
minutes. Later configures use the cached checkout.
As unit tests#
cd src
cmake --preset fuzz-smoke-test
cmake --build ../build --parallel
cd ../build && ctest -C Release -L fuzz-smoke
This finishes in seconds. Use it to confirm a new fuzz test compiles, registers, and passes its generated inputs before you start a long session.
As a coverage-guided fuzzer#
Build with the fuzz-test preset, then run one fuzz test for a duration:
cd src
cmake --preset fuzz-test
cmake --build ../build --parallel
BIN=../build/fswAlgorithms/imageProcessing/regionsOfInterest/tests/test_regionsOfInterest_fuzz
"$BIN" --list_fuzz_tests
"$BIN" --fuzz=RegionsOfInterestFuzz.fuzzRegionIdentification --fuzz_for=120s
--list_fuzz_tests prints one [*] Fuzz test: <Suite.TestName> line for each
registered test and exits. --fuzz= takes one of those names, and accepts a
substring when the substring matches exactly one test. An empty --fuzz= works
only in a binary that holds a single FUZZ_TEST. A name that matches no
FUZZ_TEST stops the binary with a non-zero exit code.
The session prints new coverage as it finds it. On a counterexample it prints the failing input and exits non-zero. See Reproducing a Finding.
Keeping a corpus between sessions#
A corpus is the set of inputs the fuzzer has found interesting, that is, the inputs that reached a branch nothing else reached. Give a session the previous corpus and it continues from that coverage instead of from nothing.
Two environment variables control this. Point both at the same directory:
mkdir -p /tmp/fuzz-corpus/RegionsOfInterestFuzz.fuzzRegionIdentification
FUZZTEST_TESTSUITE_IN_DIR=/tmp/fuzz-corpus/RegionsOfInterestFuzz.fuzzRegionIdentification \
FUZZTEST_TESTSUITE_OUT_DIR=/tmp/fuzz-corpus/RegionsOfInterestFuzz.fuzzRegionIdentification \
"$BIN" --fuzz=RegionsOfInterestFuzz.fuzzRegionIdentification --fuzz_for=60s
Give each fuzz test its own directory. A corpus file holds the serialized
parameters of one FUZZ_TEST, and two tests rarely have the same parameter
signature. Two tests that share a directory each reject the files of the other
with Unexpected intermediate representation.
The fuzzer writes corpus files as it finds them, not at exit, thus the corpus survives a crash or an interrupt.
All targets at once#
.github/scripts/run_long_fuzzers.sh runs every fuzz test in the build tree and
handles the corpus layout:
.github/scripts/run_long_fuzzers.sh ./build ./fuzz-logs --fuzz-for 120s
.github/scripts/run_long_fuzzers.sh ./build ./fuzz-logs --fuzz-for 120s --corpus .fuzztest_corpus
The first argument is the build directory, because the script runs
ctest --show-only=json-v1 -L fuzz there to find the fuzz binaries. It then
asks each binary for its test names with --list_fuzz_tests and runs them one at
a time. The duration applies to each test, not to each binary.
With --corpus it creates <corpus>/<binary>/<Suite.TestName> for each test.
It writes one log for each test to the log directory, and exits non-zero with a
list of the tests that reported a finding.
Writing a Fuzz Test#
File layout#
Put the fuzz test beside the unit tests of the module, named
test_<moduleName>_fuzz.cpp:
src/fswAlgorithms/<category>/<moduleName>/
tests/
CMakeLists.txt
test_<moduleName>.cpp # unit tests
test_<moduleName>_fuzz.cpp # fuzz tests
Do not include the unit test file from the fuzz file. The two build into separate binaries, and the include also gives the unit tests the fuzz labels.
Writing the test file#
A fuzz test file has three parts:
Includes – the algorithm header, GTest, and FuzzTest:
#include "../myAlgorithm.h" #include "gtest/gtest.h" #include <fuzztest/fuzztest.h>
Test functions – each function takes fuzzed parameters and asserts invariants. The function must not return a value; use
EXPECT_*/ASSERT_*macros to check properties:void fuzzMyAlgorithm(int32_t inputA, double inputB) { MyAlgorithm algo; auto result = algo.compute(inputA, inputB); // Assert invariants -- properties that must hold for ALL inputs EXPECT_GE(result.value, 0); EXPECT_NO_THROW(algo.reset()); }
Registration – the
FUZZ_TESTmacro registers the function and specifies input domains:FUZZ_TEST(MyAlgorithmFuzz, fuzzMyAlgorithm) .WithDomains(fuzztest::InRange(0, 1000), // inputA fuzztest::InRange(-1.0, 1.0)); // inputB
Common domain combinators:
fuzztest::InRange(lo, hi)– uniformly sample an integer or float rangefuzztest::OneOf(domain1, domain2, ...)– pick from several sub-domainsfuzztest::Arbitrary<T>()– any value of typeTfuzztest::VectorOf(domain).WithMaxSize(n)– variable-length containers
See the FuzzTest domain reference for the full list.
CMake wiring#
Add a guarded block to the module’s tests/CMakeLists.txt:
if(XMERA_ENABLE_FUZZTESTS)
fuzztest_setup_fuzzing_flags()
add_executable(test_myModule_fuzz
../myModuleAlgorithm.cpp
test_myModule_fuzz.cpp
)
target_include_directories(test_myModule_fuzz PRIVATE "../..")
target_include_directories(test_myModule_fuzz PRIVATE "..")
target_include_directories(test_myModule_fuzz PRIVATE "${CMAKE_SOURCE_DIR}")
target_link_libraries(test_myModule_fuzz PRIVATE
Eigen3::Eigen
fuzztest::fuzztest
fuzztest::fuzztest_gtest_main
)
# The labels come from xmera_label_discovered_tests, not from PROPERTIES. Refer to the
# note on that function in src/cmake/XmeraGoogleTest.cmake.
gtest_discover_tests(test_myModule_fuzz
TEST_LIST myModuleFuzzTests
)
xmera_label_discovered_tests(test_myModule_fuzz myModuleFuzzTests fuzz fuzz-smoke)
endif()
Guard the whole block with
if(XMERA_ENABLE_FUZZTESTS), so the default build does not change.Call
fuzztest_setup_fuzzing_flags()beforeadd_executable. It adds the instrumentation flags to the targets that follow it.Link
fuzztest::fuzztestandfuzztest::fuzztest_gtest_main, notGTest::gtest_main.Apply the labels with
xmera_label_discovered_tests, which needs theTEST_LISTfromgtest_discover_testsand must be called after it in the same directory.PROPERTIES LABELScannot give a test two labels; see Only one of the two labels appears.
The two labels mean different things:
fuzz-smoke– the test is quick enough for a normalctestrun.fuzz– the test is a target for a long fuzzing session.
Give a new test both. Give it fuzz alone when it becomes too slow for a normal
test run.
Choosing Invariants#
Choosing what to assert is the hard part. The usual strategies, with the form each takes in this repository:
No crash. The weakest invariant and still a useful one. A segmentation fault, an out of bounds read that the sanitizer catches, or an unhandled exception on any input is a bug. This costs nothing to get, because calling the algorithm is already part of the test.
Output bounds. EXPECT_GE(result.numberOfPixels, 0).
Conservation. The output cannot exceed a known function of the inputs, for
example EXPECT_LE(result.numberOfPixels, totalInputPixels). Where the
algorithm merges or accumulates, this is usually the strongest cheap invariant
available.
A reference oracle. Write a slow, obviously correct implementation beside the
test and compare against it. This is the strongest invariant, because it pins the
value and not only its range. The center of brightness helpers keep a
referenceUpdate that recomputes the result from the visible pixels, and the
property function compares field by field:
CenterOfBrightnessResult refResult = referenceUpdate(visible, brightnessThreshold, refState);
EXPECT_NEAR(result.centerOfBrightness[0], refResult.centerOfBrightness[0], 1e-9);
EXPECT_EQ(result.pixelsFound, refResult.pixelsFound);
EXPECT_NEAR(result.rollingAverageBrightness, refResult.rollingAverageBrightness, 1e-9);
Idempotence and round trips. Two calls with the same input give the same result; an encode followed by a decode gives the original value.
Avoid an assertion that the code cannot break. An assertion derived from the same expression the algorithm used is a tautology: it passes for every input, adds fuzzing time, and hides the absence of a real check.
Reproducing a Finding#
When the fuzzer finds a counterexample, it prints a === BUG FOUND! block with
the failing arguments, and, unless the test uses a fixture, a
=== Regression test draft block below it:
=================================================================
=== Regression test draft
TEST(RegionsOfInterestFuzz, fuzzRegionIdentificationRegression) {
fuzzRegionIdentification(
1,
0,
...
);
}
Paste that draft into the module’s unit test file. It calls the property function directly, so the assertions come with it, and the case stays covered in the normal test run after you fix the algorithm. This is the whole workflow for most findings.
For an input too large to paste, or one that needs shrinking first, use the
reproducer file. A local run writes no reproducer file unless you ask for one.
Without FUZZTEST_REPRODUCERS_OUT_DIR the session prints
[.] No reproducer output location specified - not writing the reproducer file.
and the input is lost. Set it before the run:
mkdir -p /tmp/fuzz-reproducers
FUZZTEST_REPRODUCERS_OUT_DIR=/tmp/fuzz-reproducers \
"$BIN" --fuzz=RegionsOfInterestFuzz.fuzzRegionIdentification --fuzz_for=120s
The session then prints Reproducer file was dumped at: <path>. Replay that file
against the same test:
FUZZTEST_REPLAY=/tmp/fuzz-reproducers/<file> \
"$BIN" --gtest_filter=RegionsOfInterestFuzz.fuzzRegionIdentification
FUZZTEST_REPLAY also accepts a directory, in which case it replays every file
in it. To shrink the input to the smallest one that still fails:
FUZZTEST_MINIMIZE_REPRODUCER=/tmp/fuzz-reproducers/<file> \
FUZZTEST_REPRODUCERS_OUT_DIR=/tmp/fuzz-reproducers \
"$BIN" --gtest_filter=RegionsOfInterestFuzz.fuzzRegionIdentification
Each smaller failing input is written to the reproducer directory. Replay the smallest one, take its regression test draft, and add that to the unit tests.
Troubleshooting#
The coverage does not increase#
An invariant has value only if the inputs can reach it. Two conditions stop the inputs from reaching the code, and neither is easy to see, because the test keeps reporting PASSED.
A fake that ignores its arguments. A test double can accept a window center and a size, then give the same pixels for every value. The domains for those parameters are then dead. The fuzzer continues to change the values and to store them in each corpus entry, but the result does not change. When you write a fake, make sure each argument changes what the fake gives.
A test that does not call
reset(). Some algorithms compute their window inreset(). Until the test calls it, the algorithm uses the values from its constructor. A test can set an image size of 4096 pixels, skipreset(), and leave the algorithm on its default window, which rejects most of the coordinate domain. If the default window rejects every input, the test always gets an empty result.
Watch the Edges covered count in the session output. If it does not increase
after a few seconds, check the test for these two conditions.
Unexpected intermediate representation#
The corpus directory holds files from a different FUZZ_TEST. Each test
serializes its own parameter signature, thus a test cannot parse another test’s
files and rejects them. Give each fuzz test its own corpus directory. This is what
run_long_fuzzers.sh does with <corpus>/<binary>/<Suite.TestName>.
A ctest label filter removes more than expected#
CTest compares labels with a regular expression that has no anchors. Thus
-LE fuzz excludes every test whose label contains fuzz, which includes
fuzz-smoke. Use an anchored expression, such as -LE '^fuzz$', when you mean
one label only.
Only one of the two labels appears#
gtest_discover_tests(... PROPERTIES LABELS "fuzz;fuzz-smoke") loses the second
label. It sends the properties through a command line and expands them again
without quotation marks, thus LABELS "fuzz;fuzz-smoke" becomes the three tokens
LABELS fuzz fuzz-smoke. That is an odd number of tokens for a property list,
and the second label never becomes a label.
xmera_label_discovered_tests in src/cmake/XmeraGoogleTest.cmake exists for
this. It writes a small CMake file into TEST_INCLUDE_FILES, so the labels are
applied when ctest runs, which is after gtest_discover_tests has found the
test names.
Unit tests appear under the fuzz label#
gtest_discover_tests labels every test in the binary. If the fuzz translation
unit includes the unit test one, the unit tests get the fuzz labels too. Keep the
two translation units separate.
run_long_fuzzers.sh is written to tolerate this: it uses the label only to find
the binaries, then asks each binary for its real fuzz test names with
--list_fuzz_tests.
In CI#
Pull request CI configures the fuzz-smoke-test preset and runs the whole test
suite with no label filter, thus the fuzz tests run beside the unit tests and must
stay fast. A nightly job runs run_long_fuzzers.sh for a longer duration against
a cached corpus and fails on a finding. See
.github/workflows/nightly-long-tests.yml for the detail.