Testing GPU kernels
This page's repo directory holds a suite of four tests: a golden value, a determinism check, a comparison against the CPU, and a linearity property. All four pass a reduction that never reads the last 99 elements of its input.
These are common test mistakes, not made-up examples. This module has shown a harder case already: day 27's racy kernel gave the right answer on 100 runs of 100 and is still wrong. By the end of this page you will know what each test certifies, how to rebuild the suite so it checks every path, and how to make CI report a failure.
What the harness you have been using since day 5 checks
Every program in this course has carried the same testing skeleton since day 5, and it has four parts.
An oracle that shares nothing with the kernel. The reference is a plain CPU loop, written for obvious correctness, accumulating in double where the kernel accumulates in float, with Kahan compensation on sums. The wide accumulator type and the plain loop keep the oracle independent.
An oracle that reuses the kernel's order, precision, or helpers shares the kernel's bugs. Two copies of one mistake agree.
Sizes chosen to make bugs fire. Day 5 used 611 elements because 611 is 13 x 47: no block size divides it, so the bounds guard runs on every launch instead of never. A suite that tests only sizes the block divides tests the one configuration where indexing bugs cannot show. Day 14 made the same point about block size: its race was invisible at 32 threads and produced 17,312 mismatches at 64.
Inputs where a mismatch means an index bug. The first case is always exact: small values, every intermediate representable, tolerance zero, so any disagreement is structural. Tolerance questions are real, and they are asked separately, on cases built for them.
Failures that name the first bad index, as real branches. "Wrong at
512" points at a block; "wrong" points at nothing. The branch returns
EXIT_FAILURE rather than calling assert().
Release builds define NDEBUG. The NDEBUG macro removes asserts. A check
removed from the build that CI runs cannot fail CI.
Error checking on every CUDA call belongs here too, since a test cannot judge a kernel that never launched.
Everything on this page extends that skeleton; nothing replaces it.
Diagram: what each test can reject. Three horizontal bands, each a number line of relative error with the true answer at zero, the gate drawn as a shaded window, and this page's bug marked as a dot at its actual error. Band 1, the golden snapshot: the window is a single point sitting on the dot itself. Caption "gate centred on the bug: rejects the fix, accepts the bug." Band 2, the loose oracle: a window 1e-2 wide around zero, the dot at 6.1e-3, inside. Caption "a 99-element tail is a 6.1e-3 error; a 1e-2 gate waves it through." Band 3, the scaled tolerance: a window 6.1e-5 wide around zero, the same dot far outside. Caption "4 x 2^-23 x sqrt(16,483) = 6.1e-5: room for rounding, none for the bug." Alt text: "Three tolerance windows against one bug. A golden snapshot centres its gate on the bug itself. A one percent gate hides a six-per-thousand error. A scaled gate a hundred times tighter exposes it while still allowing float rounding."
A green test proves the code and the test agree, nothing more
A large passing suite does not prove more by its size alone. A test carries evidence only if some wrong kernel would fail it. The supplied suite contains three common tests that reject nothing.
The snapshot. A "golden" value captured from the kernel under test certifies whatever the kernel did the day someone ran it. If that kernel was wrong, the test now defends the bug: fix the kernel and the suite goes red.
This project's own review history is full of the milder form, a test that recomputes the kernel's formula on the host and checks the two match. Same formula, same mistake, green forever.
The self-comparison. Running the kernel twice and demanding equal answers rejects flaky kernels and nothing else; a deterministic wrong answer passes every time, and day 27 showed even a racy kernel passing a hundred straight. Determinism is worth one test. It is not correctness.
The tolerance is too loose. A real oracle with rtol set to 1e-2 "to be safe with floats" cannot see any bug smaller than one percent, and most bugs worth catching are smaller than that. Compute the right tolerance from the arithmetic, as shown two sections down.
The supplied suite's fourth test is honest and still blind: linearity, checked only at 1,024 elements, where four whole blocks leave no tail. A sound property at the wrong size sees nothing; the case ladder is part of the test, not decoration around it.
Two suites, two variants, one program that grades the difference
Full program in
code/day66-testing/testing.cu.
What it commits to matters more than what it computes.
The bug is supplied and labelled, with the fix beside it. The kernel is the block tree from day 24; both variants launch it unchanged and differ only in the grid arithmetic:
// Deliberate bug (day 66): integer division rounds down, so the last
// n % kThreadsPerBlock elements are never covered by any block. At 1,024
// the remainder is zero and the variant is exactly right; at 611 it
// silently ignores 99 elements. sumGuarded below differs by one
// expression. Below 256 elements this grid is zero blocks and the launch
// itself fails, which is why the size ladder here stops at 611.
static float sumDropTail(const float* d_in, size_t n, float* d_partials,
float* h_partials) {
const int blocks = static_cast<int>(n / kThreadsPerBlock);
return foldPartials(d_in, n, blocks, d_partials, h_partials);
}
The thing under test is the launch plus the kernel, not the kernel body alone; this bug lives entirely in host code.
Every expected value is derivable by hand. The input is
(i % 97) * 0.5f: every value a multiple of 0.5, every total far below
2^24, so all float arithmetic in the program is exact. The buggy sum at
611 elements is 11,815.5 (two whole blocks, elements 0 through 511)
against a true 14,171:
// Tautological (test 1 of the supplied suite): this constant was captured
// from the kernel under test, so the test can only fail when the kernel
// changes. It does not certify the sum of 611 elements; it certifies
// whatever the drop-tail variant did the day someone ran it. 11,815.5 is
// the sum of elements 0..511 of this input, the two whole blocks the
// buggy grid covers. The true 611-element sum is 14,171.
constexpr float kGolden611 = 11815.5f;
static bool testGoldenSnapshot(SumFn sum, const float* d_in, float* d_partials,
float* h_partials) {
const float got = sum(d_in, kOddElems, d_partials, h_partials);
return checkRel("golden snapshot, n=611", got, kGolden611, 0.0);
}
The code computes the tolerance. The hardened oracle's limit scales with the reduction length, because rounding error grows with the number of terms folded:
// A sum of K float terms accumulates rounding error. For a tree or any
// shuffled order the per-step errors are uncorrelated, so the bound grows
// like sqrt(K), not K; the factor 4 is slack for a different but valid
// summation order. One fixed rtol is wrong at both ends: 1e-5 fails
// correct code at K in the millions, and a value loosened until every
// size passes stops being a test.
static float scaledRtol(size_t k) {
const float eps = 1.1920929e-7f; // 2^-23, float spacing at 1.0
const float grown = 4.0f * eps * std::sqrt(static_cast<float>(k));
return grown > kBaseRtol ? grown : kBaseRtol;
}
At 16,483 elements that gives 6.1e-5. It is two orders of magnitude tighter than this page's bug, yet allows a correct kernel summing in any order. Day 68 inverts the question and makes the reduction bitwise reproducible, at which point determinism stops being a property test and becomes a design constraint.
The hardened suite tests permutation invariance. It uses a fixed permutation so failures are easy to reproduce:
// Permutation invariance: a sum must not care where its elements sit.
// The permutation is a deliberate one, not a random shuffle: the first
// and last 99 elements trade places, so whatever a kernel does to the
// tail of the buffer, it now does it to different values. A random
// shuffle would also work but would make the failure a different number
// every seed, and a test you cannot reproduce exactly is a test you
// cannot debug.
static bool testSwapEnds(SumFn sum, const float* d_in, const float* d_swap,
float* d_partials, float* h_partials) {
const float plain = sum(d_in, kMaxElems, d_partials, h_partials);
const float moved = sum(d_swap, kMaxElems, d_partials, h_partials);
return checkRel("swap-ends permutation, n=16483", moved, plain,
scaledRtol(kMaxElems));
}
Linearity, sum(2x) equals 2 times sum(x), stays in the hardened suite and is expected to pass the buggy variant too, since dropping a tail is itself linear. The program prints that pass rather than hiding it.
A property test rejects one class of bug. You should know which class each test can reject.
Compute Sanitizer is a CI step, if you gate its exit code
The second program,
code/day66-testing/ci_gate.cu,
covers what the suites cannot. Its squarePastEnd kernel produces a
correct visible answer and also writes 4 bytes one element past its
output allocation (deliberate, labelled). No host-side check can see the
write, and the program exits 0.
On day 6's evidence an overrun this short is expected not to fault. The drop-tail bug is numerically wrong and memory-clean; this one is the reverse. A suite needs a check for each type of failure.
Compute Sanitizer checks memory use, but its
default exit code can hide a finding. Its --error-exitcode option defaults to 0, and
the manual says what the option is for: "The exit code Compute Sanitizer
will return if the original application succeeded but the tool detected
that errors were present. This is meant to allow Compute Sanitizer to be
integrated into automated test suites."
(https://docs.nvidia.com/compute-sanitizer/ComputeSanitizer/index.html ,
checked 2026-09-01.)
Without the flag, the tool prints its report and exits with the application's
own code. A CI step that checks $? then stays green despite an invalid write.
Use these commands:
compute-sanitizer --destroy-on-device-error kernel \
--error-exitcode 1 ./ci_gate
compute-sanitizer --leak-check full --error-exitcode 1 ./testing
The kernel-scoped destruction is part of this controlled experiment: it keeps
the context alive so ci_gate finishes with its own zero exit code and the
presence or absence of --error-exitcode is the only changed variable. Without
it, the default context teardown makes the checked synchronize fail. The
--leak-check full option turns unfreed allocations into errors as well.
Results
Run on the project's Tesla T4 (driver 595.84, CUDA 12.6 V12.6.85) on
2026-09-01. Transcripts in code/day66-testing/evidence/. Four
predictions held; the fifth held on the finding and died on the
bookkeeping.
CUDA 13.0 repeated every suite result exactly. Its corrected sanitizer
experiment used --destroy-on-device-error kernel for both ci_gate runs:
the same invalid 4-byte write was reported, the no-flag command exited 0, and
the --error-exitcode 1 command exited 1. The fixed context-lifetime control
therefore verifies the intended exit-code distinction without erasing the
original CUDA 12.6 finding.
| Suite | vs drop-tail variant | vs guarded variant |
|---|---|---|
| supplied (4 tests) | 4 of 4 PASS | n/a |
| hardened oracle ladder | FAIL at 611 and 16,483, PASS at 1,024 | all PASS |
| swap-ends permutation | FAIL | PASS |
| linearity | PASS | PASS |
| golden snapshot | PASS | FAIL |
Held. The supplied suite went 4 for 4 against the buggy variant: golden snapshot, determinism, loose oracle and linearity all PASS. A green suite certifying a kernel that drops its tail.
Held to the digit. The hardened oracle failed the buggy variant at n=611 (rel 1.662e-1 against a 1.179e-5 gate) and at n=16,483 (rel 6.111e-3 against 6.122e-5), and passed it at n=1,024, where the size divides evenly and the bug cannot show. The pass is the point: a test that only runs at 1,024 proves nothing about 611.
Held. The swap-ends permutation failed the buggy variant, got 393,106 against want 393,018, a difference of exactly 88, which is 2.239e-4 relative against a 6.122e-5 gate. The guarded variant's two orderings came back bitwise equal.
Held. Linearity passed both variants, and the golden snapshot failed the guarded variant (got 14,171.0, want 11,815.5): the tautology running backwards, a test that fails the fix because it was written from the bug.
Held after fixing the experiment. Memcheck over
./testing, the wrong-answer program, ended inERROR SUMMARY: 0 errorsand exit 0, and so did the leak-check pass with zero bytes leaked. The original CUDA 12.6 command let memcheck tear down the context, so the application's checked synchronize exited 1 even without--error-exitcode.The corrected CUDA 13.0 capture adds
--destroy-on-device-error kernel: both runs report the invalid 4-byte write, the no-flag run exits 0, and the flagged run exits- That isolates the option the lesson is testing.
The two checks cover different faults. The suite accepts ./testing and the
sanitizer agrees, yet the kernel is wrong.
The suite accepts ./ci_gate, but the sanitizer rejects it. Use both checks.
Run it yourself
You need a shell and a supported CUDA GPU. The README beside the code has both build lines and the four sanitizer commands in batch form.
compute-sanitizer ships with the toolkit and needs no root and no
performance counters, for the reason day 61
gives. NVIDIA's Colab-targeted tutorial notebooks run it
as a teaching step (https://github.com/NVIDIA/accelerated-computing-hub/blob/main/tutorials/cuda-cpp/notebooks/03.02-Kernels/03.02.04-Dev-Tools.ipynb ,
checked 2026-09-01), though this project has not reproduced that
first-hand. Compiler Explorer cannot run the full exercise because reading
tool exit codes needs a shell.
Exercise
Work in code/day66-testing/starter/testing.cu: the buggy reduction
plus only the four supplied tests. Classify the four (one sentence each:
what wrong kernel would this reject?), harden the suite per the TODO
markers until a test fails, name where the missing elements live from
the failure output alone, and only then fix the one expression.
Time: 30 to 40 minutes. Submit: your hardened testing.cu, the
four sentences, and one more naming which supplied test breaks after the
fix and why.
Check: the worked testing.cu one directory up is its own harness:
it exits 0 only if the supplied suite passes the buggy variant 4 for 4,
the hardened oracle catches it at 611, the hardened suite passes the
fixed variant, and the golden snapshot rejects the fix; on any other
outcome it prints a GATE: line naming the broken claim and returns
EXIT_FAILURE. Your hardened starter should fail before your fix with
the oracle naming 611 and 16,483, and pass after it except for the
snapshot.
Hint 1
For each test, describe a kernel that is wrong and still passes it. If you cannot, the test has teeth. Which of the four could any wrong-but-deterministic kernel pass?
Hint 2
Once the oracle fails at 611 and 16,483 but not 1,024, ask what the two failing sizes share: both are 99 mod 256. Which line of host code turns n into a number of blocks, and what does integer division do to a remainder?
Solution
The classification: the golden snapshot and the determinism check compare the kernel to itself, directly or through a captured value; the loose oracle is real but gated a hundred times too wide; linearity is sound and runs only at the one size in the file where no tail exists.
The fix is one expression, in host code, not in the kernel:
- const int blocks = static_cast<int>(n / kThreadsPerBlock);
+ const int blocks =
+ static_cast<int>((n + kThreadsPerBlock - 1) / kThreadsPerBlock);
After it, the oracle ladder passes all three sizes with zero observed error, swap-ends compares bitwise equal, and the golden snapshot fails, off by the 99-element tail it was capturing all along. Delete it, or recompute it from the oracle.
Judge a test by the wrong kernels it can reject, never by whether it passes. A test that no wrong result can fail checks nothing.
Pitfalls
Your tests recompute the kernel's formula on the host. Refactors break them, bugs do not, which is backwards. An oracle must be independent: different order, wider accumulator, no shared helpers.
Your golden values came from the kernel under test. They freeze today's behaviour, bugs included, and the failure arrives when someone fixes the kernel and the suite goes red in the fix's review. A snapshot is legitimate only when its source outranks the code under test, like a verified oracle or a slower reference implementation.
A fixed rtol of 1e-5 fails your correct million-element reduction. The opposite failure to this page's loose gate, and the usual reaction, loosening until green, lands you in that one. Scale the gate with the reduction length and print the arithmetic next to the verdict.
Your CI runs the sanitizer and stays green through real errors.
--error-exitcode defaults to 0, so the tool reports the error and then
returns the application's own exit code. Pass --error-exitcode 1 on
every sanitizer step, and add --leak-check full so unfreed allocations
gate too. Day 61 covers reading what the tool reports.
Every size in your suite is a multiple of your block size. Then no test ever runs the tail path, which is where launch arithmetic breaks. Day 5 chose 611 on purpose; keep one ragged size per remainder class your launch can produce, and one block-exact size for contrast.
Go deeper
- Compute Sanitizer User Manual, in particular the memcheck tool and the command line options table: https://docs.nvidia.com/compute-sanitizer/ComputeSanitizer/index.html (checked 2026-09-01)
- NVIDIA's accelerated-computing-hub dev-tools notebook, the sanitizer run as a teaching step in a hosted notebook: https://github.com/NVIDIA/accelerated-computing-hub/blob/main/tutorials/cuda-cpp/notebooks/03.02-Kernels/03.02.04-Dev-Tools.ipynb (checked 2026-09-01)
- GPU Puzzles, whose per-thread read and write counting is the strongest grading idea in the corpus this course drew on: https://github.com/srush/GPU-Puzzles (checked 2026-08-29)
- The harness contract this course's graded days follow: /reference/harness
Next
Day 67 gives the code you can now test a real build: days 43 and 44 as a CMake library, with an answer to which SM architectures to ship. Day 68 picks up this page's loose end: the scaled tolerance exists because summation order moves a float answer, and day 68 pins the order so the tolerance can be zero.