What is CCCL in CUDA?
The CUDA C++ Core Libraries, one repo holding Thrust, CUB and libcu++, versioned and shipped together.
The three libraries used to live in three places with three version histories. In 2023 NVIDIA merged them into one repository, https://github.com/NVIDIA/cccl , and they now release together under one version number. You already have them: CCCL ships inside the CUDA toolkit, so a plain #include <cub/cub.cuh> works with no package to install and no -l flag to pass. nvcc puts the headers on the include path itself.
The three layers divide by where you call them. Thrust is the top: STL-shaped host calls like thrust::reduce that run the whole algorithm, launches and all. CUB is the engine underneath, exposed at device scope for whole-array work and at block and warp scope for use inside a kernel you still write. libcu++ is the standard library that works on both sides of the launch: cuda::std:: types plus device primitives like cuda::atomic_ref and cuda::barrier.
Because CCCL rides in the toolkit, the version you compile against is decided by the toolkit unless you deliberately put a newer copy first on the include path. The "Mapping to CTK Versions" table in the repo README gives the full map, and the compatibility rule is one-directional: a newer CCCL works with an older toolkit, but "CCCL is never forward compatible with the CUDA Toolkit" (https://github.com/NVIDIA/cccl , checked 2026-08-30). Pinning an older CCCL than your toolkit ships is the unsupported direction, which surprises people because it is the direction dependency pinning usually goes.
The one real cost is compile time. CCCL is header only, so a single CUB include pulls a large template library into every translation unit that touches it. Day 39 found the 20 second compile cap on Compiler Explorer to be a tighter limit than the run cap for exactly this reason. What the headers buy you is measured on the same page: the library beat the course's hand-written kernels on all three algorithms it replaced.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node. Day 39's program prints the versions it was built against rather than asserting them:
CUDA runtime 12.6 ships Thrust 2.5.0 and CUB 2.5.0
Key width from cuda::std::numeric_limits: 32 bits
Both version numbers come from the library's own macros at compile time, so this is what the 12.6 toolchain on that machine actually resolved to, matching the CCCL 2.5.0 row in the repo's mapping table. The same run priced the three layers against the course's own kernels on the same card: cub::DeviceReduce::Sum 1.38 times faster than day 24's reduction, cub::DeviceScan::ExclusiveSum 2.33 times faster than day 31's scan, and cub::DeviceRadixSort::SortKeys 20.54 times faster than day 35's one-bit-per-pass sort.
Related terms
Where you meet this
- Day 39, when to stop hand-writing, the lesson that owns this term and prints the versions above.
- Day 24, parallel reduction and day 31, prefix sum, the hand-written kernels the libraries are measured against.
- Day 48, kernel fusion, the case where a chain of library calls loses to one kernel of your own.
- /setup/check-cuda-version, because your CCCL version is a fact about your toolkit.
Sources
- NVIDIA/cccl repository, for the "Mapping to CTK Versions" table and the forward-compatibility rule quoted above: https://github.com/NVIDIA/cccl (checked 2026-08-30)
- CCCL documentation: https://nvidia.github.io/cccl/unstable/ (checked 2026-08-29)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.