← Glossary
CUDA glossaryLibraries
CC 7.5

What is CCCL in CUDA?

The CUDA C++ Core Libraries, one repo holding Thrust, CUB and libcu++, versioned and shipped together.

The three libraries used to live in three places with three version histories. In 2023 NVIDIA merged them into one repository, https://github.com/NVIDIA/cccl , and they now release together under one version number. You already have them: CCCL ships inside the CUDA toolkit, so a plain #include <cub/cub.cuh> works with no package to install and no -l flag to pass. nvcc puts the headers on the include path itself.

The three layers divide by where you call them. Thrust is the top: STL-shaped host calls like thrust::reduce that run the whole algorithm, launches and all. CUB is the engine underneath, exposed at device scope for whole-array work and at block and warp scope for use inside a kernel you still write. libcu++ is the standard library that works on both sides of the launch: cuda::std:: types plus device primitives like cuda::atomic_ref and cuda::barrier.

Because CCCL rides in the toolkit, the version you compile against is decided by the toolkit unless you deliberately put a newer copy first on the include path. The "Mapping to CTK Versions" table in the repo README gives the full map, and the compatibility rule is one-directional: a newer CCCL works with an older toolkit, but "CCCL is never forward compatible with the CUDA Toolkit" (https://github.com/NVIDIA/cccl , checked 2026-08-30). Pinning an older CCCL than your toolkit ships is the unsupported direction, which surprises people because it is the direction dependency pinning usually goes.

The one real cost is compile time. CCCL is header only, so a single CUB include pulls a large template library into every translation unit that touches it. Day 39 found the 20 second compile cap on Compiler Explorer to be a tighter limit than the run cap for exactly this reason. What the headers buy you is measured on the same page: the library beat the course's hand-written kernels on all three algorithms it replaced.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node. Day 39's program prints the versions it was built against rather than asserting them:

CUDA runtime 12.6 ships Thrust 2.5.0 and CUB 2.5.0
Key width from cuda::std::numeric_limits: 32 bits

Both version numbers come from the library's own macros at compile time, so this is what the 12.6 toolchain on that machine actually resolved to, matching the CCCL 2.5.0 row in the repo's mapping table. The same run priced the three layers against the course's own kernels on the same card: cub::DeviceReduce::Sum 1.38 times faster than day 24's reduction, cub::DeviceScan::ExclusiveSum 2.33 times faster than day 31's scan, and cub::DeviceRadixSort::SortKeys 20.54 times faster than day 35's one-bit-per-pass sort.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.