What is libcu++, the CUDA standard library?
The CUDA standard library: cuda::std types that work on both sides, plus device primitives like cuda::atomic_ref, cuda::barrier and cuda::pipeline.
It is the bottom layer of CCCL, and it has two halves. cuda::std:: is the parts of std:: that make sense in device code, so the same cuda::std::numeric_limits or cuda::std::pair compiles in a kernel and on the host, and a template written against them works on both sides of the launch. The second half has no std:: counterpart because a CPU does not need one: atomics that name their scope, barriers that can be arrived at and waited on separately, and pipelines that drive asynchronous copies.
The scope is the idea worth stopping on. cuda::atomic_ref<int, cuda::thread_scope_device> is C++20's atomic_ref plus a promise about who else touches that address: thread_scope_block says only threads in this block, thread_scope_device says any thread on the GPU, thread_scope_system says the CPU may be involved too. The narrower the scope, the less the hardware must do to make the operation visible. The trap is that a too-narrow scope is a lie the card may not punish: day 27 ran a cross-block tally with block scope, the wrong scope for cross-block communication, and it produced the correct count on all 100 runs while remaining wrong by the memory model's contract. Passing tests on one card is not what correct means for atomics.
cuda::barrier is a __syncthreads() you can split into arrive and wait, doing useful work between the two, and cuda::pipeline schedules cp.async copies against compute. Both have a software path everywhere and a hardware-accelerated form, the async barrier, that needs compute capability 8.0, which a T4 does not have; days 74 and 75 are built on them.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), CCCL 2.5.0, built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node. Day 39's part 4 kernel accumulates per-block sums into one counter through a device-scope cuda::atomic_ref, and prints the check:
Key width from cuda::std::numeric_limits: 32 bits
odd keys counted on the device: 8388303, host says 8388303
The key width line is cuda::std::numeric_limits<unsigned int>::digits evaluated as a constexpr in the day 39 program, the same expression std::numeric_limits gives on the host, usable on either side of the launch. The count is 16,776,605 keys classified on the GPU with the total agreeing with the host's own pass exactly. On scopes, day 27's tally kernels are the measured companion: the device-scope version was exact on 100 runs of 1024 blocks (lowest count 262144, exactly the expected value), and the deliberately misscoped block-scope version was also exact on all 100 runs, which is the point. The wrong scope did not fail; it is still wrong.
Code
From code/day39-cccl/cccl.cu, the scope stated where the operation happens:
if (threadIdx.x == 0) {
cuda::atomic_ref<int, cuda::thread_scope_device> counter(*total);
counter.fetch_add(blockSum, cuda::memory_order_relaxed);
}
}
Related terms
Where you meet this
- Day 39, when to stop hand-writing, the lesson that owns this term and produced the output above.
- Day 27, the CUDA memory model, where the thread scopes are put on trial and the wrong one passes.
- Day 26, atomics, the plain
atomicAddthese types replace with a stated contract. - Day 28, cooperative groups, the other API for naming who synchronizes with whom.
Sources
- libcu++ documentation, for
cuda::std,cuda::atomic_ref,cuda::barrierandcuda::pipeline: https://nvidia.github.io/cccl/unstable/libcudacxx/ (checked 2026-08-29) - CUDA Programming Guide, "Asynchronous Barriers", for the compute capability 8.0 hardware acceleration line: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/async-barriers.html (checked 2026-08-30)
- CUDA Programming Guide, "Pipelines": https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/pipelines.html (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.