What is Thrust in CUDA?
An STL-shaped library of parallel algorithms over device data, where thrust::reduce replaces a kernel you would otherwise write.
Thrust is the top layer of CCCL, and it is shaped like the C++ standard library on purpose. You hand an algorithm a range and it does everything: picks the kernels, sizes the launches, allocates its scratch. The whole of a device-side sum is one host line:
float sum = thrust::reduce(thrust::device, d_in, d_in + n, 0.0f);
You never see a block size. That is the trade: maximum convenience, minimum control. A call like this returns a value to the host, so it synchronizes, and it allocates its own scratch every time. When you need the result to stay on the device, or you want to own the temp storage, the same algorithms sit one layer down in CUB with more knobs and more ceremony.
The first thing that goes wrong for new users is the execution policy. Pass a raw device pointer with no thrust::device and nothing warns you; the documentation says plainly that "Like the STL, Thrust permits this usage and it will dispatch the host path of the algorithm" (https://docs.nvidia.com/cuda/archive/12.2.0/thrust/index.html , checked 2026-08-30), and the host path then dereferences device memory on the CPU. The second thing is the belief that a generic library must be slower than your kernel. It is not generic where it counts: CUB, underneath, carries a tuning policy per architecture, so the code that runs on a T4 is not the code that runs on an H100.
What a whole-array call cannot do is run inside a kernel you already wrote, and it cannot fuse with the work either side of it, so three library calls can cost two extra round trips through global memory where one kernel of yours pays none. One more honest caveat: a floating-point reduction's answer depends on summation order, so do not expect bitwise-identical results run to run. Determinism is its own topic.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Thrust 2.5.0, built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node. Day 39 reduced the same 16,776,605 floats three ways in one process, checking every result:
| implementation | time (ms) | GB/s | vs hand |
|---|---|---|---|
| hand, day 24 version 4 | 0.349 | 192.2 | 1.00 |
thrust::reduce |
0.277 | 242.2 | 1.26 |
cub::DeviceReduce::Sum |
0.254 | 264.4 | 1.38 |
The hand row is day 24's version 4, the fourth rung of the reduction ladder, not day 25's final version; day 39 explains that choice and leaves porting the last rungs as its exercise. The gap between the two library rows is what Thrust's convenience costs here: its reduce returns the value to the host, so it synchronizes and allocates scratch per call, while CUB writes to a device pointer you own. Which one you want depends on where the answer is needed, and that difference is a real property of the interfaces rather than a measurement artifact.
Related terms
Where you meet this
- Day 39, when to stop hand-writing, the lesson that owns this term and produced the table above.
- Day 24, parallel reduction and day 25, the ladder the measurement is honest about.
- Day 40, the PageRank capstone, which wants your own CSR kernels first and a Thrust version second, on the same graph.
- Day 48, kernel fusion, the case against calling the library for everything.
- Colab setup, enough card to run the day 39 comparison yourself.
Sources
- Thrust documentation: https://nvidia.github.io/cccl/unstable/thrust/ (checked 2026-08-29)
- Thrust guide, for the host-path dispatch rule quoted above: https://docs.nvidia.com/cuda/archive/12.2.0/thrust/index.html (checked 2026-08-30)
- NVIDIA/cccl repository, for versioning and the toolkit mapping: https://github.com/NVIDIA/cccl (checked 2026-08-30)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.