← Glossary
CUDA glossaryLibraries
CC 7.5

What is Thrust in CUDA?

An STL-shaped library of parallel algorithms over device data, where thrust::reduce replaces a kernel you would otherwise write.

Thrust is the top layer of CCCL, and it is shaped like the C++ standard library on purpose. You hand an algorithm a range and it does everything: picks the kernels, sizes the launches, allocates its scratch. The whole of a device-side sum is one host line:

float sum = thrust::reduce(thrust::device, d_in, d_in + n, 0.0f);

You never see a block size. That is the trade: maximum convenience, minimum control. A call like this returns a value to the host, so it synchronizes, and it allocates its own scratch every time. When you need the result to stay on the device, or you want to own the temp storage, the same algorithms sit one layer down in CUB with more knobs and more ceremony.

The first thing that goes wrong for new users is the execution policy. Pass a raw device pointer with no thrust::device and nothing warns you; the documentation says plainly that "Like the STL, Thrust permits this usage and it will dispatch the host path of the algorithm" (https://docs.nvidia.com/cuda/archive/12.2.0/thrust/index.html , checked 2026-08-30), and the host path then dereferences device memory on the CPU. The second thing is the belief that a generic library must be slower than your kernel. It is not generic where it counts: CUB, underneath, carries a tuning policy per architecture, so the code that runs on a T4 is not the code that runs on an H100.

What a whole-array call cannot do is run inside a kernel you already wrote, and it cannot fuse with the work either side of it, so three library calls can cost two extra round trips through global memory where one kernel of yours pays none. One more honest caveat: a floating-point reduction's answer depends on summation order, so do not expect bitwise-identical results run to run. Determinism is its own topic.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), Thrust 2.5.0, built with nvcc -O3 -arch=sm_75. Captured 2026-08-30 on the project's verification node. Day 39 reduced the same 16,776,605 floats three ways in one process, checking every result:

implementation time (ms) GB/s vs hand
hand, day 24 version 4 0.349 192.2 1.00
thrust::reduce 0.277 242.2 1.26
cub::DeviceReduce::Sum 0.254 264.4 1.38

The hand row is day 24's version 4, the fourth rung of the reduction ladder, not day 25's final version; day 39 explains that choice and leaves porting the last rungs as its exercise. The gap between the two library rows is what Thrust's convenience costs here: its reduce returns the value to the host, so it synchronizes and allocates scratch per call, while CUB writes to a device pointer you own. Which one you want depends on where the answer is needed, and that difference is a real property of the interfaces rather than a measurement artifact.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-08-30, and this entry stays a draft until a named author and a different named reviewer sign it.