← Glossary
CUDA glossaryHardware
CC 7.5

What is PCIe to a CUDA program?

The bus between host memory and the GPU, roughly an order of magnitude slower than the GPU's own memory.

Every byte your program moves between host and device crosses PCIe, and the bus's throughput is set by its generation and lane count, not by anything your code does. The T4 is a PCIe Gen3 x16 card (https://www.nvidia.com/en-us/data-center/tesla-t4/ , checked 2026-09-01), and on the project's node the best measured host-to-device rate was 12.2 GB/s from pinned memory. The same program measured the same data being read on the card at 244.8 GB/s effective. That factor of twenty, measured on one machine in one binary, is the number every design decision about copies comes back to.

Three consequences follow. First, minimize crossings, not just bytes: day 9's classic finding is a kernel 24x faster than the CPU losing end to end because the round trip cost more than the compute. Second, when you must cross, cross well: pageable copies ran at 2.8 to 4.4 GB/s on this node against pinned's 9.3 to 12.2, so the allocation choice is worth up to 4x before any overlap. Third, never compute across the bus by accident: a kernel dereferencing mapped host memory pays PCIe rates per access, and day 53 measured that mistake at 9.93x the device-resident time.

The bus also shapes multi-GPU work: peer traffic without NVLink routes over PCIe at these same rates, which is the free-tier reality day 91 measures on Kaggle's T4 pair.

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 53 measured both sides of the bus and the card in one program:

path rate
host to device, pageable, 256 MiB 4.4 GB/s
host to device, pinned, 256 MiB 12.2 GB/s
on-device read, same kernel's data resident 244.8 GB/s effective
kernel reading over the bus (zero-copy) 24.7 GB/s effective, 9.93x slower

The ratios are the entry: the bus at its measured best delivers a twentieth of the card's own bandwidth, and the naive pageable path a fiftieth. The zero-copy row is the honest nuance; overlapped streaming reads can beat the raw copy rate (24.7 against 12.2) while still losing badly to residency, which is why "copy once, compute resident" is the default this course teaches.

Diagram

timeline-host-device, preset bus-vs-card: one lane for PCIe transfers and one for on-device traffic, bars scaled to the measured rates so the width difference is the argument.

Alt text: "The same 64 MiB drawn to scale: crossing PCIe takes an order of magnitude longer than moving it inside the card."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.