What is PCIe to a CUDA program?
The bus between host memory and the GPU, roughly an order of magnitude slower than the GPU's own memory.
Every byte your program moves between host and device crosses PCIe, and the bus's throughput is set by its generation and lane count, not by anything your code does. The T4 is a PCIe Gen3 x16 card (https://www.nvidia.com/en-us/data-center/tesla-t4/ , checked 2026-09-01), and on the project's node the best measured host-to-device rate was 12.2 GB/s from pinned memory. The same program measured the same data being read on the card at 244.8 GB/s effective. That factor of twenty, measured on one machine in one binary, is the number every design decision about copies comes back to.
Three consequences follow. First, minimize crossings, not just bytes: day 9's classic finding is a kernel 24x faster than the CPU losing end to end because the round trip cost more than the compute. Second, when you must cross, cross well: pageable copies ran at 2.8 to 4.4 GB/s on this node against pinned's 9.3 to 12.2, so the allocation choice is worth up to 4x before any overlap. Third, never compute across the bus by accident: a kernel dereferencing mapped host memory pays PCIe rates per access, and day 53 measured that mistake at 9.93x the device-resident time.
The bus also shapes multi-GPU work: peer traffic without NVLink routes over PCIe at these same rates, which is the free-tier reality day 91 measures on Kaggle's T4 pair.
Measured
Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 53 measured both sides of the bus and the card in one program:
| path | rate |
|---|---|
| host to device, pageable, 256 MiB | 4.4 GB/s |
| host to device, pinned, 256 MiB | 12.2 GB/s |
| on-device read, same kernel's data resident | 244.8 GB/s effective |
| kernel reading over the bus (zero-copy) | 24.7 GB/s effective, 9.93x slower |
The ratios are the entry: the bus at its measured best delivers a twentieth of the card's own bandwidth, and the naive pageable path a fiftieth. The zero-copy row is the honest nuance; overlapped streaming reads can beat the raw copy rate (24.7 against 12.2) while still losing badly to residency, which is why "copy once, compute resident" is the default this course teaches.
Diagram
timeline-host-device, preset bus-vs-card: one lane for PCIe transfers and one for on-device traffic, bars scaled to the measured rates so the width difference is the argument.
Alt text: "The same 64 MiB drawn to scale: crossing PCIe takes an order of magnitude longer than moving it inside the card."
Related terms
Where you meet this
- Day 53, pinned memory, the lesson that owns this term and measured every row above.
- Day 9, why your GPU code looks slower than your CPU, the bus's most-asked question.
- Day 49, the measured roofline, where the card-side 244.8 GB/s ceiling comes from the same class of measurement.
- Day 54, double buffering, hiding the bus behind compute.
Sources
- NVIDIA Tesla T4 product page, for the PCIe Gen3 x16 interface: https://www.nvidia.com/en-us/data-center/tesla-t4/ (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.