What is GPU RAM?
The DRAM attached to a GPU, usually GDDR on consumer cards or stacked HBM on datacenter accelerators.
CUDA calls this device memory or global memory depending on whether the subject is allocation or the address space. The physical chips are the card's RAM: GDDR6 on a Tesla T4, GDDR7 on current consumer Blackwell cards, and HBM stacks on accelerators such as A100 and H100. Capacity decides what fits. Width, transfer rate and memory clocks set the published bandwidth ceiling.
The published ceiling is not an application guarantee. Protocol overhead, request mix, clocks and access pattern all reduce the useful rate. A coalesced copy is the clean baseline because it asks for long, adjacent reads and writes; a strided kernel can use the same allocation and reach a small fraction of that number. Caches can also make a kernel appear to exceed a DRAM ceiling when most requests never reach RAM, which is why profiler bytes and compulsory bytes are different columns.
Measured
NVIDIA's current Tesla T4 page publishes 320+ GB/s (https://www.nvidia.com/en-us/data-center/tesla-t4/, checked 2026-09-01). On the project's T4 with driver 595.84 and CUDA 12.6, day 49 measured a coalesced copy ceiling of 245.1 GB/s against that 320 GB/s published ceiling. The same run measured 4.2 MB reaching DRAM for a naive matmul whose compulsory-byte model asked for 1,074.8 MB, DRAM/ask 0.0039. The missing traffic was cache reuse, not impossible RAM bandwidth.
Related terms
Where you meet this
- Day 11, memory coalescing, where one allocation produces six bandwidth rows.
- Day 49, roofline, the profiler-backed DRAM-byte comparison.
- Out of memory, the allocation failure tied to capacity rather than bandwidth.
Sources
- NVIDIA Tesla T4 product page, for the published 320+ GB/s ceiling: https://www.nvidia.com/en-us/data-center/tesla-t4/ (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.