CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

out of memory: cudaMalloc returned CUDA error 2

CUDA error 2 comes from cudaMalloc, not PyTorch. On a T4, a failed allocation left free memory unchanged and size_t arithmetic did not wrap.

The driver could not find a contiguous block of device memory the size you asked for, so cudaMalloc returned without giving you a pointer.

Enum cudaErrorMemoryAllocation
Code 2
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Scope. This page is the C and C++ side: cudaMalloc returned 2. If your paste begins torch.OutOfMemoryError: CUDA out of memory. Tried to allocate, that is a different error with a different fix and it is not covered here.

Strings a reader might paste

  • Runtime: out of memory
  • Driver API and Numba: CUDA_ERROR_OUT_OF_MEMORY
  • Enum name, which is what you get from cudaGetErrorName: cudaErrorMemoryAllocation. There is no cudaErrorOutOfMemory in the runtime enum, whatever a rival error index lists; CUDA_ERROR_OUT_OF_MEMORY is the driver API's name for the same value.

What the run actually shows

From code/errors/out-of-memory/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:

cudaMemGetInfo                     -> 0 (cudaSuccess: no error)
free 15527116800 bytes, total 15637086208 bytes
cudaMalloc(100 GB)                 -> 2 (cudaErrorMemoryAllocation: out of memory)
cudaMemGetInfo after failure       -> 0 (cudaSuccess: no error)
free 15527116800 bytes, total 15637086208 bytes
n*n*4 with n=400000000 wraps to 640000000000000000

Two things in that transcript are worth more than the error itself. First, the free figure is identical before and after the failed request: 15527116800 both times. A failed cudaMalloc is clean, it takes nothing and leaks nothing, so if your free memory is falling across a loop the allocations that are succeeding are the problem, not the one that failed. Second, free is already below total on the first line, before this program allocated anything at all. That gap is the CUDA context, which day 9 measured costing real time as well as real memory.

Cause 1: the allocation genuinely does not fit

The measured case above. The program asks for 100ull << 30 bytes on a card whose total is 15637086208 bytes, and the driver says no.

void* p = nullptr;
report("cudaMalloc(100 GB)", cudaMalloc(&p, 100ull << 30));

Fix: check the return value of every cudaMalloc and compare the request against cudaMemGetInfo before you make it. An unchecked cudaMalloc hands a null pointer to a kernel, and that arrives later as an illegal memory access was encountered from a line that has nothing to do with the allocation.

Cause 2: the size arithmetic, and the overflow that did not happen

The folklore version of this cause says an int overflow turns a 4 GB request into a wrapped, nonsense number. Our repro tests the arithmetic in size_t and prints the result:

size_t n = 400000000;      // 400M floats
size_t bytes = n * n * 4;

The transcript's last line reads n*n*4 with n=400000000 wraps to 640000000000000000, and that value is exactly n * n * 4 computed correctly. Nothing wrapped. In 64-bit size_t the product is representable, so the program asks for a truthful 640 quadrillion bytes and gets a truthful refusal. The overflow story is real only when the intermediate type is narrower than the result:

int n = 40000;
size_t bytes = n * n * sizeof(float);   // n * n is int arithmetic first

Fix: widen before you multiply, size_t bytes = (size_t)n * n * sizeof(float);. The rule is that the cast goes on the first operand, not on the assignment, because the assignment happens after the damage.

Cause 3: something else is holding the memory

A second process, a notebook kernel nobody shut down, or the display. The context gap visible on the first transcript line is the benign version of this; a stale python process is the annoying one.

Fix: nvidia-smi --query-compute-apps=pid,used_memory --format=csv names the processes and what each holds.

Confirm which one you have

Print cudaMemGetInfo immediately before the allocation and again after it fails, exactly as the repro does. Three readings tell you three different stories: a request larger than total is cause 1 and no amount of freeing will help; a request that is a wild number is cause 2; a free far below total with a small request is cause 3, and the process list names the culprit.

compute-sanitizer --tool memcheck --leak-check full ./a.out reports device allocations that were never freed, which is how you find the loop that is eating the card (https://docs.nvidia.com/compute-sanitizer/ComputeSanitizer/index.html , checked 2026-08-29).

Prevention

  • Allocate once outside the loop and reuse the buffer. When the sizes genuinely vary, stream-ordered allocation and a memory pool are the supported answer, and day 55 measures what that change buys.
  • Check every allocation. Day 6's CUDA_CHECK macro exists for exactly this class of silent failure.

Related errors

The lesson

Day 55 replaces per-iteration cudaMalloc with a pool and measures the difference; day 9 prices the context this error's free-memory gap. Day 5 owns the vector-add walkthrough that breaks allocation on purpose. The full code table is on the error hub, and a machine with no working toolkit belongs at /setup/install-cuda.


Written 2026-09-01. Transcript from code/errors/out-of-memory/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.