CUDA ERROR
initialization error: CUDA error 3
CUDA error 3 means the runtime could not build a usable context. On a T4, fork() after CUDA initialization failed only in the child.
The CUDA runtime tried to build a usable device context for this process and could not.
| Enum | cudaErrorInitializationError |
| Code | 3 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
initialization error - Two near neighbours that mean something different:
driver shutting down(4), which you get when the process is already exiting, andinvalid device context(201). If you are seeing 4, your CUDA call is running from a static destructor aftermainreturned, and the fix is to free device resources before you return rather than to debug the driver.
The measured cause: fork()
The repro initialises CUDA in the parent, forks, and allocates in both halves. From code/errors/initialization-error/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:
cudaFree(0) in the parent -> 0 (cudaSuccess: no error)
cudaMalloc in the fork()ed child -> 3 (cudaErrorInitializationError: initialization error)
cudaFree(0) in the parent -> 0 (cudaSuccess: no error)
cudaMalloc in the parent after wait -> 0 (cudaSuccess: no error)
Four lines, and the shape of them is the lesson. The parent works before the fork and works again after reaping the child. The child, which inherited a fully initialised CUDA state, fails on its first real call. A CUDA context belongs to the process that created it, and fork copies the address space without copying the driver's side of that relationship, so the child holds a handle to something it does not own.
Nothing is wrong with the machine here. nvidia-smi would be perfectly happy throughout, which is why this error sends so many people to reinstall a driver that was never broken.
Cause 1: a fork after CUDA was initialised
The measured case above, and the one to check first if your program uses processes at all. On Linux, Python's multiprocessing defaults to the fork start method, so importing a library that touches CUDA before you create a Pool produces exactly this.
cudaFree(nullptr); // initialises the context
pid_t pid = fork();
if (pid == 0) { cudaMalloc(&d, 4); } // 3, in the child
Fix: use the spawn start method, or do not touch CUDA before forking. Initialise the device inside each worker after the fork, never before it.
Cause 2: the driver is broken or mismatched
The same mess that produces the NVML version mismatch, seen from inside your application instead of from nvidia-smi.
Fix: prove it before you chase it. Run nvidia-smi first; if that fails too, this page is not your page and the driver is. If nvidia-smi is healthy and your process still gets 3, go back to cause 1.
Cause 3: a CUDA call after main returned
A static or global object whose destructor frees device memory. By then the runtime has torn itself down, and the usual code is 4, driver shutting down, rather than 3.
Fix: free device resources explicitly at the end of main, and keep CUDA handles out of objects with static storage duration.
Confirm which one you have
Make cudaFree(0) the first CUDA call in main and check its return value. It forces context creation at a line you chose, so the failure appears in your code instead of inside whichever library happened to touch the GPU first. Our repro uses exactly that call as its probe, in both the parent and the child, which is what makes the transcript above readable at all.
Then, if the probe succeeds and a later call fails, count your processes. A success followed by a failure in a child is cause 1 and nothing else.
Prevention
- Decide where the context is created and make it one explicit line early in the program. An implicit context created by a library import is the reason this error is usually reported from a stack you did not write.
- If the program forks, initialise CUDA after the fork, in each child.
Related errors
no CUDA-capable device is detected(100): there was nothing to build a context on.invalid device ordinal(101): the device you named does not exist in this process.out of memory(2): the context that this error could not create is itself what the free-memory gap on that page pays for.
The lesson
Day 59 covers host threads feeding the GPU, which is where the process and context rules start to matter; day 3 owns the install, and the environment half is at /setup/install-cuda. The full code table is on the error hub.
Written 2026-09-01. Transcript from code/errors/initialization-error/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.