CUDA ERROR
too many resources requested for launch: causes and fix
CUDA error 701 means a launch exceeds SM resources. On a T4, two classic triggers returned error 1 instead; 62 registers at 1024 threads was legal.
Threads per block times registers per thread exceeded the SM's register file, or the block asked for more resources than one SM can ever place, so no block could run.
| Enum | cudaErrorLaunchOutOfResources |
| Code | 701 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- Runtime:
too many resources requested for launch - Driver API and Numba:
CUDA_ERROR_LAUNCH_OUT_OF_RESOURCES - Real learner reports: https://stackoverflow.com/questions/26201172/cuda-too-many-resources-requested-for-launch and https://forums.developer.nvidia.com/t/334995 (both checked 2026-08-29).
What we measured, before the folklore
We tried to produce 701 on purpose and could not, and the two failures we hit instead are the ones you are statistically more likely to be holding. All from code/errors/too-many-resources-requested/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6.
A __launch_bounds__(256) kernel launched with 1024 threads returns error 1, not 701, rejected at the launch and not sticky:
peek after 1024-thread launch -> 1 (cudaErrorInvalidValue: invalid argument)
cudaDeviceSynchronize -> 0 (cudaSuccess: no error)
(The transcript's next line, peek after 256-thread launch -> 1, is the peek-versus-get trap, not a second failure: cudaPeekAtLastError does not clear the stored status, so the legal 256-thread launch appears to fail. cudaGetLastError in the second program reads 0 there. Day 6 measures this distinction.)
A register-heavy kernel at 1024 threads simply ran. ptxas compiled our hog to 62 registers:
ptxas info : Used 62 registers, used 0 barriers, 368 bytes cmem[0]
...
regHog uses 62 registers per thread
cudaGetLastError after 1024-thread launch -> 0 (cudaSuccess: no error)
62 times 1024 is 63,488, under the 64 K (65,536) registers per block and per SM that compute capability 7.5 allows (https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/compute-capabilities.html , checked 2026-08-29), so the launch is legal. To hit 701 through registers you need the product over the file, and on this toolkit ptxas kept our kernel under the line on its own. The inventory's classic "65 registers at 1024 threads" arithmetic is right; getting a real compiler to emit it is the hard part, which is worth knowing when a tutorial hands you a repro that does not repro.
Oversized dynamic shared memory is error 1 here too, not 701. Recaptured in the same evidence file from day 13's program:
53248 bytes of dynamic shared memory, no opt-in: invalid argument
53248 bytes after cudaFuncSetAttribute: no error
So when do you actually get 701?
Per the runtime documentation, when the launch geometry is legal in itself but the resources it implies cannot fit: typically a register product over the file on a build where ptxas was forced high (a -maxrregcount raise, heavy double-precision code, or a different architecture's allocation granularity). We state that from the docs, not from a captured run; on this card and toolkit our attempts landed on error 1 instead, and the page will gain a real 701 transcript when one occurs naturally in the course.
Fixes, in order of preference
- Drop to 512 or 256 threads per block.
- Add
__launch_bounds__(N)so ptxas caps the register count for the launch you intend, and note the measured flip side above: launching wider than the bound is error 1 at launch. - Pass
-maxrregcount, last, because it caps every kernel in the unit and trades spills for occupancy. Day 17 measures that trade.
Confirm it
No sanitizer needed. Compile with nvcc -Xptxas -v and read the Used N registers line (real one quoted above), or query it at run time with cudaFuncGetAttributes as our second repro does. Multiply by your block size and compare against cudaDevAttrMaxRegistersPerBlock. If the product is under the limit and the launch still fails, your error is probably 1 or 9, not 701; check the exact string.
Related errors
invalid configuration argument(9): the geometry itself is illegal (over 1024 threads, zero grid).invalid argument(1): what both of our attempted 701 triggers actually returned on this card.out of memory(2): global memory exhaustion, a different budget entirely.
The lesson
Day 17 owns registers and spills; day 45 explains why the fix that clears this error can also make the kernel slower. The full code table is on the error hub; setup-side failures start at /setup/install-cuda.
Written 2026-09-01. Transcripts from code/errors/too-many-resources-requested/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.