CUDA ERROR
device-side assert triggered: causes and fix
CUDA error 710 names the file, line, block and thread. A T4 capture also shows why a build with -DNDEBUG can pass silently.
An assert() inside a kernel fired, the kernel died, and unlike every other crash on this site the error comes with a message that names the exact thread.
| Enum | cudaErrorAssert |
| Code | 710 |
| Source | https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29) |
Strings a reader might paste
- From the kernel itself, before any runtime error, measured in
code/errors/device-side-assert-triggered/evidence/run-2026-09-01.txton a Tesla T4 (sm_75), driver 595.84, CUDA 12.6:repro.cu:12: void checkedWrite(float *, int): block: [0,0,0], thread: [8,0,0] Assertioni < nfailed. - Runtime:
device-side assert triggered - PyTorch:
RuntimeError: CUDA error: device-side assert triggeredortorch.AcceleratorError: ..., with the async suffix, and on 2.7 and earlier the lineCompile withTORCH_USE_CUDA_DSAto enable device-side assertions.The env var isPYTORCH_USE_CUDA_DSA; the build macro isTORCH_USE_CUDA_DSA; the error names the build macro, which is why so many people set the wrong one (https://github.com/pytorch/pytorch/blob/main/c10/cuda/CUDADeviceAssertionHost.cpp , checked 2026-08-29).
The measured run: one assert line per failing thread, then the code
Our repro launches 32 threads at an 8-element buffer with assert(i < n) in the kernel. All 24 out-of-range threads print, thread by thread, and then the runtime reports 710 and the context is dead:
repro.cu:12: void checkedWrite(float *, int): block: [0,0,0], thread: [8,0,0] Assertion `i < n` failed.
...
repro.cu:12: void checkedWrite(float *, int): block: [0,0,0], thread: [31,0,0] Assertion `i < n` failed.
cudaMalloc (8 floats) -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize -> 710 (cudaErrorAssert: device-side assert triggered)
next call (cudaFree) -> 710 (cudaErrorAssert: device-side assert triggered)
Two things to read off that. The assert message names the file, the line, the block and the thread, so start there, not at the stack trace. And 710 is sticky: the cudaFree after it returns 710 too, so only the first message means anything and only a new process gets CUDA back.
The trap: -DNDEBUG deletes the assert and the bug sails through
Same source, rebuilt with -DNDEBUG, same 24-threads-out-of-range launch. From the same evidence file:
== run of the -DNDEBUG build ==
cudaMalloc (8 floats) -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize -> 0 (cudaSuccess: no error)
next call (cudaFree) -> 0 (cudaSuccess: no error)
Not a different error. No error. The out-of-bounds write the assert was guarding happened silently into whatever sits past the 8 floats, inside a live allocation, so not even a 700 fired. Release builds define NDEBUG by default in most build systems, which means an assert-based correctness gate does not exist in the build that ships. This project's own linter bans runtime assert() in checked-in kernels for exactly this reason; use a branch that records the failure instead.
Causes, ranked
- An index out of range inside a PyTorch op. An embedding, gather, scatter or loss fed a label of 10 in a 10-class problem, or a token id past the vocabulary. The overwhelming majority of search traffic for this string. Fix: validate on the host; re-run the failing op on CPU, where the error names the bad index.
- Your own
assert()firing, as in the measured run above. Fix: read the printed message, then fix the launch or the data it names. - A library assert you did not write, from cuDNN or a fused kernel, usually a shape mismatch. Fix:
CUDA_LAUNCH_BLOCKING=1to pin the real call site, then check the shapes at that call.
Confirm it
Usually you do not need a tool, because the assert message already carries the coordinates. If the output was swallowed, compute-sanitizer --tool memcheck ./a.out reprints the assertion text; in cuda-gdb the signal is CUDA_EXCEPTION_12, Warp Assert, precise and per warp (https://docs.nvidia.com/cuda/cuda-gdb/index.html , checked 2026-08-29).
Prevention
- Validate indices on the host before they reach the device.
- Do not use
assert()as a shipping correctness gate; the measured-DNDEBUGrun above is what that buys you. A branch plus an error flag survives every build type. Day 6 shows the checking pattern.
Related errors
an illegal memory access was encountered(700): what the same bug produces once the write lands outside a live allocation.unspecified launch failure(719): the same family, minus the information.invalid configuration argument(9): a launch refused before any assert could run.
The lesson
Day 6 owns the error-reporting mechanics this page leans on; day 64 measures kernel asserts and device printf ordering. The full code table is on the error hub; a machine that cannot run any kernel starts at /setup/install-cuda.
Written 2026-09-01. Transcript from code/errors/device-side-assert-triggered/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.