What is cuda-gdb?
The CUDA-aware debugger that can stop inside a kernel, select a GPU thread and inspect its device-side state.
cuda-gdb extends the familiar GDB model with CUDA grids, blocks, warps and lanes. You can break on a kernel source line, switch focus to a named device thread and print its local variables. NVIDIA's manual documents the CUDA focus syntax and device inspection commands (https://docs.nvidia.com/cuda/cuda-gdb/index.html, checked 2026-09-01).
Build with -G when you need reliable device source debugging. That build changes optimization, so it is diagnostic, not the binary to benchmark. Use Compute Sanitizer first for bad memory or synchronization, then cuda-gdb when you need to see why one thread computed the wrong address.
Measured
On a Tesla T4 with driver 595.84 and CUDA 12.6, day 63 stopped all 2,560 threads at one source line and focused block (1,0,0), thread (2,3,0). The thread had loaded source index 201 but computed destination 669; the correct row * width + col was 201. A second thread showed the same transposition, 965 versus 209. The release build took 0.0037 ms per launch and the standalone -G build took 0.0064 ms, 1.7x slower.
Related terms
Where you meet this
- Day 63, cuda-gdb, the complete batch transcript.
- Day 64, device printf and assert, two lighter-weight ways to expose device state.
- Illegal memory access, where a debugger can inspect the failing thread.
Sources
- NVIDIA CUDA-GDB documentation, for focus and inspection commands: https://docs.nvidia.com/cuda/cuda-gdb/index.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.