CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

invalid device function: causes and fix

CUDA error 98 means the kernel symbol is missing for this device. On a T4, the classic no-rdc reproduction still linked and ran under CUDA 12.6.

The runtime went looking for this kernel's compiled function on this device and did not find one, even though some device code loaded.

Enum cudaErrorInvalidDeviceFunction
Code 98
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

One paragraph of arithmetic hygiene: a widely indexed competitor error list publishes this as error 8. It is 98; 8 is cudaErrorProfilerAlreadyStopped. Both values are in our own runtime dump, code/errors/_strings/evidence/run-2026-09-01.txt, printed by cudaGetErrorName on a real device, and in the runtime API page linked above.

Strings a reader might paste

  • Runtime: invalid device function
  • The report usually arrives as a story: "worked yesterday, broke after I added a second .cu file" or "broke on the new GPU".

What we measured: the folklore repro does not fail on CUDA 12.6

The standard explanation for this error blames separate compilation: a kernel defined in one translation unit, launched from another, built without -rdc=true. We built exactly that, a worker kernel defined in kernels.cu and declared and launched in main.cu. From code/errors/invalid-device-function/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:

== build ==
nvcc -std=c++17 -O3 -arch=sm_75 -o repro main.cu kernels.cu   # NO -rdc=true

== run ==
cudaMalloc                         -> 0 (cudaSuccess: no error)
peek after launch                  -> 0 (cudaSuccess: no error)
cudaDeviceSynchronize              -> 0 (cudaSuccess: no error)

The -rdc=true build in the same file runs identically. So a __global__ kernel in another .cu file does not need relocatable device code just to be launched; each translation unit gets whole-program compilation and the launch resolves through the host-side stub. What genuinely needs -rdc=true is a __device__ function called across translation units, and that failure is a link error at build time, not a 98 at run time. We publish the negative result because half the answers on this error tell you to flip a flag that, in this shape, changes nothing.

When 98 actually appears

From the documentation and the mechanism, not from a captured run (this page gains a transcript when the course produces one naturally):

  1. The binary has an image for your device but not this symbol in it. A template kernel whose instantiation only exists in a translation unit compiled for a different architecture list, or a kernel behind an #ifdef that is off in the build that ran. The runtime API description for 98 is "the requested device function does not exist or is not compiled for the proper device architecture" (source in the identity box, checked 2026-08-29).
  2. A per-kernel architecture mismatch inside a fat binary. Practically a cousin of no kernel image is available; which of the two you get depends on whether any image matched the device at all. Our measured 209 run on the same node shows the whole-image case.
  3. A function pointer or name passed to the runtime that is not a device function. cudaFuncGetAttributes or cudaLaunchKernel handed a host function.

Confirm it

cuobjdump -symbols ./a.out | grep <kernel>

If the mangled name is absent, the symbol was compiled out (cause 1); if present, compare architectures per the no-kernel-image page. Note from our session: cuobjdump is not installed with every toolkit packaging (the evidence for the 209 page records cuobjdump: command not found on this node), so check it exists before building a workflow on it.

Prevention

  • One explicit -gencode list shared by every translation unit in the target; a per-file list that drifts is how one symbol goes missing.
  • If you use -rdc=true for real cross-unit device calls, set it project-wide (CUDA_SEPARABLE_COMPILATION ON in CMake) rather than per file.

Related errors

The lesson

Day 46 explains images, PTX and SASS; day 67 owns separable compilation in CMake. The full code table is on the error hub; toolchain setup lives at /setup/install-cuda.


Written 2026-09-01. Transcripts from code/errors/invalid-device-function/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.