CUDA errorsdraft
Reproduced 2026-09-01

CUDA ERROR

invalid device symbol: causes and fix

CUDA error 13 means the runtime did not recognise a device symbol. On a T4, the often-blamed ampersand succeeded and a host array failed.

The runtime looked up the thing you passed to a symbol API and found no device variable registered under it.

Enum cudaErrorInvalidSymbol
Code 13
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

  • Runtime: invalid device symbol
  • The near neighbour is named symbol not found (500), which is what cudaGetSymbolAddress returns for a name that does not exist at all. 13 means "this is not a symbol"; 500 means "no symbol by that name".

The measured surprise: the ampersand is not the bug

Every second answer about this error blames a stray & in front of the symbol name. We tested that directly. From code/errors/invalid-device-symbol/evidence/run-2026-09-01.txt, Tesla T4 (sm_75), driver 595.84, CUDA 12.6:

cudaMemcpyToSymbol(filter, ...)    -> 0 (cudaSuccess: no error)
cudaMemcpyToSymbol(&filter, ...)  [spurious &] -> 0 (cudaSuccess: no error)

Both succeeded. The reason is ordinary C: filter is declared __constant__ float filter[16], and for an array the address of the array and the address of its first element are the same number, so the runtime received an identical pointer both times and resolved the same symbol.

The documented contract still stands, and it is the one to follow: pass the device variable itself, not its address, because the API resolves the symbol through the compiler's registration table rather than by dereferencing a pointer you hand it. Our capture only shows that CUDA 12.6 on this card tolerates the address-of form for an array; that tolerance is not something to write code against, and it will not compile at all without a cast to force it through.

A second program in the same session found what does. It declares a plain host array next to the real __constant__ one and passes that:

cudaMemcpyToSymbol(filter, ...)          -> 0 (cudaSuccess: no error)
cudaMemcpyToSymbol(hostSide, ...)  [host array, not a symbol] -> 13 (cudaErrorInvalidSymbol: invalid device symbol)

That is the honest one-line rule this page exists to publish: error 13 is about what the address refers to, not about how you wrote the expression. A pointer the compiler never registered as device data fails, whether or not you took its address correctly.

Cause 1: the target is not a device symbol at all

The measured failure. A host array, a local buffer or a plain global passed where a __constant__ or __device__ variable belongs.

__constant__ float filter[16];
float hostSide[16];
cudaMemcpyToSymbol(hostSide, h, sizeof(h));   // 13: hostSide is host data

Fix: pass the device variable itself. If the destination is ordinary device memory from cudaMalloc, you want cudaMemcpy, not the symbol API; the two are not interchangeable and the symbol call is the one that knows about the compiler's registration table.

Cause 2: separate compilation without device linking

A __constant__ variable defined in one translation unit and referenced from another, built without -rdc=true. Each unit gets its own copy or none at all, and the symbol the host side resolves is not the one the kernel reads.

Fix: -rdc=true across the project, or CUDA_SEPARABLE_COMPILATION ON in CMake. Note that our separate-compilation repro for invalid device function showed a plain kernel launch across two files working fine without that flag under CUDA 12.6, so do not reach for -rdc=true on faith. It is cross-unit device data and cross-unit __device__ functions that need it, not every multi-file build.

Cause 3: a symbol referenced by string name

Old tutorials pass "filter" as a string. That form was removed from the API years ago, so code copied from a 2010 answer fails against a modern toolkit.

Fix: pass the variable, not its name in quotes.

Confirm which one you have

cuobjdump -symbols ./a.out | grep filter

If the name is present, the variable really is a device symbol and you are passing the wrong thing at the call site, which is cause 1. If it is absent, the definition never reached this binary and you are in cause 2. One caution from our own session: cuobjdump is not present in every toolkit packaging, so confirm the tool exists before you build a habit around it.

Prevention

  • Keep __constant__ definitions in a header included by every unit that touches them, or accept -rdc=true deliberately and write down why.
  • Check the return value of the symbol copy. It is a plain cudaError_t and it fails synchronously, so unlike a launch failure there is no excuse for finding out later.

Related errors

  • invalid device function (98): the same registration table, missing a kernel instead of a variable.
  • invalid argument (1): the catch-all for a bad size or a null pointer in the same call.
  • named symbol not found (500): the lookup-by-name cousin, with no page of its own yet.

The lesson

Day 18 measures constant memory winning on broadcast reads and losing badly on per-lane reads, which is the reason to be using this API at all. Day 46 explains what the symbol table in a binary actually holds. The full code table is on the error hub, and toolkit setup lives at /setup/install-cuda.


Written 2026-09-01. Transcripts from code/errors/invalid-device-symbol/evidence/run-2026-09-01.txt, captured 2026-09-01 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.