CUDA errorsdraft
Documented

CUDA ERROR

invalid argument: causes and fix

cudaErrorInvalidValue is code 1. On a T4, a null destination and 53248 bytes of dynamic shared memory failed; the wrong cudaMemcpyKind did not.

One argument you passed is outside the range the API accepts, and because code 1 is the runtime's catch-all, the whole job of this page is narrowing it down by which call raised it.

Enum cudaErrorInvalidValue
Code 1
Source https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html (CUDA 13.3, checked 2026-08-29)

Strings a reader might paste

  • Runtime: invalid argument
  • Driver API: CUDA_ERROR_INVALID_VALUE
  • PyTorch: CUDA error: invalid argument

Unlike error 700, code 1 is reported synchronously, by the call that made it. The transcript below proves it: the failing cudaMemcpy itself returned the code, and both peek and get saw the same status. If your traceback names a call, that call really is the guilty one.

Cause 1: a null pointer or a zero size where the API forbids it

The commonest shape is a size computed from an empty container, or a pointer variable passed to cudaMalloc instead of its address. Here is the measured behaviour of a null destination, from code/day06-error-checking/evidence/run-2026-08-30.txt, captured on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6:

part 1: one failed call, and who can still see it
  cudaMemcpy returned              cudaErrorInvalidValue          invalid argument
  cudaPeekAtLastError              cudaErrorInvalidValue          invalid argument
  cudaPeekAtLastError, again       cudaErrorInvalidValue          invalid argument
  cudaGetLastError                 cudaErrorInvalidValue          invalid argument
  cudaPeekAtLastError, after get   cudaSuccess                    no error

part 2: the same context after a non-sticky error
  all 611 elements match, so the context survived it

Two useful facts in those seven lines. cudaGetLastError clears the status and cudaPeekAtLastError does not, which is the whole difference between them. And the error is not sticky: the same context ran a correct copy and a correct kernel launch immediately afterwards. A code 1 does not poison anything; fix the argument and carry on in the same process.

Cause 2: dynamic shared memory above the default per-block limit

Requesting more shared memory in the launch's third parameter than the device's default cap does not return a resource error; on this card it returns invalid argument at the launch. Measured in code/day13-shared-memory/evidence/run-2026-08-30.txt, same T4, where the default cap is 49152 bytes and the opt-in ceiling is 65536:

53248 bytes of dynamic shared memory, no opt-in: invalid argument
53248 bytes after cudaFuncSetAttribute:         no error

Fix: query cudaDevAttrMaxSharedMemoryPerBlockOptin, then opt in with cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, bytes). Day 13 does exactly this and shows the before and after.

Cause 3: the wrong cudaMemcpyKind, which mostly does not fail

Every checklist tells you a reversed direction flag returns an error. We tried it, and on this hardware it does not. The first run of day 6's program forced the error by tagging a host-to-device copy cudaMemcpyDeviceToHost. From the superseded-run section of code/day06-error-checking/evidence/run-2026-08-30.txt:

  the reversed copy succeeded, so part 1 shows nothing on this device

  part 1: one failed call, and who can still see it
    cudaMemcpy returned              cudaSuccess                    no error

The copy also produced correct data. With unified addressing the runtime reads the real direction from the pointer attributes and treats the kind argument as a hint it can overrule. So the direction flag will not catch your mistake for you, and when a reversed copy does fail (older setups, pointers the runtime cannot attribute), the code you get is 1 rather than 21, invalid copy direction for memcpy. Either way, do not rely on it: use cudaMemcpyDefault, or read the argument order out loud, destination first.

Confirm which one you have

Because code 1 is synchronous, instrumentation is the whole diagnosis. Wrap every call in the CUDA_CHECK macro from day 6 and the failing line names itself. The canonical Stack Overflow question on the macro has 178,060 views (https://stackoverflow.com/questions/14038589/what-is-the-canonical-way-to-check-for-errors-using-the-cuda-runtime-api , checked 2026-08-29). If the failing call is a launch, print the grid, the block and the shared-memory request next to the device limits before it.

Prevention

  • Never pass a size you did not compute in size_t.
  • cudaMalloc(&d, bytes), not cudaMalloc(d, bytes). The compiler cannot catch the second one through the void** conversion.
  • Treat any dynamic shared request above 48 KiB as an opt-in feature, because on most cards it is.

Related errors

The lesson

Day 6 is the lesson this page's main transcript comes from, and its whole subject is who sees an error and when. The full code table lives on the error hub; environment problems that produce code 1 on the very first call usually belong at /setup/install-cuda instead.


Written 2026-09-01. Transcripts quoted from code/day06-error-checking/evidence/run-2026-08-30.txt and code/day13-shared-memory/evidence/run-2026-08-30.txt, both captured 2026-08-30 on a Tesla T4 (sm_75), driver 595.84, CUDA 12.6 (V12.6.85). Author and reviewer: not yet assigned; this page does not publish until both are named.