What is separate compilation in CUDA?
Compiling device code in multiple translation units and joining it with a device-link step before the final host link.
Ordinary host linking cannot resolve a call from one .cu file to a device function in another. nvcc first emits relocatable device code, then performs a device link, then completes the host link. In CMake, CUDA_SEPARABLE_COMPILATION requests that pipeline for a target (https://cmake.org/cmake/help/v3.28/prop_tgt/CUDA_SEPARABLE_COMPILATION.html, checked 2026-09-02).
The tradeoff is an optimization boundary. Unless link-time optimization restores it, the compiler cannot treat a cross-file device call exactly like a local function it can inline. Inspect the final binary and resource report instead of assuming the boundary is free or automatically expensive.
Measured
On a Tesla T4 build with CUDA 12.6 and CMake 4.4.3, day 67 disabled relocatable device code and ptxas stopped with Unresolved extern function '_Z7fmaStepfff', exit 255. With separate compilation enabled, both kernels passed. cuobjdump -res-usage reported 40 registers for the external-call kernel and 45 for the local version; the work moved into a separate 24-register callee rather than disappearing.
Related terms
Where you meet this
- Day 67, CMake and CUDA, the multi-file build and failed control.
- CMake cannot find CUDA, the toolchain setup failure that happens earlier.
- Day 87, a PyTorch custom op, a later multi-file CUDA extension.
Sources
- NVIDIA nvcc manual, for relocatable device code and device linking: https://docs.nvidia.com/cuda/archive/12.6.3/cuda-compiler-driver-nvcc/index.html (checked 2026-09-01)
- CMake target property reference, for
CUDA_SEPARABLE_COMPILATION: https://cmake.org/cmake/help/v3.28/prop_tgt/CUDA_SEPARABLE_COMPILATION.html (checked 2026-09-02)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.