← Glossary
CUDA glossaryLibraries
CC 7.5

What is separate compilation in CUDA?

Compiling device code in multiple translation units and joining it with a device-link step before the final host link.

Ordinary host linking cannot resolve a call from one .cu file to a device function in another. nvcc first emits relocatable device code, then performs a device link, then completes the host link. In CMake, CUDA_SEPARABLE_COMPILATION requests that pipeline for a target (https://cmake.org/cmake/help/v3.28/prop_tgt/CUDA_SEPARABLE_COMPILATION.html, checked 2026-09-02).

The tradeoff is an optimization boundary. Unless link-time optimization restores it, the compiler cannot treat a cross-file device call exactly like a local function it can inline. Inspect the final binary and resource report instead of assuming the boundary is free or automatically expensive.

Measured

On a Tesla T4 build with CUDA 12.6 and CMake 4.4.3, day 67 disabled relocatable device code and ptxas stopped with Unresolved extern function '_Z7fmaStepfff', exit 255. With separate compilation enabled, both kernels passed. cuobjdump -res-usage reported 40 registers for the external-call kernel and 45 for the local version; the work moved into a separate 24-register callee rather than disappearing.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came from the verification node on 2026-09-01.