CUDA from Python: cuda.core, CuPy, Numba and PyCUDA
Four Python libraries can launch the kernel from day 5. A wall-clock measurement of the first call mostly compares their compilers, not their launch overhead.
CuPy compiles a
RawKernel "at an invocation of the
__call__() method"
(https://docs.cupy.dev/en/stable/reference/generated/cupy.RawKernel.html ,
checked 2026-09-01) and PyCUDA's SourceModule runs
nvcc in a subprocess before it returns anything at all
(https://github.com/inducer/pycuda/blob/main/pycuda/compiler.py , read
2026-09-01). Those paths can cost tens or hundreds of milliseconds before
the GPU does the measured work. The comparison below uses one set of buffers,
checks each output bit for bit, and reports compilation separately from
launches.
Where each library compiles, and when
The kernel is seven lines. Three libraries compile the same CUDA source, while Numba compiles an equivalent Python function. They differ in compiler, compile time, and whether compilation is explicit.
cuda.core is NVIDIA's binding. Compilation is explicit:
Program(...), then .compile("cubin"), then get_kernel(...), with NVRTC
running in your process.
cuda-python is now a metapackage; at 13.3.1 it
requires cuda-bindings~=13.3.1, cuda-core~=1.0.0 and
cuda-pathfinder~=1.1 (https://pypi.org/pypi/cuda-python/json , read
2026-09-01), so the import is cuda.core; CUDA 12 needs cuda-core[cu12].
cuda.core 1.0.0 removed the old experimental namespace: "Removed the deprecated
cuda.core.experimental namespace"
(https://nvidia.github.io/cuda-python/cuda-core/latest/release/1.0.0-notes.html
, checked 2026-09-01).
CuPy is an array library first and a kernel launcher second. RawKernel
defers compilation to the first call and then keeps the result: "The
compiled binary is also cached into a file under the
$HOME/.cupy/kernel_cache/ directory with a hashed file name. The cached
binary is reused by other processes."
The default backend is NVRTC; nvcc
is a constructor argument.
Numba takes a Python kernel instead of CUDA source. @cuda.jit compiles
it through NVVM to PTX at the first launch.
Its opt-in disk cache documentation says the version loaded is chosen by "the compute capability of the device the kernel is first launched on in the current run" (https://nvidia.github.io/numba-cuda/user/caching.html , checked 2026-09-01). This is the only row whose kernel does not come from the shared CUDA file.
PyCUDA wraps the driver API and compiles
eagerly: SourceModule(...) writes your source to a temporary file, runs
nvcc --cubin on it, and caches the result by checksum under
platformdirs.user_cache_dir("pycuda", "pycuda"). It therefore needs a full
toolkit at run time.
Diagram: four front ends, four moments of compilation. Four horizontal bands. Left of each is the Python process on a timeline with two marks, "you build the kernel object" and "first launch"; right of each is the compiler that runs, drawn as a box on the timeline where it fires. Band 1, cuda.core: NVRTC box on the build mark. Caption "Explicit. You write the compile call." Band 2, CuPy: build mark empty, NVRTC box on the first launch, with an arrow down to a disk labelled
$HOME/.cupy/kernel_cache/. Caption "Deferred, then cached on disk for the next process." Band 3, Numba: build mark empty, NVVM box on the first launch, labelled "once per argument type". Caption "Deferred, and typed." Band 4, PyCUDA: annvccbox on the build mark drawn outside the process boundary. Caption "Eager, and a separate program." Alt text: "Four Python front ends compile one kernel at four different moments. cuda.core compiles when you ask, PyCUDA spawns nvcc straight away, CuPy and Numba both wait for the first launch."
Separate compile time from launch time
Every launch moves 3n floats and performs one addition per element, so the kernel is limited by memory bandwidth. With 16 MiB per buffer, the GPU runs for hundreds of microseconds while the host can prepare the next launch. Day 48 measured 2.668 us of launch overhead for an empty kernel; a Python wrapper adds host work around that launch.
Python costs appear during first-call compilation and as per-call host work for small inputs. The timing columns separate these costs. When four libraries launch the same code, different steady-state times indicate extra work inside at least one timed region.
Choose a library
Installation requirements can rule out a library before its API matters.
| cuda.core | CuPy | Numba | PyCUDA | |
|---|---|---|---|---|
| For | calling your own CUDA from Python with nothing in the way | array work, with RawKernel as the escape hatch |
writing the kernel itself in Python | maintaining code that already uses it |
| Install | cuda-core[cu12], wheels |
cupy-cuda12x, one wheel |
numba-cuda[cu12], wheels |
sdist only, pip compiles it, needs nvcc |
| Kernel source | your .cu text |
your .cu text |
Python | your .cu text |
| Compiles | when you call .compile() |
first __call__, cached to disk by default |
first launch, and the disk cache is opt-in (cache=True) |
in the SourceModule constructor, via nvcc |
| Takes arrays from | raw device pointers, plus __cuda_stream__ for a foreign stream |
__cuda_array_interface__ and DLPack, both ways |
__cuda_array_interface__ via as_cuda_array |
any argument exposing __cuda_array_interface__ |
The last row supports the next two lessons. CuPy "supports
importing any array- or tensor-like objects that support the DLPack protocol
through cupy.from_dlpack()" and says PyTorch tensors also carry
__cuda_array_interface__, "so zero-copy data exchange between CuPy and
PyTorch can be achieved at no cost"
(https://docs.cupy.dev/en/stable/user_guide/interoperability.html , checked
2026-09-01). That interface lets a custom PyTorch
operator share device storage with CuPy.
For new code, use cuda.core to launch your own kernels or CuPy when you also
need arrays. PyCUDA 2026.1 shipped in February 2026 and remains a short path
from a .cu string to a running kernel. Unlike the other three, pip builds
its C++ extension and each run needs a toolkit.
One kernel, read out of the file once
Full programs: code/day85-python/vector_add.cu and code/day85-python/four_ways.py.
The kernel comes from day 5 instead of
day 11's strided probe. One IEEE-754 addition
per element has no fusion or reduction-order difference, so all four
libraries must produce identical bits. The check can use np.array_equal
without a tolerance.
One copy of the source. The Python program reads the kernel from the
snippet markers in the .cu file instead of copying it into Python.
def kernel_source() -> str:
"""The kernel text from vector_add.cu, wrapped for the JIT front ends.
Reading it out of the .cu by marker is what makes "the same kernel" true
rather than a claim. extern "C" turns off name mangling, so all three
front ends can look the kernel up as plain "vectorAdd".
"""
src = (Path(__file__).resolve().parent / "vector_add.cu").read_text()
region = re.search(r"^// snippet: kernel$(.*?)^// end snippet$",
src, re.M | re.S)
if region is None:
raise SystemExit("no `kernel` snippet region in vector_add.cu")
return 'extern "C" {\n' + region.group(1).strip("\n") + "\n}\n"
One allocator. CuPy owns all three buffers, and the other rows launch against its pointers. This removes buffer placement as a source of timing differences. The cuda.core path shows the binding steps:
from cuda.core import (Device, EventOptions, LaunchConfig, Program,
ProgramOptions, launch)
dev = Device()
dev.set_current()
stream = dev.create_stream()
build = time.perf_counter()
options = ProgramOptions(std="c++17", arch=f"sm_{dev.arch}")
program = Program(src, code_type="c++", options=options)
module = program.compile("cubin", name_expressions=("vectorAdd",))
kernel = module.get_kernel("vectorAdd")
build_ms = (time.perf_counter() - build) * 1.0e3
config = LaunchConfig(grid=BLOCKS, block=THREADS)
args = (d_a.data.ptr, d_b.data.ptr, d_out.data.ptr, np.uint64(N))
# CuPy filled the inputs on a different stream from this one.
dev.sync()
Two clocks. build ms and first ms use time.perf_counter because a
compiler does not run on the GPU and a
CUDA event cannot see it. mean ms is each
library's own event wrapper around ten launches after three warm-ups, which
is the same measurement vector_add.cu makes with cudaEventElapsedTime.
Numba uses an equivalent Python kernel, so it stays in its own row:
@nbcuda.jit
def vector_add(a, b, out):
i = nbcuda.grid(1)
if i < out.size:
out[i] = a[i] + b[i]
The harness fills the output buffer with NaN between rows. A launcher that does nothing then fails instead of inheriting the previous output, the same silent failure guarded against on day 39.
Results
Verified on hardware. Both Python processes and the C++ reference ran on a Tesla T4 (driver 580.173.02, CUDA 12.6), and every output matched bit for bit. The working Python stack was CuPy 13.6.0 with NumPy 1.26.4, cuda-core 1.1.1 with cuda-bindings 12.9.7, numba-cuda 0.30.4 and PyCUDA 2026.1 built from source.
The native vector_add.cu baseline was separately rebuilt with CUDA 13.0
V13.0.88 and passed all 4,194,915 elements at 0.1945 ms and 258.8 GB/s,
effectively unchanged from its CUDA 12.6 row. No Python binding was rerun in
that environment. The Python results and package compatibility claims below
remain tied to their recorded CUDA 12.x stack.
| Row | build ms | first ms | mean ms | GB/s | bits match |
|---|---|---|---|---|---|
| cuda.core | 44.7 | 0.4 | 0.1963 | 256.4 | yes |
| CuPy RawKernel | 0.0 | 38.9 | 0.1977 | 254.6 | yes |
Numba @cuda.jit |
0.1 | 297.2 | 0.2032 | 247.7 | yes |
| PyCUDA SourceModule | 465.6 | 0.8 | 0.1980 | 254.2 | yes |
vector_add.cu |
n/a | n/a | 0.1948 | 258.5 | yes |
The second Python process retained caches:
| Row | second build ms | second first ms | second mean ms |
|---|---|---|---|
| cuda.core | 44.5 | 0.3 | 0.1962 |
| CuPy RawKernel | 0.0 | 0.9 | 0.1969 |
Numba @cuda.jit |
0.1 | 267.6 | 0.2026 |
| PyCUDA SourceModule | 8.3 | 0.8 | 0.1981 |
What the run settled
- Held. All four Python libraries produced bit-identical output, and the C++ reference matched its CPU check across all 4,194,915 elements.
- Held. The Python steady-state means ranged from 0.1963 to 0.2032 ms, within 4.3 percent of C++ at 0.1948 ms. Each path moved the same bytes for the same operation, and the wrappers did not change steady-state GPU throughput once compilation was excluded.
- Held. cuda.core spent 44.7 ms in explicit build, PyCUDA spent 465.6 ms there, and deferred compilation appeared in CuPy's 38.9 ms and Numba's 297.2 ms first calls.
- Refuted. CuPy's first call fell from 38.9 to 0.9 ms, and the captured cache directory contains compiled cubins. It was not the only cell that moved: PyCUDA's build fell from 465.6 to 8.3 ms, showing its cache also survived the process. Numba still spent 267.6 ms on its second first call.
- Triggered and resolved. The original package set failed. CuPy 14.2
requires NumPy 2 or newer, while numba-cuda 0.30.4 imports the removed
np.row_stack. CuPy 13.6.0 plus NumPy 1.26.4 is the compatible pair. The retained failure transcripts and corrected README are part of the evidence contract.
The two time columns measure different costs and should remain separate.
Run it yourself
The Compiler Explorer embed runs only
vector_add.cu; it has no Python environment for the four-way comparison.
Follow the README for the Python run.
See the hardware access options if you need a
compatible environment.
pip install pycuda builds a C++ extension from the 2026.1 source tarball and
can take several minutes, while
the other three packages use wheels. Use the listed versions because the
latest CuPy and NumPy combination is not compatible with the current
numba-cuda release.
Exercise
Add a fifth and sixth row to the table using CuPy's own array path rather
than a raw kernel: c = a + b, and then cupy.add(a, b, out=c). Time both
exactly the way the other rows are timed, warm-up and all, and say in one
sentence what the difference between the two measures.
Time: 25 to 40 minutes. Submit: the two new mean times, the
RawKernel mean time beside them, and the sentence.
Check: the harness compares every row against the same NumPy reference
with np.array_equal and prints the first differing index, so a row that
allocated the wrong shape or wrote nowhere fails rather than looking fast.
It prints all six rows whether they pass or not, because the times are the
answer and the check is only the gate.
Hint 1
Both new rows run the same kernel on the same bytes. So write down everything each call does that is not the launch, and mark which of those things happens once and which happens on every call.
Hint 2
One call must obtain a new 16 MiB array before writing its result. Identify
where that memory comes from on the tenth call and what out= changes.
Solution
c = a + b obtains an output array on every call, so its host work includes
a trip to CuPy's memory pool that neither the out= form nor the raw kernel
pays. The two new rows measure this fixed per-call cost. At 16 MiB per
buffer, the kernel runs for hundreds of microseconds, so the three rows
should converge.
Shrink n until the ranking reappears. Day 9 describes the same crossover between per-call overhead and per-byte cost.
Pitfalls
Your benchmark ranks compilers. If you time one call of a RawKernel or a
@cuda.jit function, most of what you measured was
JIT compilation.
Warm up, time with events, and report the first call separately when it matters. Day 9 explains why.
The same script is faster the second time. CuPy
wrote the compiled binary to $HOME/.cupy/kernel_cache/ and says the cache
"is reused by other processes". Nothing about your code changed.
Either report both runs or clear the cache, and say which you did.
You write import pycuda.autoinit and your CuPy pointers stop working.
autoinit calls Device.make_context(), which creates a second context on
the same device, and every other library here lives in the primary one.
pycuda.autoprimaryctx calls retain_primary_context() instead and the two
share (both files at
https://github.com/inducer/pycuda/tree/main/pycuda , read 2026-09-01).
You install cuda-python and get bindings for the wrong toolkit. The
metapackage pulls the 13 line. For CUDA 12, install cuda-core[cu12],
which resolves cuda-bindings[all]==12.* and the matching NVRTC wheel.
Day 44 lost a session to the mirror image of this, libcublas.so.13.
pip install pycuda runs for minutes and then fails. There is no wheel,
so pip compiles a C++ extension against your
toolkit headers and needs nvcc on PATH. So does every later run, because
SourceModule shells out to it.
You time work on one stream with events recorded on another. NVIDIA's own cuda.core example carries the comment "cupy runs on a different stream from stream, so sync before accessing" (https://github.com/NVIDIA/cuda-python/blob/v13.3.1/cuda_core/examples/vector_add.py , read 2026-09-01). Sync at the boundary, or hand one library the other's stream.
Go deeper
- cuda.core, "10 minutes to cuda.core", the compile-and-launch example this page's row follows: https://nvidia.github.io/cuda-python/cuda-core/latest/getting-started.html
- CuPy user guide, "Raw kernels" and "Interoperability", for
RawKerneland the DLPack rules: https://docs.cupy.dev/en/stable/user_guide/kernel.html and https://docs.cupy.dev/en/stable/user_guide/interoperability.html - Numba-CUDA, "Writing CUDA Kernels" and "Installation": https://nvidia.github.io/numba-cuda/user/kernels.html
- PyCUDA documentation,
pycuda.driver.EventandSourceModule: https://documen.tician.de/pycuda/ - CUDA C++ Best Practices Guide 8.1, "Using CUDA GPU Timers": https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html#using-cuda-gpu-timers
Next
Day 86 keeps Python but replaces CUDA source with Triton
for softmax and matrix multiplication. Its GPU path needs compute capability
8.0, while TRITON_INTERPRET=1 covers correctness without a supported GPU.
Day 87 uses the interoperability interface to wrap a CUDA kernel as a PyTorch
operator.