What is a PyTorch custom operator?
A user-defined operation registered with PyTorch's dispatcher, with explicit device implementations and optional fake and autograd registrations.
A raw kernel launcher sees pointers and sizes. A PyTorch tensor also carries a device, dtype, shape, strides, current CUDA stream and autograd history. A custom operator is the contract that connects those worlds. The C++ TORCH_LIBRARY schema names the operation; per-device implementations tell the dispatcher where it can run; register_fake describes output metadata without reading data; and register_autograd supplies the backward formula.
Those pieces solve different failures. A CUDA implementation makes eager execution possible. A fake implementation lets torch.compile reason about shapes without launching the kernel. An autograd registration keeps gradients flowing. Omitting any one can leave the forward numerically correct while compilation or training is still broken.
The operator must preserve framework execution rules. It should select the inputs' device, allocate the output there, and launch on PyTorch's current stream rather than creating or assuming another one. It should raise an exception for an invalid contract instead of calling std::exit, and it should not insert a device-wide synchronization merely to make errors easier to observe.
Measured
On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 87 registered a WMMA FP16-input, FP32-output matrix multiply. Four invalid calls raised, two non-square forward cases matched float64 CPU references, CPU and CUDA opcheck passed, float64 gradcheck passed, and a deliberately wrong backward formula raised GradcheckError. The extension was built with both sm_75 and compute_75 targets.
The same harness then ran with CUDA_VISIBLE_DEVICES=. The contract, forward and CUDA opcheck rows skipped as designed, while CPU opcheck, gradcheck and the deliberately wrong-formula check still passed. That split matters: it proves the backward formula can be reviewed without pretending the CUDA kernel ran.
Related terms
Where you meet this
- Day 87, custom PyTorch operators, which builds and verifies the operator above.
- Day 88, profiling a PyTorch model, where a registered fused operator appears by name in framework and CUDA traces.
- Day 69, CUDA portability, for the cubin and PTX targets embedded in an extension.
Sources
- PyTorch, “Custom C++ and CUDA Operators”: https://docs.pytorch.org/tutorials/advanced/cpp_custom_ops.html (checked 2026-09-01)
torch.libraryAPI reference: https://docs.pytorch.org/docs/2.13/library.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified results were captured on 2026-09-02; publication still requires named author and reviewer sign-off.