What is a CUDA green context?
A CUDA execution context provisioned with a selected group of streaming multiprocessors, restricting its work to that SM partition.
A normal context schedules blocks across the whole GPU. A green context instead owns a resource descriptor built from one or more SM groups returned by the driver API. Work launched on its stream cannot borrow an SM outside that group, and work in a neighbouring green context cannot take one of its SMs.
The driver chooses legal group sizes. You ask cuDevSmResourceSplitByCount for a minimum, first with a null output array to simulate the split, then use the returned groups to build a descriptor and call cuGreenCtxCreate. The CUDA 12.6 reference describes 2-SM granularity for compute capability 7.x, 4-SM minimum with multiples of 2 for 8.x, and groups of 8 from 9.0 onward, while warning that these are architecture guidelines rather than constants to hard-code.
An SM partition is an isolation promise, not a concurrency guarantee or an automatic speedup. NVIDIA documents that disjoint green-context kernels may still fail to overlap because they share resources beyond SMs. The application must measure both the latency-sensitive job and total completion time.
Measured
On a 40-SM Tesla T4 (driver 580.173.02, CUDA 12.6), day 94 asked for 1 SM and received 20 groups of 2; asking for 3 produced 10 groups of 4. Its scheduled run then created two equal 20-SM green contexts.
With two ordinary whole-device streams, the short workload finished at 21.488 ms and everything finished at 21.491 ms. With the two green contexts, the short workload finished at 3.044 ms and everything at 22.004 ms. All schedules matched the closed-form reference and the alone runs bit for bit. This run demonstrates overlap on this T4; it does not turn overlap into an API guarantee.
Diagram: one 40-SM card shown first as one shared pool, then split into two fixed 20-SM bands. The short job waits behind batch blocks in the shared pool but retains its own SM band in the partitioned case.
Related terms
Where you meet this
- Day 94, sharing a GPU, which produces the split table and scheduling comparison above.
- Day 51, CUDA streams, where concurrent work still shares the whole device.
Sources
- CUDA Programming Guide, Green Contexts: https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/green-contexts.html (checked 2026-09-01)
- CUDA 12.6 Driver API, Green Contexts: https://docs.nvidia.com/cuda/archive/12.6.0/cuda-driver-api/group__CUDA__GREEN__CONTEXTS.html (checked 2026-09-01)
Byline
Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.