← Glossary
CUDA glossaryExecution model
CC 7.5

What is a CUDA event?

A marker you record into a stream, used either to time GPU work or to make one stream wait on another.

cudaEventRecord(ev, stream) drops a marker into a stream's queue; the event completes when everything before it in that stream has finished. Two uses follow from that one mechanism. Timing: record an event before and after some work, and cudaEventElapsedTime gives the gap by the GPU's clock, which is how this course times every kernel from day 9 on, because a host-side std::chrono around an asynchronous launch times the launch call, not the kernel. Ordering: cudaStreamWaitEvent(streamB, ev) makes stream B hold at that point until the event completes, which is how independent streams express "B needs A's result" without synchronizing the whole device.

The ordering use is the one that scales. A fork-join, one producer feeding several parallel branches feeding a combiner, is a diamond of dependencies, and events express exactly the edges: the branches wait on the producer's event, the join waits on each branch's. Everything not connected by an edge stays free to overlap. The alternative people reach for first, cudaDeviceSynchronize() between phases, adds edges to everything and drags the host into the middle of each step.

One flag matters more than it looks: events for ordering should be created with cudaEventDisableTiming, which makes them cheaper to record, and the API enforces the division: asking such an event for a time is an error, not a zero (https://docs.nvidia.com/cuda/archive/12.6.3/cuda-runtime-api/group__CUDART__EVENT.html , checked 2026-09-01).

Measured

Tesla T4, driver 595.84, CUDA 12.6 (V12.6.85), built with nvcc -std=c++17 -O3 -arch=sm_75 -lineinfo, captured 2026-09-01. Day 52 runs a four-branch fork-join three ways:

arrangement result
diamond, event-ordered across streams 4.826 ms, 0 mismatches
same work, one stream, serial 8.366 ms, 0 mismatches
streams with the waits removed 986400 of 1048576 elements wrong

The event-ordered diamond runs 1.73x faster than the serial version because the four branches genuinely overlap, and both produce exact results. The third row is why the edges are not optional: without the waits the combiner races the branches, and the corruption count is not even stable run to run. The same program asked cudaEventElapsedTime for a time on its cudaEventDisableTiming fork event and got cudaErrorInvalidResourceHandle, the exact string, so the two event roles do not mix silently.

Diagram

timeline-host-device, preset fork-join-events: four branch lanes starting at the producer's event marker and a join lane starting where all four branch markers complete, against a serial lane of the same six boxes end to end.

Alt text: "Events let four branch kernels start together after the producer and let the join start after the last branch. The same work in one stream runs the boxes end to end at 1.73 times the cost."

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. The numbers came off the verification node on 2026-09-01, and this entry stays a draft until a named author and a different named reviewer sign it.