← Glossary
CUDA glossaryPerformance
CC 7.5

What is online softmax?

A one-pass softmax recurrence that rescales its running sum whenever the maximum increases, allowing normalization without storing all logits first.

Stable softmax normally finds a row maximum, computes exponentials and their sum, then normalizes. Online softmax merges the first two passes. For a running maximum m and sum l, a larger new value changes m; the old sum is multiplied by exp(old_m - new_m) before the new exponential is added. The invariant remains the sum of exponentials relative to the current maximum.

Attention uses that recurrence tile by tile. A kernel can update the weighted-value accumulator alongside m and l, then rescale both when a key tile raises the maximum. The N-by-N score matrix never reaches global memory. This is the numerical idea behind FlashAttention, combined with tiling and kernel fusion.

Deleting traffic does not guarantee a speedup at every size. An online kernel can use more shared memory, lower occupancy and reread key/value tiles. Its summation order also differs from a materialized softmax, so correctness needs an independent reference and tolerance rather than bitwise equality.

Measured

On a Tesla T4 (driver 580.173.02, CUDA 12.6), day 97 compared a three-kernel score-matrix path with one online kernel. At n=256, online took 0.3230 ms versus naive's 0.0682 ms. At n=1024 the times were 1.3584 and 0.7836 ms. At n=4096, online won: 6.7860 ms versus 7.4483 ms, a naive/online ratio of 1.098. Maximum errors against the double reference were 1.310e-09 and 1.453e-09. The online kernel used 32,896 shared bytes, one block per SM and 25 percent occupancy.

Related terms

Where you meet this

Sources

Byline

Written by: pending. Reviewed by: pending. Written on: pending. Last checked on: pending. Verified numbers were captured on 2026-09-02; publication still requires named author and reviewer sign-off.