FlashAttention-3 on NVIDIA Hopper/Blackwell: FP8 Asynchronous Tensor-Core Warps
A deep CUDA GPU kernel engineering guide to FlashAttention-3. We dissect warp-specialized asynchronous Tensor Core operations, block-sparse softmax quantization, overlapping GMEM-SMEM transfers via TMA, and achieving 850 TFLOPs/sec.
9/2/202624 min read