CUDA & GPU Systems

FlashAttention-3 on NVIDIA Hopper/Blackwell: FP8 Asynchronous Tensor-Core Warps

A deep CUDA GPU kernel engineering guide to FlashAttention-3. We dissect warp-specialized asynchronous Tensor Core operations, block-sparse softmax quantization, overlapping GMEM-SMEM transfers via TMA, and achieving 850 TFLOPs/sec.

Sachin Sharma
Sachin SharmaCreator
Sep 2, 2026
3 min read
FlashAttention-3 on NVIDIA Hopper/Blackwell: FP8 Asynchronous Tensor-Core Warps
Featured Resource
Quick Overview

A deep CUDA GPU kernel engineering guide to FlashAttention-3. We dissect warp-specialized asynchronous Tensor Core operations, block-sparse softmax quantization, overlapping GMEM-SMEM transfers via TMA, and achieving 850 TFLOPs/sec.

Sachin Sharma

Sachin Sharma

Software Developer & Mobile Engineer

Building digital experiences at the intersection of design and code. Sharing weekly insights on engineering, productivity, and the future of tech.