Research questionHow can FP4 attention exploit Blackwell tensor cores when softmax overhead dominates?Shrinking attention matrix products to FP4 does not guarantee faster execution when softmax conversion and on-chip dependencies become the bottleneck. Causal training also requires quantization choices that preserve usable gradients.