PyTorch published work on Jagged Flash Attention, the kernel behind Meta’s Generative Ads Model on NVIDIA Blackwell B200, built with TLX so developers get hardware-aware control layered on Triton’s tile-based model. Reported benchmarks compare against Flash Attention 4 from May 2026 on the jagged shapes that matter for GEM: roughly thirteen percent faster forward and about fifty percent faster backward in bfloat16 on B200. Implementation is tied to Meta’s ads_model_kernel_library, which gives operators a concrete reference for variable-length attention on Blackwell without waiting on a generic FA4 path for those layouts. The comparison is workload-specific to GEM jagged tensors on B200; the packet does not claim leadership across all attention shapes, dtypes, or GPU generations.
PyTorch reports Jagged Flash Attention for Meta GEM on NVIDIA Blackwell B200 built with TLX, with hardware-aware control on top of Triton. On performance, it outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass, benchmarked in bfloat16 on B200, with code in Meta ads_model_kernel_library.
Key takeaway
TLX-backed Jagged Flash Attention for Meta GEM on B200 beats FA4 (May 2026) on GEM jagged shapes, with larger backward gains than forward.
What happened
PyTorch reports Jagged Flash Attention for Meta’s Generative Ads Model on NVIDIA Blackwell B200, implemented with TLX (Triton Low-level Extensions) for explicit, hardware-aware control on top of Triton’s high-level, tile-based programming model.
On performance, PyTorch states the kernel outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM by about thirteen percent on the forward pass and about fifty percent on the backward pass, benchmarked in bfloat16 on B200, with code in Meta ads_model_kernel_library.
Evidence
JFA is the attention kernel behind Meta’s Generative Ads Model on Blackwell B200, built with TLX.
PyTorch · attributed
Jagged Flash Attention (JFA) — the attention kernel behind Meta’s Generative Ads Model (GEM) — on NVIDIA Blackwell (B200), built with TLX (Triton Low-level Extensions)
Reported speedups versus FA4 (May 2026) on GEM-relevant jagged shapes in bfloat16 on B200.
PyTorch · attributed
it outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass, benchmarked in bfloat16 on B200
Reference implementation is published in Meta’s ads model kernel library.
PyTorch · attributed
with code in Meta ads_model_kernel_library
Why it matters
Teams running Triton on Blackwell for irregular-sequence attention get a documented kernel stack and library hook for GEM-style jagged workloads instead of assuming FA4 is already optimal there.
Limits and uncertainties
Benchmarks are described for GEM jagged shapes on B200 in bfloat16 versus FA4 (May 2026); broader shape, dtype, or hardware coverage is not stated in the packet.
The link summary excerpt in the packet is truncated and does not add full methodology or workload definitions beyond the cited performance claims.
Practical implications
Inspect Meta ads_model_kernel_library when porting or tuning jagged attention on Blackwell with TLX rather than defaulting to stock FA4 for GEM-like layouts.
Profile forward and backward separately: reported backward gains (~50%) exceed forward (~13%) on the cited GEM jagged benchmarks.
What to watch
Whether PyTorch or Meta publish additional JFA benchmarks outside GEM jagged shapes or beyond B200 bfloat16 settings cited here.
How FA4 revisions after May 2026 compare on the same jagged GEM workloads if vendors rebaseline kernels.