Skip to main content
LLMgram · AI News · 2026-10-02

PyTorch TLX Jagged Flash Attention outruns FA4 on Blackwell GEM jagged workloads

PyTorch TLX Jagged Flash Attention outruns FA4 on Blackwell GEM jagged workloads

PyTorch published work on Jagged Flash Attention, the kernel behind Meta’s Generative Ads Model on NVIDIA Blackwell B200, built with TLX so developers get hardware-aware control layered on Triton’s tile-based model. Reported benchmarks compare against Flash Attention 4 from May 2026 on the jagged shapes that matter for GEM: roughly thirteen percent faster forward and about fifty percent faster backward in bfloat16 on B200. Implementation is tied to Meta’s ads_model_kernel_library, which gives operators a concrete reference for variable-length attention on Blackwell without waiting on a generic FA4 path for those layouts. The comparison is workload-specific to GEM jagged tensors on B200; the packet does not claim leadership across all attention shapes, dtypes, or GPU generations.

Sources

PyTorch TLX Jagged Flash Attention outruns FA4 on Blackwell GEM jagged workloads

PyTorch TLX Jagged Flash Attention outruns FA4 on Blackwell GEM jagged workloads

PyTorch reports Jagged Flash Attention for Meta GEM on NVIDIA Blackwell B200 built with TLX, with hardware-aware control on top of Triton. On performance, it outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass, benchmarked in bfloat16 on B200, with code in Meta ads_model_kernel_library.

Key takeaway

TLX-backed Jagged Flash Attention for Meta GEM on B200 beats FA4 (May 2026) on GEM jagged shapes, with larger backward gains than forward.

What happened

PyTorch reports Jagged Flash Attention for Meta’s Generative Ads Model on NVIDIA Blackwell B200, implemented with TLX (Triton Low-level Extensions) for explicit, hardware-aware control on top of Triton’s high-level, tile-based programming model.

On performance, PyTorch states the kernel outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM by about thirteen percent on the forward pass and about fifty percent on the backward pass, benchmarked in bfloat16 on B200, with code in Meta ads_model_kernel_library.

Evidence

  • JFA is the attention kernel behind Meta’s Generative Ads Model on Blackwell B200, built with TLX.

    PyTorch · attributed

    Jagged Flash Attention (JFA) — the attention kernel behind Meta’s Generative Ads Model (GEM) — on NVIDIA Blackwell (B200), built with TLX (Triton Low-level Extensions)

  • Reported speedups versus FA4 (May 2026) on GEM-relevant jagged shapes in bfloat16 on B200.

    PyTorch · attributed

    it outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass, benchmarked in bfloat16 on B200

  • Reference implementation is published in Meta’s ads model kernel library.

    PyTorch · attributed

    with code in Meta ads_model_kernel_library

Why it matters

Teams running Triton on Blackwell for irregular-sequence attention get a documented kernel stack and library hook for GEM-style jagged workloads instead of assuming FA4 is already optimal there.

Limits and uncertainties

Benchmarks are described for GEM jagged shapes on B200 in bfloat16 versus FA4 (May 2026); broader shape, dtype, or hardware coverage is not stated in the packet.

The link summary excerpt in the packet is truncated and does not add full methodology or workload definitions beyond the cited performance claims.

Practical implications

Inspect Meta ads_model_kernel_library when porting or tuning jagged attention on Blackwell with TLX rather than defaulting to stock FA4 for GEM-like layouts.

Profile forward and backward separately: reported backward gains (~50%) exceed forward (~13%) on the cited GEM jagged benchmarks.

What to watch

Whether PyTorch or Meta publish additional JFA benchmarks outside GEM jagged shapes or beyond B200 bfloat16 settings cited here.

How FA4 revisions after May 2026 compare on the same jagged GEM workloads if vendors rebaseline kernels.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell