Skip to content

BF16 gradients rise late in training

A paper reports a 1,000-fold increase in gradient norm after 25 billion tokens in a transformer trained with FlashAttention-3 in BF16.

By Wendelmaques ·

Source: Gradientes de BF16 no FlashAttention-3 podem crescer no fim do treinamento (arxiv.org). Text prepared with AI from this source.

What happened and what to do

Published on October 4, 2026, the paper reports that a 450-million-parameter transformer trained with FlashAttention-3 in BF16 showed a 1,000-fold increase in gradient norm after 25 billion tokens. Training ended with a higher loss than the FP32 attention configuration, without NaNs. Recomputing the attention backward pass for two layers in FP32 removed the reported problem.

Companies training models with fused BF16 attention can assess stability across long runs by comparing gradient norms and loss against an FP32 configuration. A practical implementation could log these metrics by training step, alert on late-stage deviations, and test selective FP32 backward recomputation before adopting the mitigation in production.

How the consultancy can help

Wendelmaques can diagnose exposure in the training pipeline, define comparative tests, and implement gradient and loss monitoring as well as a scoped FP32 mitigation for the client’s environment, with agreed implementation and operation.

Next step

Send a short description of the model, training setup, and observed issue to receive a scoped proposal.

Consulting for your project

Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.

Quoted per project

Request a proposal