Dust explores gradient-free transformer training
Published on October 5, 2026, the paper presents Dust, a zeroth-order method using activation perturbations, and reports scaling experiments and comparisons with backpropagation.
Source: Dust usa perturbações de ativação para treinar transformers sem gradientes (x.com). Text prepared with AI from this source.
What happened and what to do
The paper, published on October 5, 2026, describes Dust, a zeroth-order optimization method that uses parallel activation perturbations as a “virtual population” to pretrain transformers. It reports experiments on how the approach behaves as model size and compute increase, as well as comparisons with backpropagation and an evolutionary-strategy method.
A company evaluating alternatives for model training could implement a comparison pilot on its own infrastructure: define tasks and metrics, record compute consumption and results, and track training progress in a dashboard. The practical goal is to determine whether the approach fits the business’s specific constraints and objectives, without assuming it replaces backpropagation in every setting.
How the consultancy can help
Wendelmaques can assess the opportunity and infrastructure constraints, define an evaluation scope, and implement and operate a training and monitoring pilot with agreed metrics.
Next step
Send a short description of your case to receive a scoped proposal for diagnosis, a pilot, and possible operation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal