TL;DR
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Q Labs Research reports that Dust, a zeroth-order training method, reached results competitive with backpropagation in its transformer language-model pretraining experiments. The authors say Dust uses token-level activation perturbations to estimate updates in parallel, but matching backprop closely required larger populations and substantially more compute.
Q Labs Research has reported that its method, Dust, can pretrain transformer language models without backpropagation and achieve results the authors describe as competitive with backprop in their experiments. The October 2026 report matters because it presents a way to estimate training updates from forward passes alone, while its findings remain experimental and rely on larger perturbation populations that require more compute.
Dust is a zeroth-order optimization method: it perturbs a model’s activations, measures how each perturbation affects loss, and combines perturbations weighted by their loss changes to estimate an update. Instead of perturbing model weights and separately evaluating each candidate, Dust perturbs activations at every token. The authors call this a virtual population, because one forward pass processes many token-level perturbations in parallel.
In the report’s pretraining experiments, Q Labs says Dust can approximate backpropagation more closely as the population grows and that it exceeds backprop in some tested settings. The authors also report that a 243-million-parameter model outperformed a model 120 times smaller at most population sizes they tested. These are results reported by the research team; the supplied material does not give enough detail to independently assess the training runs or reproduce the comparisons.
The report says Dust’s gradient estimates remained well aligned with backpropagation across tested scales, including experiments up to 1 billion tokens. It also compares Dust with EGGROLL, an evolution-strategy method that perturbs weights. Q Labs estimates that, from one million tokens onward, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL. That figure is an extrapolation in the report, not a measured general-purpose advantage over backpropagation.
A Different Route to Training
Backpropagation computes how changes to model parameters affect a loss, and it underpins the training of modern neural networks. Dust tests whether a less direct search procedure can train transformer language models at a competitive level. If later work confirms the reported results, activation perturbations could give researchers another way to study learning systems and the role of compute in training.
The practical case is not settled by competitiveness alone. The authors say Dust approaches backprop more closely at larger population sizes, which means more computation. Its reported efficiency advantage concerns EGGROLL, a different method, and does not establish that Dust is cheaper or faster than backprop on comparable hardware and tasks. The tradeoff between parallel forward-pass work and the missing backward pass will matter to any assessment of real-world usefulness.
high performance GPU for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Weight Search to Activations
Earlier evolution-strategy approaches search by perturbing model weights. The report says this can be expensive because each population member must be represented and evaluated. Dust instead draws on node perturbation, applying changes to activations, and uses token positions as population members that can be handled in parallel during a forward pass.
Q Labs frames the work against a common view that zeroth-order methods do not scale well to large networks. Its experiments challenge that expectation within the settings tested: the report says larger models were more population-efficient, and that alignment with backprop’s gradient estimates improved as the population grew. The report does not establish how those observations extend to other architectures, datasets, or large production training runs.
“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”
— Q Labs Research, in its report summary
large compute server for machine learning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The supplied report material does not specify enough experimental detail to determine how Dust compares with backprop under matched compute budgets, hardware, and training setups. It also does not establish whether the reported competitiveness holds across different language-model benchmarks or at larger scales. The claimed efficiency gap over EGGROLL is explicitly an extrapolation, and its comparison is with EGGROLL rather than backpropagation.
The report’s claims about better generalization or advantages in compute-rich settings are possibilities raised by the authors, not outcomes established by the described results. Further evidence would be needed to show whether Dust improves final model quality, training cost, or practical deployment outcomes.
transformer model training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Replication and Larger Runs
The report presents experimental findings but does not specify a next release, independent replication, or a date for follow-up results. The next useful evidence would include reproducible training details and comparisons that account for total compute, hardware, model quality, and performance across multiple tasks. Until then, Dust is a research result that challenges assumptions about zeroth-order training, rather than an established replacement for backpropagation.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Dust?
Dust is Q Labs Research’s zeroth-order method for training transformer language models. It perturbs activations and estimates updates from changes in loss, without computing a backpropagation gradient.
Does Dust outperform backpropagation?
The authors say Dust was competitive with backpropagation in their experiments and exceeded it in some settings. They also say closer approximation required larger populations and substantially more compute.
How does Dust compare with EGGROLL?
Q Labs estimates Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from one million tokens onward. The report describes this as an extrapolation, and it is not a comparison with backpropagation.
Has Dust been shown to work at every model scale?
No. The report describes tests up to one billion tokens and reports results for a 243-million-parameter model, but the supplied material does not establish performance across all model sizes, datasets, or training conditions.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
