Blog
DUST estimates the gradient of a blurred loss
DUST pretrains transformers with forward passes alone. On average, it follows the gradient backprop would compute on a blurred copy of the loss.
Pick one layer of a language model, add a little Gaussian noise to its output, and run the forward pass to see whether the loss went up or down. Repeat this a few thousand times with fresh noise and average the noise draws, weighting each by how much it raised the loss: the weighted average estimates the gradient, and an optimizer can use it as it would use backprop’s.
Methods of this kind are called zeroth order, because they learn from values of the loss alone and run entirely on forward passes.
Zeroth-order methods
The idea goes back to Kiefer and Wolfowitz (1952), who estimated a gradient by nudging one coordinate at a time. Forty years later Spall (1992) nudged every coordinate at once along a random direction, Williams (1992) carried the method into reinforcement learning, and Widrow and Lehr (1990) moved the noise onto the neurons, a variant now called node perturbation. Evolution strategies later scaled it to over a thousand parallel workers, and MeZO uses it today to fine-tune language models.
Because it runs on forward passes, the method fits in the memory needed to run the model, and it can train through steps that backprop finds hard, such as a discrete choice, a program in the loop, or a recurrence too long to unroll.
Each forward pass, however, returns a single number, and in \(d\) dimensions a random direction is almost at right angles to the gradient, so the noise in the estimate grows with \(d\) (Nesterov and Spokoiny, 2017). A model with a hundred million weights would therefore need a population of roughly a hundred million, and most researchers concluded that the method suits small networks and fine-tuning while pretraining belongs to backprop (Lillicrap et al., 2020).
Per-token activation noise
DUST (paper, GitHub, tweet), from Samip Dahal and colleagues at Q Labs, applies this idea to pretraining transformers. It puts the noise on the activations, as node perturbation does, and draws fresh noise at every token, so a sequence of 2,048 tokens holds 2,048 separate guesses and one forward pass tests them all, an arrangement the authors call a virtual population.
DUST finishes ahead of backprop at 100k and 1M training tokens, and by the authors’ estimate it is a thousand to ten thousand times as efficient as EGGROLL, the strongest evolution strategy that perturbs weights. Larger models also do better at a fixed population: at most population sizes a 243M-parameter model beats one 120 times smaller, the reverse of what the dimension argument predicts.
The gradient estimator
For a linear layer \(y_t = W x_t\), DUST jitters the output at every token, \(y_t \to y_t + \sigma a_t\) with \(a_t \sim \mathcal N(0, I)\), and runs the forward pass. Because the losses before token \(t\) are already fixed, a jitter at \(t\) reaches the loss at \(t\) and, through attention, every loss after it, so DUST scores each jitter by the losses that follow:
\[F_t = \sum_{s \ge t} \gamma^{\,s-t}\, \ell_s .\]
The discount \(\gamma\) sets how far the credit reaches: at \(\gamma = 0\) a jitter answers for its own token alone, and at \(\gamma = 1\) it answers for every token after it, as it does in backprop.
With \(K\) independent draws, DUST compares each draw’s score with the population average \(\bar F_t\) and weights its noise by the difference:
\[\hat g_t = \frac{1}{K\sigma} \sum_{i=1}^{K} \bigl(F_t^{(i)} - \bar F_t\bigr)\, a_{i,t} .\]
Draws that scored above the average pull the estimate toward their noise and draws below it push the estimate away, so the sum is a noisy estimate of the error signal at the layer’s output. Multiplying it by the layer’s input and summing over tokens gives a weight gradient, \(\hat G_W = \sum_t \hat g_t\, x_t^\top\), the same outer product backprop builds from the same input; the difference is that backprop computes the output error with the chain rule, while DUST measures it.
Expectation of the estimator
The estimate has a simple expectation:
\[\begin{aligned} \mathbb E\bigl[\hat g_t\bigr] &= \frac{K-1}{K}\, \nabla_{y_t} F_{t,\sigma}(y), \\[4pt] F_{t,\sigma}(y) &= \mathbb E_A\bigl[F_t(y + \sigma A)\bigr], \end{aligned}\]
and its variance falls as \(1/K\). On average, then, DUST follows the gradient of a blurred loss, the loss you would get if noise of size \(\sigma\) stayed on the layer’s output for good, scaled down by \((K-1)/K\).
The proof starts from the fact that each draw’s noise is independent of the other draws, so it correlates with its own score alone, and Stein’s lemma turns that correlation into \(\sigma\) times the gradient of the blurred score. Each draw is also one of the \(K\) terms in the average it is compared with, so it cancels \(1/K\) of its own signal, and together these two facts give the theorem.
Two limits follow from this result: as \(\sigma\) shrinks, the blurred gradient approaches the ordinary one with an error of order \(\sigma^2\), and with \(\gamma = 1\) the ordinary gradient at \(y_t\) is backprop’s output error, so the expected weight update matches \(\partial L/\partial W\). Although DUST uses forward passes and averages alone, in expectation it lands on the chain rule’s answer, which makes it backprop on a blurred loss, and on a shortsighted one when \(\gamma < 1\).
The \((K-1)/K\) factor
The factor \((K-1)/K\) comes from the baseline, because DUST compares each draw with an average that includes it, much as a test graded on a curve lets your own score help set the curve. Your score pulls the class average toward you, so you lose \(1/K\) of your margin, and in a class of two you lose half.
With 16,384 draws the factor is \(1 - 1/16{,}384\), close enough to one, but the released code centers rewards within small chunks of draws that share a forward pass, two to sixteen draws at a time (DRAW_CHUNK_SIZES). At a population of 256 each chunk holds two draws, so the direct-draw estimates come out at half the gradient they aim for, a discrepancy the tuned learning rate absorbs. Comparing each draw with the average of the other draws removes it, since that comparison multiplies the estimate by \(K/(K-1)\).
Consequences
Cosine similarity and population size
Q Labs measure the cosine between DUST’s estimate and backprop’s gradient as the population grows and fit \(\cos(K) = c_{\max}/\sqrt{1 + c/K}\), the curve traced by any estimator whose mean stays fixed while its variance falls as \(1/K\). The theorem gives both constants a meaning: the ceiling \(c_{\max}\) is the cosine between backprop’s gradient and the blurred, discounted gradient DUST aims at, and \(c\) is the noise-to-signal ratio of a single draw. In a toy simulation a single value of \(c\) fits every population from 16 to 16,384 and matches the ratio measured on single draws. The announcement describes backprop-like gradients as emerging from a large population, whereas the theorem places them in the expectation at every population size, leaving the population the job of averaging away the noise around them.
The blur as a curvature penalty
For small \(\sigma\) the blurred loss is the loss plus a curvature penalty,
\[F_{t,\sigma}(y) \approx F_t(y) + \frac{\sigma^2}{2}\, \operatorname{tr} \nabla_y^2 F_t(y) ,\]
so DUST descends the loss plus a penalty on the trace of its Hessian with respect to the layer’s activations, which favors flat regions, a known effect of noise injection (Bishop, 1995; Orvieto et al., 2022). The paper’s first open question is whether DUST picks up “higher-order curvature that pulls the search toward flat regions,” and this term is a natural candidate. It is also easy to test by training with backprop while adding the same noise to the activations, a run that follows the same blurred gradient in expectation, minus the Monte Carlo noise. If that run beats clean backprop, the blur explains the edge; if it ties or loses, the explanation lies in the shortsighted credit or in the noise of the estimate itself.
Non-differentiable networks
The identity behind the theorem differentiates the Gaussian and leaves \(F\) alone, so the blurred loss stays smooth even when the network has jumps and kinks, and DUST’s target remains well defined for networks with argmax, sampling, tool calls, or a program in the loop. Zeroth-order methods have always promised that freedom, and the theorem states what DUST optimizes when it uses it.
Relation to RetNet and Gated DeltaNet-2
The discount also has a structural reading. Written as a recurrence, the score runs from the last token back to the first,
\[F_t = \ell_t + \gamma\, F_{t+1} ,\]
and the same recurrence run forward is the memory of RetNet (Sun et al., 2023), \(S_t = \gamma\, S_{t-1} + k_t v_t^\top\), in which each token adds the outer product of its key and value to a running sum that fades by \(\gamma\) for every token that passes. DUST’s credit is that memory run backward, with each token’s loss in place of its key and value and a single number for the state.
The resemblance goes deeper than the shared recurrence. The gradient of the score sums the gradients of the later losses, each discounted by its distance,
\[\nabla_{y_t} F_t = \sum_{s \ge t} \gamma^{\,s-t}\, \frac{\partial \ell_s}{\partial y_t} ,\]
and each of those gradients is a sum over routes through the network from token \(t\) to token \(s\). A route moves forward in time only where a layer mixes tokens, as attention does, so its crossings from token to token add up to \(s - t\) however many layers it passes through, and damping each crossing by \(\gamma\) for every token it spans multiplies every route by exactly \(\gamma^{s-t}\). The gradient of the score is therefore the gradient backprop would compute for the total loss if every crossing were damped this way on the way back, with the forward pass left alone. In a layer of softmax attention, where token \(s\) reads token \(t\) with weight \(A_{st}\), the backward pass sees that weight as \(\gamma^{s-t} A_{st}\), which is the attention matrix multiplied entry by entry by RetNet’s decay mask. At \(\gamma = 0\) only the diagonal survives, and each token answers for its own loss as if no token could read another; at \(\gamma = 1\) the attention comes back whole.
Recurrent layers take the damping in an even simpler form. Gated DeltaNet-2 (Hatamizadeh et al., 2026), a recent linear-attention layer from NVIDIA, replaces the growing cache of keys and values with a memory of fixed size that it edits in place: at each token it decays the memory channel by channel, erases what is stored along the token’s key, and writes the new value, each under its own gate,
\[S_t = \bigl(I - k_t (b_t \odot k_t)^\top\bigr) \operatorname{Diag}(\alpha_t)\, S_{t-1} + k_t (w_t \odot v_t)^\top ,\]
with \(\alpha_t\) setting the decay, \(b_t\) the erase, and \(w_t\) the write. A route through this memory crosses from one token to the next through one copy of the update at a time, so damping each crossing by \(\gamma\) turns \(\operatorname{Diag}(\alpha_t)\) into \(\operatorname{Diag}(\gamma \alpha_t)\): for this layer, DUST’s discount amounts to backprop through the same layer forgetting an extra 2% per token at \(\gamma = 0.98\), a shift of \(\ln \gamma \approx -0.02\) in the log-decay the layer already learns. On average, then, DUST follows backprop on a blurred loss through a network whose memory fades by an extra factor of \(\gamma\) per token on the way back.
Which activations get the discount
The released code gives the long discount only to activations that later tokens read: attention keys, values, value embeddings, and their gates get \(\gamma = 0.98\), while queries and every other layer answer for their own token alone. The rule follows memory rather than layer type, and Gated DeltaNet-2 sharpens it: its decay and erase gates store nothing, yet a jitter on either changes what every later token reads, so they would need the discount as well.
Summary
On average, DUST’s estimate is \((K-1)/K\) times backprop’s error signal on a blurred, discounted loss, which leaves three knobs between DUST and backprop: the noise scale \(\sigma\) blurs the loss, the discount \(\gamma\) cuts the credit short, and the population \(K\) sets the noise. Turning \(\sigma\) down, setting \(\gamma\) to one, and correcting for \(K\) recovers backprop’s gradient.
The first two knobs may well help training, while the third costs compute, and reducing that cost is the open problem. Kiefer and Wolfowitz wrote down the first of these estimators in 1952, and seventy-four years later one of them has pretrained a transformer.
The figures use a toy one-dimensional loss and a made-up attention pattern, and compute the estimator, the blur, the \((K-1)/K\) factor, and the discounted attention from the formulas above.