Urgent.News

What's breaking now, across thousands of outlets.

AI

Dust: Pretraining Transformers Without Backpropagation

The article discusses the potential of a new learning algorithm called Dust, which aims to replace the traditional backpropagation method in training deep learning models. Backpropagation, while effective in low-compute scenarios, becomes less efficient in high-compute regimes and limits the exploration of possible architectures.

Dust is a zeroth-order optimization algorithm that perturbs activations instead of weights, assigning rewards to these perturbations based on their impact on the loss. This approach allows for parallel evaluation of perturbations, significantly increasing the computational scale compared to traditional methods. Dust operates by adding Gaussian noise to the output of each linear layer in a transformer model, independently at every token, and evaluating the impact on the loss.

This virtual population approach avoids the need to materialize each member, drastically reducing computational costs. The algorithm is designed to be competitive with backpropagation on pretraining transformer models, although it does not aim to be compute-efficient enough for immediate practical implementation. The authors also note that Dust can be used to train more flexible neural network architectures that are difficult to optimize using backpropagation, such as those with external program loops or transformers processed over multiple steps.

While Dust shows promise in theory, the authors caution that it is not yet optimized for computational efficiency and that the development of new neural network types that can leverage Dust's capabilities is a topic for future research.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at qlabs.sh →

More in AI

More from Monday 5 October →