search
backpropagation
Trends
- 1
Researchers at QLabs have published Dust, a method for pretraining transformer models without using backpropagation, the algorithm that underpins almost all modern AI training. The work, described on the company's research page, is drawing attention from machine learning practitioners debating whether such approaches could reduce the heavy memory and compute costs of training large models.
- 2Dust: Pretraining Transformers Without Backpropagation●Dust: Pretraining Transformers Without Backpropagation https://qlabs.sh/research/dust # HackerNews # Tech # AI
A research project called Dust claims a method for pretraining transformer models without backpropagation, the algorithm at the core of modern deep learning. If the results hold up, the approach could challenge assumptions about how large models must be trained, but independent verification and details of its performance are not yet established.
- 3
A technical article explaining the softmax function and how to compute its derivative has drawn attention among developers and machine-learning practitioners. Softmax converts a vector of raw scores into a probability distribution and is central to classification models and neural network outputs. The piece walks through the mathematics of its gradient, including the coupling between outputs, which often trips people up when deriving backpropagation by hand.