Recurrent neural networks (RNNs) were the standard deep learning approach to sequences — text, speech, time series — before transformers.
How RNNs Work
An RNN reads a sequence one element at a time, maintaining a hidden state that summarises what it has seen so far. At each step it combines the new input with the previous state.
The Vanishing Gradient Problem
When training on long sequences, the signal from early steps fades as it is propagated back through many time steps. Plain RNNs therefore struggle to learn long-range dependencies.
LSTMs and GRUs
Long short-term memory networks add gates that control what information to keep, forget and output, allowing them to remember across longer spans. Gated recurrent units (GRUs) are a simpler variant with similar performance. For years they powered machine translation, speech recognition and text generation.
Limitations
- Sequential processing is hard to parallelise, so training is slow.
- Very long dependencies remain difficult.
Why Transformers Took Over
Transformers process all positions in parallel using attention, letting each element directly relate to every other. They train faster on modern hardware and handle long-range context better, so they replaced RNNs for most language tasks.
Where RNNs Still Appear
Small, efficient RNNs remain useful for some streaming and on-device applications, and for certain time-series problems. Recent state-space models revisit the recurrent idea with better long-range performance.