Neural networks are the technology behind modern image recognition, speech recognition and language models. The core idea is surprisingly simple.
The Artificial Neuron
A neuron takes several numbers as input, multiplies each by a weight, adds them up with a bias, and passes the result through an activation function. The weights and biases are the parameters learned during training.
Layers
Neurons are arranged in layers. The input layer receives the features; one or more hidden layers transform them; the output layer produces the prediction. A network with many hidden layers is called deep.
Why Activation Functions Matter
Without a non-linear activation function, stacking layers would be no more powerful than a single linear equation. Functions such as ReLU (which replaces negative values with zero) let networks model complex, curved relationships.
How Networks Learn
Training uses backpropagation and gradient descent: the network makes predictions, a loss function measures the error, and the error is traced back through the layers to work out how each weight should change. Repeating this over many examples gradually improves the network.
Specialised Architectures
- Convolutional networks (CNNs) for images.
- Recurrent networks (RNNs) for sequences, now largely replaced by transformers.
- Transformers for language and increasingly for images and audio.
When to Use Them
Neural networks excel with large amounts of unstructured data. For small tabular datasets, simpler methods such as gradient boosting are often as good and easier to explain.