Convolutional neural networks (CNNs) are designed for grid-like data such as images. They were the engine of the deep learning revolution in computer vision.
Convolutions
A convolutional layer slides small filters — for example 3×3 pixels — across the image. Each filter detects a pattern, such as an edge at a particular angle, wherever it appears. Because the same filter is reused across the image, CNNs need far fewer parameters than fully connected networks and naturally recognise patterns regardless of position.
Pooling
Pooling layers downsample feature maps (for example, taking the maximum in each 2×2 block), making the network more robust to small shifts and reducing computation.
A Hierarchy of Features
Stacked layers learn increasingly abstract features: edges and colours in early layers, textures and shapes in the middle, object parts and whole objects near the end.
Landmark Architectures
LeNet, AlexNet, VGG, ResNet (which introduced skip connections enabling very deep networks), Inception and EfficientNet. Vision transformers now rival CNNs, but CNNs remain widely used and efficient.
Beyond Classification
CNN backbones power object detection, segmentation, image similarity and medical imaging, and apply to audio spectrograms and some time series.
In Practice
Start from a pretrained CNN and fine-tune it, rather than training from scratch. Prepare images exactly as the model expects: size, colour order and normalisation.