Lesson 1 of 4
How computers see images
Pixels, convolutions and why pretrained networks transfer so well.
12 min 3-question quiz 3 guides to read next
To a computer, a colour image is a grid of numbers: height × width × 3 colour channels (red, green, blue), usually 0–255 each. A 224 × 224 photo is over 150,000 numbers.
Convolutional neural networks (CNNs) process that grid with small filters — typically 3 × 3 pixels — that slide across the image looking for a pattern. Early layers learn to detect edges and colour blobs; middle layers combine them into textures and parts such as eyes or wheels; late layers respond to whole objects.
ResNet ("residual network") added skip connections that let information bypass layers. That made it practical to train much deeper networks, and ResNet-50 — 50 layers deep — became one of the most widely used vision models.
Why pretrained models transfer
A network trained on ImageNet (over a million photos in 1,000 categories) has learned general visual features — edges, textures, shapes — that are useful for almost any image task. That is why you can take such a model and, with relatively few examples, teach it to recognise things it never saw: defects on a production line, plant diseases, types of document.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Convolutional Neural Networks
How CNNs recognise patterns in images using filters and pooling, and why pretrained CNNs are so widely reused.
2 min read
-
Image Classification With Pretrained Models
Classifying images without training from scratch: choosing a pretrained model, preparing images correctly and reading the output.
2 min read
-
Transfer Learning
Reusing a model trained on one task as the starting point for another — the technique that makes deep learning practical with limited data.
2 min read