Lesson 1 of 4
How speech recognition works
From sound waves to spectrograms to text.
12 min 3-question quiz 3 guides to read next
Audio arrives as a waveform: thousands of air-pressure measurements per second. Whisper expects 16,000 samples per second (16 kHz), and converts each 30-second window into a log-Mel spectrogram — a picture of which frequencies are loud at each moment, on a scale that matches human hearing.
Whisper is an encoder–decoder model:
- The encoder reads the spectrogram and builds a representation of what was said.
- The decoder writes text one token at a time, like a language model, guided by the encoder's output.
Special tokens at the start of the decoder's output tell it the language, whether to transcribe (same language) or translate (into English), and whether to include timestamps.
OpenAI trained Whisper on 680,000 hours of audio collected from the web, which is why it copes with accents, background noise and technical vocabulary better than many older systems. It comes in several sizes — tiny, base, small, medium, large — trading speed for accuracy.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Speech Recognition With AI
How automatic speech recognition works, choosing a model, and measuring accuracy with word error rate.
2 min read
-
Attention and the Transformer Architecture
The mechanism behind modern language models: self-attention, multi-head attention and positional information, explained intuitively.
2 min read
-
Multimodal AI Models
Models that understand and generate across text, images, audio and video — what they can do and how to use them.
2 min read