Quantisation stores model weights with fewer bits — for example 8 or 4 bits instead of 16 — reducing memory and speeding up inference.
Why It Matters
A model's memory needs are roughly its parameter count times bytes per parameter. Quantising from 16 to 4 bits cuts memory by about three-quarters, letting larger models run on smaller hardware.
Types
- Post-training quantisation: convert a trained model directly. Quick and common.
- Quantisation-aware training: train with quantisation in mind for better quality.
Common Formats
Various formats and methods exist for different hardware and runtimes, such as GGUF for local inference tools and GPU-oriented methods for servers.
Quality Trade-Offs
- 8-bit quantisation usually has minimal quality loss.
- 4-bit is often acceptable, with some degradation.
- Lower precisions degrade more noticeably, especially for reasoning and smaller models.
Evaluate
Test quantised models on your tasks; average benchmark results can hide losses on specific skills.
Beyond Weights
The key-value cache can also be quantised to save memory on long contexts.