Skip to content

Quantisation of Language Models

How reducing numerical precision makes models smaller and faster, and the quality trade-offs involved.

Editorial team 1 min read

Quantisation stores model weights with fewer bits — for example 8 or 4 bits instead of 16 — reducing memory and speeding up inference.

Why It Matters

A model's memory needs are roughly its parameter count times bytes per parameter. Quantising from 16 to 4 bits cuts memory by about three-quarters, letting larger models run on smaller hardware.

Types

  • Post-training quantisation: convert a trained model directly. Quick and common.
  • Quantisation-aware training: train with quantisation in mind for better quality.

Common Formats

Various formats and methods exist for different hardware and runtimes, such as GGUF for local inference tools and GPU-oriented methods for servers.

Quality Trade-Offs

  • 8-bit quantisation usually has minimal quality loss.
  • 4-bit is often acceptable, with some degradation.
  • Lower precisions degrade more noticeably, especially for reasoning and smaller models.

Evaluate

Test quantised models on your tasks; average benchmark results can hide losses on specific skills.

Beyond Weights

The key-value cache can also be quantised to save memory on long contexts.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026