Skip to content

t-SNE and UMAP for Visualising High-Dimensional Data

How t-SNE and UMAP project complex data into two dimensions for exploration, and how to avoid misreading the plots.

Editorial team 1 min read

t-SNE and UMAP are techniques for visualising high-dimensional data — embeddings, gene expression, images — as 2D or 3D scatter plots.

How They Work

Both try to keep nearby points in the original space close together in the plot. They focus on local neighbourhoods rather than preserving global distances.

t-SNE

Produces clear, well-separated clusters. Slower on large datasets, and results depend on the perplexity setting.

UMAP

Generally faster, scales better and often preserves more global structure. Key settings are the number of neighbours and minimum distance.

Reading the Plots Carefully

  • Cluster sizes don't mean much.
  • Distances between clusters may not reflect real distances.
  • Different random seeds and settings produce different pictures.
  • Apparent clusters can appear in random data with some settings.

Good Practice

  • Try several settings and seeds.
  • Colour points by known labels to interpret.
  • Use for exploration and hypothesis generation, not as proof.

Versus PCA

PCA is linear and preserves global variance; it's more faithful but often less visually separated. Using PCA first to reduce dimensions, then UMAP, is common.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026