t-SNE and UMAP are techniques for visualising high-dimensional data — embeddings, gene expression, images — as 2D or 3D scatter plots.
How They Work
Both try to keep nearby points in the original space close together in the plot. They focus on local neighbourhoods rather than preserving global distances.
t-SNE
Produces clear, well-separated clusters. Slower on large datasets, and results depend on the perplexity setting.
UMAP
Generally faster, scales better and often preserves more global structure. Key settings are the number of neighbours and minimum distance.
Reading the Plots Carefully
- Cluster sizes don't mean much.
- Distances between clusters may not reflect real distances.
- Different random seeds and settings produce different pictures.
- Apparent clusters can appear in random data with some settings.
Good Practice
- Try several settings and seeds.
- Colour points by known labels to interpret.
- Use for exploration and hypothesis generation, not as proof.
Versus PCA
PCA is linear and preserves global variance; it's more faithful but often less visually separated. Using PCA first to reduce dimensions, then UMAP, is common.