Skip to content

Lesson 1 of 4

Free preview

What an embedding is

Meaning as a position in space, and cosine similarity.

12 min 3-question quiz 3 guides to read next

On this page
  1. Measuring closeness
  2. What embeddings are good for

An embedding is a list of numbers — a vector — that represents a piece of text so that texts with similar meaning get similar vectors. The all-MiniLM-L6-v2 model turns any sentence into 384 numbers.

Think of each text as a point in a 384-dimensional space. "How do I reset my password?" and "I forgot my login details" share almost no words, but they land close together. "Password requirements for new accounts" is nearby; "Our office opening hours" is far away.

Measuring closeness

The usual measure is cosine similarity: the cosine of the angle between two vectors. It ranges from −1 to 1, and higher means more similar. If vectors are normalised to length 1, cosine similarity is simply their dot product, which is very fast to compute.

What embeddings are good for

  • Semantic search: find passages that answer a question, even with different wording.
  • Clustering: group similar support tickets or documents.
  • Deduplication: spot near-identical records.
  • Retrieval for LLMs: the "R" in RAG.

Keep in mind that an embedding model has its own limits: all-MiniLM-L6-v2 is trained on English and truncates input after 256 word pieces, so long documents must be split into chunks.

Check your understanding

3 questions · pass with 2 correct

1. What does an embedding represent?
2. If vectors are normalised to length 1, cosine similarity equals…
3. all-MiniLM-L6-v2 truncates input after 256 word pieces. What follows?

You'll see your score; enrol to have it count towards your certificate.

Enrol for free to save your progress, unlock every lesson and earn a certificate.

Sign in to enrol

Further reading

Guides that go deeper on this lesson.

  • Embeddings Explained

    Embeddings turn text, images and other data into vectors where similar meanings sit close together. How they work and what they're used for.

    2 min read

  • Choosing an Embedding Model for RAG

    The factors that matter when picking an embedding model for retrieval: language, domain, dimensions, input length, cost and licence.

    2 min read

  • Vector Databases and Similarity Search

    How vector indexes find the nearest embeddings quickly, the main options, and when you don't need one.

    2 min read