Skip to content

Embeddings, semantic search and RAG

Turn text into vectors, build a semantic search index, and ground an LLM's answers in your own documents.

Free on glitchdata advanced 4 lessons 1 hr 10 min

What you'll learn

  • Explain what embeddings are and how similarity is measured
  • Build a small semantic search index in Python
  • Chunk documents sensibly for retrieval
  • Design a RAG pipeline and know where it fails

About this course

Retrieval-augmented generation (RAG) is the most common way to make a language model answer questions about your own documents. This course builds it from the ground up: embeddings, similarity search, chunking and finally the retrieval-plus-generation loop.

The examples use the all-MiniLM-L6-v2 model from the hub. Install with pip install sentence-transformers numpy.

Before you start

  • Comfortable with Python
  • Prompting large language models

Course content

4 lessons · 1 hr 10 min

  1. 1
    What an embedding is

    Meaning as a position in space, and cosine similarity.

    Free preview 12 min
  2. 2
    Semantic search in twenty lines

    Embed a set of passages and rank them against a question.

    20 min
  3. 3
    Chunking and indexing documents

    How to split long documents so retrieval finds the right piece.

    18 min
  4. 4
    Retrieval-augmented generation

    Put retrieval and generation together, and evaluate the result.

    20 min

What learners say

Sign in and enrol to leave a review.

No reviews yet — be the first once you have worked through it.

More in Generative AI