Lesson 1 / 28

From Words to Coordinates

Understand the central idea: similar meaning, nearby vectors.

Closeness stands for similarity

An embedding is a list of numbers (a vector) that represents an item, such as a sentence, a document, an image or a product, so that items with similar meaning get vectors that are close together and unrelated items get vectors that are far apart. Computers cannot compare the meaning of two sentences directly, but they can compare two lists of numbers quickly. Once text is embedded, tasks such as semantic search ("find passages about holidays even if they say vacation"), recommendation, clustering, duplicate detection, classification and retrieval for LLMs (RAG) become geometry: find the nearest vectors. Typical text embeddings have a few hundred to a few thousand dimensions.

Meaning as a point in space

An embedding maps text, images or other items to a vector so that similar items sit close together.

Four ideas: vector, distance, model, dimensions.
Figure 1.1 — Vector, distance, model and dimensions.

A map of ideas

On a city map, nearby places are close in real life. An embedding is a map of meaning: sentences about leave land in one neighbourhood and sentences about travel in another.

Embeddings are not the text

A vector cannot be turned back into the exact text. Always store the original text (or a pointer to it) next to its vector.

Quick check: What should be true of the embeddings of two sentences with similar meaning?

  • Their vectors are far apart
  • Their vectors are identical
  • Their vectors are close together
  • They have the same length in characters
Answer

Their vectors are close together — Embedding models are trained so that similar meaning gives nearby vectors.