Lesson 3 / 28
Similarity Measures: Dot, Cosine and Euclidean
Compare vectors with the three common measures and know when they agree.
Direction, size, or both
Three measures dominate. Dot product multiplies matching components and adds them: it grows with both the angle between vectors and their lengths. Cosine similarity divides the dot product by both lengths, so it measures direction only (1 = same direction, 0 = unrelated, -1 = opposite). Euclidean (L2) distance is the straight-line distance (smaller = closer). For unit-length (normalised) vectors, all three agree on ranking: cosine = dot, and L2² = 2 - 2·cosine. Many embedding models output normalised vectors, or are meant to be used with cosine; follow the model's documentation, and normalise yourself if unsure so that a fast dot product gives cosine ranking.
The three measures on small vectors, run
I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. Vector b is a twice-as-long copy of a, and c is perpendicular to a. Dot product gives 50 for a·b but 0 for a·c; cosine says a and b point the same way (1.0) while a and c are unrelated (0.0); L2 says b is 5 away and c is 7.07 away. After normalising, dot equals cosine.
import numpy as np
a = np.array([3.0, 4.0]); b = np.array([6.0, 8.0]); c = np.array([4.0, -3.0])
def dot(x, y): return float(x @ y)
def cos(x, y): return float(x @ y / (np.linalg.norm(x) * np.linalg.norm(y)))
def l2(x, y): return float(np.linalg.norm(x - y))
print("dot a.b, a.c :", dot(a, b), dot(a, c))
print("cosine a,b / a,c:", round(cos(a, b), 3), round(cos(a, c), 3))
print("L2 a,b / a,c :", round(l2(a, b), 3), round(l2(a, c), 3))
an, bn = a / np.linalg.norm(a), b / np.linalg.norm(b)
print("after normalising, dot == cosine:", round(dot(an, bn), 3), "==", round(cos(a, b), 3))
Output:
dot a.b, a.c : 50.0 0.0 cosine a,b / a,c: 1.0 0.0 L2 a,b / a,c : 5.0 7.071 after normalising, dot == cosine: 1.0 == 1.0
When the rankings agree, run
I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. On 1,000 random vectors with very different lengths, raw dot product and raw L2 pick different neighbours from cosine, because length distorts them. After normalising to unit length, L2 gives exactly the cosine ranking.
import numpy as np
rng = np.random.default_rng(0)
docs = rng.normal(size=(1000, 32)) * rng.uniform(0.5, 5.0, size=(1000, 1)) # vectors with very different lengths
q = rng.normal(size=32)
by_dot = np.argsort(-(docs @ q))[:5]
by_l2 = np.argsort(np.linalg.norm(docs - q, axis=1))[:5]
dn = docs / np.linalg.norm(docs, axis=1, keepdims=True); qn = q / np.linalg.norm(q)
by_cos = np.argsort(-(dn @ qn))[:5]
dn_l2 = np.argsort(np.linalg.norm(dn - qn, axis=1))[:5]
print("dot (raw) :", by_dot.tolist())
print("L2 (raw) :", by_l2.tolist())
print("cosine :", by_cos.tolist())
print("L2 on unit vecs:", dn_l2.tolist(), "| same as cosine:", dn_l2.tolist() == by_cos.tolist())
Output:
dot (raw) : [353, 868, 15, 456, 130] L2 (raw) : [25, 226, 20, 590, 550] cosine : [859, 36, 25, 226, 20] L2 on unit vecs: [859, 36, 25, 226, 20] | same as cosine: True
Match the metric to the model
Use the similarity the model was trained for (usually cosine). In a vector store, configure the same metric when you create the index.
Quick check: For unit-length vectors, what is true?
- L2 gives the opposite ranking
- Cosine is always zero
- Dot product is meaningless
- Dot product equals cosine similarity, and L2 gives the same ranking
Answer
Dot product equals cosine similarity, and L2 gives the same ranking — Normalising removes length, leaving only direction.