Lesson 15 / 28

Clustering, Topics and De-duplication

Use embeddings to group similar items and find near-duplicates.

Neighbours are groups and duplicates

Embeddings are useful beyond search. Clustering (for example k-means or HDBSCAN) groups vectors into topics, so you can summarise a large pile of feedback, tickets or documents. Near-duplicate detection flags pairs whose cosine similarity is above a high threshold (such as 0.95 or 0.98) to remove repeated content or merge duplicate tickets; choose the threshold by inspecting examples, since it depends on the model. Classification can be done by training a small classifier on embeddings or by finding the nearest labelled example. Anomaly detection looks for vectors far from every cluster. In all cases, look at real examples from each group, because clusters are mathematical groupings and may not match the categories people care about.

Grouping sentences by topic, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. K-means with 3 clusters on 4-dimension embeddings of the 12 sentences separates security, leave and travel sentences cleanly. (With 6 dimensions the same method mixed a few leave and travel sentences, so representation size matters.)

DOCS = [
 "Employees get 24 days of annual leave and may take a vacation any time of the year",
 "Unused vacation days and leave can be carried over to the next year",
 "Public holidays are paid days off and do not count as annual leave",
 "Sick leave needs a doctor note after three days of absence",
 "Flights must be booked fourteen days in advance for business travel",
 "Hotel stays on a business trip are capped at six thousand rupees per night",
 "Travel expenses need receipts and a trip report within thirty days",
 "Taxi and train tickets are reimbursed for a business trip",
 "Use a password manager and enable two factor authentication on every account",
 "Report a lost laptop or phone to the security team within twenty four hours",
 "Never share a password and always lock your laptop screen",
 "Phishing emails should be reported to the security team immediately",
]
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize
from sklearn.cluster import KMeans

X = TfidfVectorizer(stop_words="english").fit_transform(DOCS)
E = normalize(TruncatedSVD(n_components=4, random_state=0).fit_transform(X))
labels = KMeans(n_clusters=3, n_init=10, random_state=0).fit_predict(E)
groups = {}
for doc, lab in zip(DOCS, labels):
    groups.setdefault(int(lab), []).append(doc.split()[0] + " " + doc.split()[1] + " " + doc.split()[2])
for lab in sorted(groups):
    print("cluster", lab, "->", groups[lab])

Output:

cluster 0 -> ['Use a password', 'Report a lost', 'Never share a', 'Phishing emails should']
cluster 1 -> ['Employees get 24', 'Unused vacation days', 'Public holidays are', 'Sick leave needs']
cluster 2 -> ['Flights must be', 'Hotel stays on', 'Travel expenses need', 'Taxi and train']

Near-duplicate detection, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. Only the pair a and a_copy (cosine 0.999) is above the 0.98 threshold and is flagged. The pairs at 0.707 and 0.742 are related but distinct.

import numpy as np
V = {"a": [1.0, 0.0], "a_copy": [0.99, 0.05], "b": [0.0, 1.0], "c": [0.7, 0.7]}
names = list(V); M = np.array([V[n] for n in names]); M /= np.linalg.norm(M, axis=1, keepdims=True)
S = M @ M.T
for i in range(len(names)):
    for j in range(i + 1, len(names)):
        flag = "DUPLICATE" if S[i, j] > 0.98 else ""
        print(f"{names[i]:7} {names[j]:7} cosine={S[i, j]:.3f} {flag}")

Output:

a       a_copy  cosine=0.999 DUPLICATE
a       b       cosine=0.000 
a       c       cosine=0.707 
a_copy  b       cosine=0.050 
a_copy  c       cosine=0.742 
b       c       cosine=0.707

Quick check: How should a near-duplicate threshold be chosen?

  • By inspecting real pairs for your model, since scales differ
  • Always 0.5
  • Always exactly 1.0
  • At random
Answer

By inspecting real pairs for your model, since scales differ — Similarity distributions vary by embedding model and data.