Lesson 25 / 28
Classification, Routing and Few-Shot Labelling
Use embeddings as features for cheap, fast classifiers.
A small classifier on top of vectors
Embeddings turn text classification into a simple problem. Embed labelled examples and train a lightweight classifier (logistic regression, k-nearest neighbours) on the vectors: it needs far fewer labels than training a text model from scratch, runs in microseconds, and is easy to retrain. A zero-training variant is nearest-prototype routing: embed a few example sentences per category, average them into a centroid, and assign a new text to the nearest centroid, with a similarity threshold below which the answer is "unknown / send to a human". This works well for intent routing, tagging and triage. Check accuracy per class with a confusion matrix on held-out examples, and watch for classes that the embedding does not separate.
Nearest-centroid routing (illustrative)
Assumes embed() returns unit vectors; not run here because it needs a real embedding model.
import numpy as np
examples = {
"billing": ["I was charged twice", "refund my invoice"],
"technical": ["the app crashes on start", "cannot log in"],
}
centroids = {}
for label, texts in examples.items():
c = np.mean([embed(t) for t in texts], axis=0)
centroids[label] = c / np.linalg.norm(c)
def route(text, threshold=0.45):
v = embed(text)
label, score = max(((l, float(v @ c)) for l, c in centroids.items()), key=lambda x: x[1])
return label if score >= threshold else "unknown" # below threshold: send to a humanQuick check: Why include a similarity threshold in nearest-centroid routing?
- To avoid embeddings
- To make vectors longer
- So off-topic inputs become "unknown" instead of a forced wrong label
- To speed up training
Answer
So off-topic inputs become "unknown" instead of a forced wrong label — Nearest-neighbour methods always return something; a threshold allows abstaining.