पाठ 15 / 28

Clustering, विषय और De-duplication

समान items समूहित करने और लगभग-duplicates खोजने को embeddings उपयोग करें।

पड़ोसी समूह और duplicates हैं

Embeddings खोज के अलावा भी उपयोगी हैं। Clustering (जैसे k-means या HDBSCAN) vectors को विषयों में समूहित करती है, ताकि आप feedback, tickets या दस्तावेज़ों के बड़े ढेर का सार कर सकें। Near-duplicate पहचान उन जोड़ियों को चिह्नित करती है जिनकी cosine similarity ऊँची सीमा (जैसे 0.95 या 0.98) से ऊपर हो, दोहरी सामग्री हटाने या duplicate tickets मिलाने के लिए; सीमा उदाहरण देखकर चुनें, क्योंकि वह मॉडल पर निर्भर है। वर्गीकरण embeddings पर छोटा classifier प्रशिक्षित करके या निकटतम labelled उदाहरण खोजकर हो सकता है। Anomaly पहचान हर cluster से दूर vectors खोजती है। हर मामले में हर समूह के असली उदाहरण देखें, क्योंकि clusters गणितीय समूह हैं और उन श्रेणियों से मेल नहीं खा सकते जिनकी लोगों को परवाह है।

विषय से वाक्य समूहित करना, चलाकर

मैंने यह Python virtual environment में numpy 2.5.3, scikit-learn 1.9.1 और faiss-cpu 1.15.1 के साथ चलाया, निश्चित random seeds के साथ ताकि संख्याएँ दोहराई जाएँ। 12 वाक्यों के 4-आयामी embeddings पर 3 clusters वाला k-means सुरक्षा, leave और travel वाक्यों को साफ़ अलग करता है। (6 आयामों के साथ वही तरीक़ा कुछ leave और travel वाक्य मिला देता था, इसलिए प्रतिनिधित्व का आकार मायने रखता है।)

DOCS = [
 "Employees get 24 days of annual leave and may take a vacation any time of the year",
 "Unused vacation days and leave can be carried over to the next year",
 "Public holidays are paid days off and do not count as annual leave",
 "Sick leave needs a doctor note after three days of absence",
 "Flights must be booked fourteen days in advance for business travel",
 "Hotel stays on a business trip are capped at six thousand rupees per night",
 "Travel expenses need receipts and a trip report within thirty days",
 "Taxi and train tickets are reimbursed for a business trip",
 "Use a password manager and enable two factor authentication on every account",
 "Report a lost laptop or phone to the security team within twenty four hours",
 "Never share a password and always lock your laptop screen",
 "Phishing emails should be reported to the security team immediately",
]
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize
from sklearn.cluster import KMeans

X = TfidfVectorizer(stop_words="english").fit_transform(DOCS)
E = normalize(TruncatedSVD(n_components=4, random_state=0).fit_transform(X))
labels = KMeans(n_clusters=3, n_init=10, random_state=0).fit_predict(E)
groups = {}
for doc, lab in zip(DOCS, labels):
    groups.setdefault(int(lab), []).append(doc.split()[0] + " " + doc.split()[1] + " " + doc.split()[2])
for lab in sorted(groups):
    print("cluster", lab, "->", groups[lab])

Output:

cluster 0 -> ['Use a password', 'Report a lost', 'Never share a', 'Phishing emails should']
cluster 1 -> ['Employees get 24', 'Unused vacation days', 'Public holidays are', 'Sick leave needs']
cluster 2 -> ['Flights must be', 'Hotel stays on', 'Travel expenses need', 'Taxi and train']

Near-duplicate पहचान, चलाकर

मैंने यह Python virtual environment में numpy 2.5.3, scikit-learn 1.9.1 और faiss-cpu 1.15.1 के साथ चलाया, निश्चित random seeds के साथ ताकि संख्याएँ दोहराई जाएँ। सिर्फ़ जोड़ी a और a_copy (cosine 0.999) 0.98 सीमा से ऊपर है और चिह्नित होती है। 0.707 और 0.742 वाली जोड़ियाँ संबंधित पर अलग हैं।

import numpy as np
V = {"a": [1.0, 0.0], "a_copy": [0.99, 0.05], "b": [0.0, 1.0], "c": [0.7, 0.7]}
names = list(V); M = np.array([V[n] for n in names]); M /= np.linalg.norm(M, axis=1, keepdims=True)
S = M @ M.T
for i in range(len(names)):
    for j in range(i + 1, len(names)):
        flag = "DUPLICATE" if S[i, j] > 0.98 else ""
        print(f"{names[i]:7} {names[j]:7} cosine={S[i, j]:.3f} {flag}")

Output:

a       a_copy  cosine=0.999 DUPLICATE
a       b       cosine=0.000 
a       c       cosine=0.707 
a_copy  b       cosine=0.050 
a_copy  c       cosine=0.742 
b       c       cosine=0.707

त्वरित जाँच: Near-duplicate सीमा कैसे चुननी चाहिए?

  • अपने मॉडल के असली जोड़े देखकर, क्योंकि पैमाने अलग हैं
  • हमेशा 0.5
  • हमेशा ठीक 1.0
  • Random
Answer

अपने मॉडल के असली जोड़े देखकर, क्योंकि पैमाने अलग हैं — Similarity वितरण embedding model और डेटा के अनुसार बदलते हैं।