Lesson 26 / 28
Monitoring and Cost
Watch quality, latency, freshness and spending.
Four dashboards
Track quality: recall@k on a rolling golden set or sampled queries against exact search, "no result" rate, and user feedback. Track performance: p50/p95/p99 latency, queries per second, error rate, queue depth, CPU and memory, cache hit rate. Track freshness: ingestion lag and the delay between write and searchability. Track cost: memory per vector, embedding fees (indexing and per query), storage, replicas, and cost per 1,000 queries. Alert on regressions such as a recall drop after an upgrade, latency spikes at specific filters, memory approaching limits, replica lag, and failed ingestion batches. Review the slowest and worst-recall queries regularly; they point to missing indexes, bad filters or data problems. Before each engine or index parameter change, rerun the benchmark and the golden set.
Quick check: Which metric catches a silent loss of search quality after an upgrade?
- Recall@k on a golden set or sampled queries vs exact search
- Server uptime alone
- CPU temperature
- Number of tables
Answer
Recall@k on a golden set or sampled queries vs exact search — Uptime can be perfect while results quietly get worse.