Lesson 26 / 28

Monitoring and Cost

Watch quality, latency, freshness and spending.

Four dashboards

Track quality: recall@k on a rolling golden set or sampled queries against exact search, "no result" rate, and user feedback. Track performance: p50/p95/p99 latency, queries per second, error rate, queue depth, CPU and memory, cache hit rate. Track freshness: ingestion lag and the delay between write and searchability. Track cost: memory per vector, embedding fees (indexing and per query), storage, replicas, and cost per 1,000 queries. Alert on regressions such as a recall drop after an upgrade, latency spikes at specific filters, memory approaching limits, replica lag, and failed ingestion batches. Review the slowest and worst-recall queries regularly; they point to missing indexes, bad filters or data problems. Before each engine or index parameter change, rerun the benchmark and the golden set.

Quick check: Which metric catches a silent loss of search quality after an upgrade?

  • Recall@k on a golden set or sampled queries vs exact search
  • Server uptime alone
  • CPU temperature
  • Number of tables
Answer

Recall@k on a golden set or sampled queries vs exact search — Uptime can be perfect while results quietly get worse.