# Index Size, Build Time and IVFFlat — Vector Databases

Source: https://www.geekswithgeeks.com/en/vector-databases/i-size

> Account for the memory and disk an index needs, and know the alternative.

## The index can be bigger than the vectors

An HNSW index stores graph links in addition to the vectors, so it is often **larger than the raw data** and works best when it fits in memory (PostgreSQL's shared buffers and the OS cache). Building it on millions of rows can take a long time and a lot of memory, and heavy inserts and deletes add maintenance cost. **IVFFlat** is the other index type in pgvector: it clusters vectors into lists at build time (so it needs data present before you build it, and a sensible number of `lists`), searches the nearest `probes` lists, builds faster and uses less memory, but usually gives lower recall at the same speed than HNSW and may need re-building after the data distribution changes. Rule of thumb: start with **HNSW** unless build time or memory forces **IVFFlat**.

## Table and index sizes, run

I ran this SQL on PostgreSQL 16 with the pgvector extension, version 0.8.6, in a Docker container. For 20,000 vectors of 32 dimensions, the raw vector data is about 2.6 MB, the table about 4 MB, and the HNSW index about 9 MB, so the index alone is several times the raw vectors. At a million vectors of 768 dimensions the same effect is gigabytes.

```sql
SELECT 'rows' AS what, count(*)::text AS value FROM big
UNION ALL SELECT 'table size (MB)', round(pg_table_size('big') / 1e6)::text
UNION ALL SELECT 'hnsw index size (MB)', round(pg_relation_size('big_hnsw') / 1e6)::text
UNION ALL SELECT 'raw vectors: 20000 x 32 x 4 bytes (MB)', round(20000 * 32 * 4 / 1e6, 1)::text;
```

Output:

```
                  what                  | value 
----------------------------------------+-------
 rows                                   | 20000
 table size (MB)                        | 4
 hnsw index size (MB)                   | 9
 raw vectors: 20000 x 32 x 4 bytes (MB) | 2.6
(4 rows)
```

## Build after bulk loading

Creating the index after loading the data is usually much faster than inserting into an existing index.

**Quiz:** Which index type usually gives better recall at the same speed in pgvector?

- [ ] IVFFlat
- [x] HNSW
- [ ] Neither; they are identical
- [ ] A B-tree

*Answer:* HNSW. IVFFlat trades some recall for faster builds and smaller memory.
