Build a RAG chatbot on your own laptop with Ollama and pgvector
Index your own docs, store embeddings in Postgres, and answer questions with a local model. No API keys, no cloud bill, about an hour of work.
On this page
Most "chat with your docs" tutorials start with an API key and end with a monthly bill. This one builds a RAG chatbot that runs entirely on your machine. By the end you will have two small Python scripts that answer questions about any folder of Markdown files, using only open-source tools: Ollama for the models and Postgres with pgvector for search.
Why run RAG locally?
A language model only knows what it was trained on. Retrieval-augmented generation (RAG) fixes that by finding the most relevant pieces of your documents and handing them to the model together with the question. Running it locally means your documents never leave your laptop, and you can experiment as much as you like without watching a usage meter.
If you can explain the answer from three paragraphs of your own docs, a small local model is usually enough.
You need about 8 GB of free RAM for the chat model used here. On a smaller machine, swap in a lighter model; the code does not change.
What you'll build
- An ingest script that splits Markdown files into chunks and stores an embedding for each chunk.
- A Postgres table with the
vectortype, so similarity search is one SQL query. - An ask script that retrieves context and calls a local model.
An embedding is a list of numbers that represents the meaning of a piece of text. Texts with similar meaning get similar numbers, which is what lets a database find "the paragraphs closest to this question" without any keyword matching.
Set up Ollama and Postgres
- Install Ollama and check that
ollama --versionprints a version. - Pull a chat model and an embedding model.
- Start Postgres with the pgvector extension in Docker, then install the Python packages.
ollama pull llama3.2
pulling manifest
success
ollama pull nomic-embed-text
success
docker run -d --name pgvector -e POSTGRES_PASSWORD=dev -p 5432:5432 pgvector/pgvector:pg16
pip install ollama "psycopg[binary]" pgvector numpyThe copy button on a terminal block copies only the commands, never the output lines, so you can paste straight into your terminal.
If port 5432 is already taken by another Postgres on your machine, map a different host port, for example -p 5433:5432, and change the connection string below to match.
Store embeddings in pgvector
Create the table
The nomic-embed-text model returns 768 numbers per chunk, so the column is vector(768). Connect with any SQL client (or docker exec -it pgvector psql -U postgres) and run:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE docs (
id bigserial PRIMARY KEY,
source text NOT NULL,
body text NOT NULL,
embedding vector(768)
);The source column keeps the file name next to each chunk, so later you can show where an answer came from.
Chunk and embed your docs
Split each file into overlapping chunks of about 800 characters, embed each chunk, and insert it. The overlap means a sentence cut at a chunk boundary still appears whole in one of the two neighbouring chunks.
from pathlib import Path
import numpy as np
import ollama
import psycopg
from pgvector.psycopg import register_vector
conn = psycopg.connect("postgresql://postgres:dev@localhost:5432/postgres", autocommit=True)
register_vector(conn)
def embed(text: str) -> list[float]:
return ollama.embed(model="nomic-embed-text", input=text)["embeddings"][0]
def chunk(text: str, size: int = 800, overlap: int = 120):
for start in range(0, len(text), size - overlap):
yield text[start:start + size]
# one row per chunk
for path in Path("docs").glob("*.md"):
for piece in chunk(path.read_text()):
conn.execute(
"INSERT INTO docs (source, body, embedding) VALUES (%s, %s, %s)",
(path.name, piece, np.array(embed(piece))),
)Put a few Markdown files in a docs folder next to the script and run python ingest.py. Re-running it inserts the chunks again, so empty the table first with TRUNCATE docs; when you re-ingest.
If you switch embedding models later, the vector size changes. Re-create the table and re-ingest everything; one column cannot mix sizes, and vectors from different models are not comparable anyway.
Which embedding model should you pick? For most docs, start with the default and only change it if answers miss obvious matches.
| Model | Dimensions | Good for |
|---|---|---|
nomic-embed-text | 768 | A solid default for long technical docs |
mxbai-embed-large | 1024 | Higher retrieval quality, slower to ingest |
all-minilm | 384 | Tiny and fast, good for quick prototypes |
Ask questions over your docs
Embed the question the same way, fetch the four closest chunks with the <=> cosine distance operator, and pass them to the model as context.
import sys
import numpy as np
import ollama
import psycopg
from pgvector.psycopg import register_vector
conn = psycopg.connect("postgresql://postgres:dev@localhost:5432/postgres")
register_vector(conn)
question = sys.argv[1] if len(sys.argv) > 1 else "How do I rotate the API keys?"
query = np.array(ollama.embed(model="nomic-embed-text", input=question)["embeddings"][0])
rows = conn.execute(
"SELECT source, body FROM docs ORDER BY embedding <=> %s LIMIT 4",
(query,),
).fetchall()
context = "\n\n".join(f"[{source}]\n{body}" for source, body in rows)
reply = ollama.chat(model="llama3.2", messages=[
{"role": "system", "content": f"Answer only from this context. If the answer is not there, say so.\n\n{context}"},
{"role": "user", "content": question},
])
print(reply["message"]["content"])
print("\nSources:", ", ".join(sorted({source for source, _ in rows})))Run it with a question in quotes, for example python ask.py "How do I reset my password?". The answer is printed first, followed by the files the context came from.
"Answer only from this context" does more for accuracy than any model upgrade.
When answers are wrong
Most bad answers come from retrieval, not from the model. Before you try a bigger model, print rows and read what was actually retrieved:
- The right chunk is missing. Try smaller chunks, a larger
LIMIT, or a stronger embedding model from the table above. - The right chunk is there but cut in half. Increase the overlap.
- The model ignores the context. Keep the system prompt short and explicit, and put the context in the system message as shown.
Turning this script into a chatbot with a web UI is the focus of Hour 11 of Learning Development with AI, and serving it to a team comes in Hour 12.
Where to go next
You now have the core of every RAG system: ingest, retrieve, generate. From here, the biggest wins come from better chunking (split on headings instead of character counts), showing sources next to each answer, and putting one gateway in front of all your models so you can swap them freely. The LiteLLM gateway tutorial covers that last step, and the comparison of local runners helps if Ollama is not the right fit for your machine.