Lesson 2 / 25
How an Assistant Knows About Your Brand
Distinguish knowledge learned in training from information retrieved live from the web.
Two sources of knowledge
A language model learns patterns from a huge training dataset up to a cutoff date; what it says about your brand from memory reflects how often and how consistently you appeared in that data. Many assistants can also search the web or your own documents at question time (retrieval, as in RAG), then write an answer from the pages they fetched, often citing them. So visibility has two parts: being well represented in the data the model learned from, and being easy to find and use when it searches. Changes you make today can show up quickly in retrieval, but only slowly (or never) in a model's built-in memory.
Which lever affects which source
Treat this as a simplified model. Real systems mix both and change often.
Source How it works What you can influence
Training memory learned before a cutoff date long-term presence: widely cited, consistent facts
Live retrieval search + fetch pages at question crawlability, clear pages, fresh content, good titles
User-provided pasted docs / connected tools accurate public docs and help pages people pasteAsk the assistant what it knows
Try asking an assistant with web search off, then on. The difference shows how much your visibility depends on retrieval versus memory, and what it currently gets wrong about you.
Quick check: Which source of knowledge can your recent website changes influence fastest?
- Live retrieval at question time
- Training data already frozen at a past cutoff
- The model's weights directly
- Neither can be influenced
Answer
Live retrieval at question time — Retrieval reads the current web, whereas memory reflects older training data.