Lesson 6 / 27
Roles, System Prompts and Conversation History
Build multi-turn conversations correctly.
You send the whole history every time
Messages alternate between the user and the assistant; a system prompt sets standing instructions (role, rules, format). To continue a conversation, append the assistant's reply and the next user message to your list and send the entire list again. This means cost and latency grow with length, and the list can exceed the context window. Manage it deliberately: keep the system prompt, keep the newest turns, summarise or drop old ones, and store the full log separately for audit. Keep user-supplied text out of the system prompt where possible, and never put secrets in any prompt.
Where the system prompt goes, run
I ran this in a Python virtual environment with the anthropic 1.11.0 and openai 3.22.1 SDKs against a small local stand-in server (shown in the testing topic, saved as mock.py). The server returns canned replies, so no key, network or real model is involved: it proves how the SDK builds requests and handles replies, not what a real model would say. The same instruction is sent two ways. Anthropic receives it in a top-level system field; OpenAI receives it as the first message with role system.
import mock, anthropic, openai
srv = mock.start(); base = f"http://127.0.0.1:{srv.server_address[1]}" # local stand-in server, not a real API
a = anthropic.Anthropic(api_key="k", base_url=base)
a.messages.create(model="demo-model", max_tokens=50, system="You are a terse assistant.",
messages=[{"role": "user", "content": "Hi"}])
print("anthropic body:", mock.STATE["log"][-1]["body"])
o = openai.OpenAI(api_key="k", base_url=base + "/v1")
o.chat.completions.create(model="demo-model", messages=[
{"role": "system", "content": "You are a terse assistant."}, {"role": "user", "content": "Hi"}])
print("openai body: ", mock.STATE["log"][-1]["body"])
Output:
anthropic body: {'max_tokens': 50, 'messages': [{'role': 'user', 'content': 'Hi'}], 'model': 'demo-model', 'system': 'You are a terse assistant.'}
openai body: {'messages': [{'role': 'system', 'content': 'You are a terse assistant.'}, {'role': 'user', 'content': 'Hi'}], 'model': 'demo-model'}Trimming history to a budget, run
I ran this plain-Python (standard library only) example. The function keeps the system message and the newest turns that fit in 130 characters: 9 messages become 4 (the system message plus turns 6, 7 and 8). Real code would count tokens with the provider's tokenizer rather than characters.
def trim(messages, max_chars):
"""Keep the system message, then the newest turns that fit."""
system, rest = messages[0], messages[1:]
kept, used = [], len(system["content"])
for m in reversed(rest):
used += len(m["content"])
if used > max_chars: break
kept.append(m)
return [system] + list(reversed(kept))
chat = [{"role": "system", "content": "Be brief."}]
for i in range(1, 9):
chat.append({"role": "user" if i % 2 else "assistant", "content": f"turn {i} " + "x" * 30})
short = trim(chat, 130)
print(len(chat), "messages ->", len(short))
print([" ".join(m["content"].split()[:2]) for m in short])
Output:
9 messages -> 4 ['Be brief.', 'turn 6', 'turn 7', 'turn 8']
Keep a full log elsewhere
Send the model a trimmed history but store the complete conversation in your database for support and audits.
Quick check: To continue a conversation, what must you send?
- Nothing; the server remembers
- Only the newest message
- Only the API key
- The relevant history plus the new message
Answer
The relevant history plus the new message — The API is stateless; context is whatever you send.