Lesson 13 / 29

Prompting, RAG or Fine-Tuning: A Decision Process

Choose the cheapest approach that meets the target, and know what fine-tuning costs.

Escalate only when the evaluation says so

A practical ladder: (1) better prompts and examples, since they are cheapest and fastest to iterate; (2) retrieval (RAG) when the model lacks facts that change or are private; (3) a stronger or different model; (4) fine-tuning when you need consistent style, format or a narrow skill, lower latency or cost through a smaller specialised model, or behaviour that long prompts cannot achieve. Fine-tuning is not a way to add fresh knowledge cheaply; it needs quality training data (hundreds to thousands of reviewed examples in the right format), an evaluation set, a training run, and ongoing maintenance when the base model changes. Parameter-efficient methods such as LoRA train small adapter matrices instead of all weights, cutting cost sharply: for one 4096 by 4096 matrix, a rank-8 adapter has 65,536 trainable parameters against 16.8 million (0.39%). Always compare the fine-tuned model with the best prompted baseline on the same evaluation set, because the baseline is often good enough.

How small is a LoRA adapter, run

I ran this with plain Python 3 (standard library only). For a 4096 x 4096 matrix, adapters of rank 4, 8, 16 and 64 have 0.20%, 0.39%, 0.78% and 3.12% of the full matrix's parameters. Adapting 4 matrices in each of 32 layers at rank 8 trains about 8.4 million parameters. The layer and matrix counts are an example, not a specific model.

def full_params(d_in, d_out): return d_in * d_out
def lora_params(d_in, d_out, r): return r * (d_in + d_out)           # two thin matrices: (d_in x r) and (r x d_out)

d = 4096
for r in (4, 8, 16, 64):
    full, lora = full_params(d, d), lora_params(d, d, r)
    print(f"rank {r:2d}: {lora:>9,} trainable vs {full:,} in the full matrix -> {lora / full:.2%}")

layers, matrices_per_layer = 32, 4                                  # e.g. adapt 4 attention matrices in each of 32 layers
total = layers * matrices_per_layer * lora_params(d, d, 8)
print(f"adapting {matrices_per_layer} matrices in {layers} layers at rank 8: {total:,} trainable parameters ({total / 1e6:.1f} M)")

Output:

rank  4:    32,768 trainable vs 16,777,216 in the full matrix -> 0.20%
rank  8:    65,536 trainable vs 16,777,216 in the full matrix -> 0.39%
rank 16:   131,072 trainable vs 16,777,216 in the full matrix -> 0.78%
rank 64:   524,288 trainable vs 16,777,216 in the full matrix -> 3.12%
adapting 4 matrices in 32 layers at rank 8: 8,388,608 trainable parameters (8.4 M)

Checking fine-tuning data before training, run

I ran this with plain Python 3 (standard library only). A validator checks each JSONL line in a chat-style format: valid JSON, a messages list with at least two turns, known roles, a final assistant answer that is not empty. Only the first example is accepted; the others fail for different reasons. Real providers define their own exact formats, so follow their documentation.

import json

def check_example(line):
    try: ex = json.loads(line)
    except json.JSONDecodeError as e: return f"bad JSON: {e.msg}"
    msgs = ex.get("messages")
    if not isinstance(msgs, list) or len(msgs) < 2: return "needs a messages list with at least 2 turns"
    if any(m.get("role") not in {"system", "user", "assistant"} for m in msgs): return "unknown role"
    if msgs[-1]["role"] != "assistant": return "last turn must be the assistant answer to learn"
    if not msgs[-1].get("content", "").strip(): return "empty assistant answer"
    return "ok"

lines = [
    '{"messages": [{"role": "user", "content": "Refund status?"}, {"role": "assistant", "content": "It is processed within 5 days."}]}',
    '{"messages": [{"role": "user", "content": "Hi"}]}',
    '{"messages": [{"role": "user", "content": "Hi"}, {"role": "assistant", "content": ""}]}',
    '{"messages": [{"role": "robot", "content": "Hi"}, {"role": "assistant", "content": "Hello"}]}',
    '{messages: nope}',
]
for i, l in enumerate(lines, 1): print(i, check_example(l))

Output:

1 ok
2 needs a messages list with at least 2 turns
3 empty assistant answer
4 unknown role
5 bad JSON: Expecting property name enclosed in double quotes

Compare with the best prompted baseline

A fine-tuned model must beat a well-prompted strong model to justify its cost.

Quick check: Fine-tuning is best suited to which need?

  • Adding facts that change daily
  • Consistent style, format or a narrow skill
  • Avoiding any evaluation
  • Replacing access control
Answer

Consistent style, format or a narrow skill — Use retrieval for changing facts and fine-tuning for stable behaviour.