Lesson 6 / 27
Logits, Softmax and Temperature
Turn raw scores into probabilities and control randomness.
From scores to a distribution
The network outputs one raw score, a logit, for every token in the vocabulary. Softmax exponentiates and normalises them so they are positive and sum to 1. Temperature divides the logits first: below 1 sharpens the distribution (more predictable), above 1 flattens it (more varied). At temperature near 0 the model almost always picks the top token ("greedy"). Subtracting the maximum before exp is a standard trick for numerical stability.
Temperature in action, run
I ran this plain-Python (standard library only) example. At T=0.5 "Paris" takes 95% of the probability; at T=2.0 it takes 56% and the other tokens get a real chance.
import math
def softmax(logits, temperature=1.0):
scaled = [x / temperature for x in logits]
m = max(scaled)
exps = [math.exp(x - m) for x in scaled]
total = sum(exps)
return [e / total for e in exps]
tokens = ["Paris", "London", "banana", "the"]
logits = [4.0, 2.5, -1.0, 1.0]
for t in (0.5, 1.0, 2.0):
probs = softmax(logits, t)
print("T =", t, [f"{p:.3f}" for p in probs])
Output:
T = 0.5 ['0.950', '0.047', '0.000', '0.002'] T = 1.0 ['0.781', '0.174', '0.005', '0.039'] T = 2.0 ['0.563', '0.266', '0.046', '0.126']
Low temperature for facts, higher for ideas
Use low temperature for extraction and factual tasks; raise it for brainstorming. Low temperature reduces variation but does not guarantee correctness.
Quick check: What does raising the temperature do?
- Deletes tokens
- Makes the model larger
- Guarantees correct answers
- Flattens the distribution, making output more varied
Answer
Flattens the distribution, making output more varied — Higher temperature spreads probability across more tokens.