Evaluating N-grams with Perplexity
On this page
Evaluating N-grams with Perplexity
Recall first
If a model assigns higher probability to the correct test sequence, should its perplexity be higher or lower? What must be held constant when comparing two models?
First principles
For a test sequence W = w₁…w_N, the cross-entropy in base 2 is
H(W) = −(1/N) Σᵢ log₂ P(wᵢ | history).
Perplexity is
PP(W) = 2^{H(W)} = P(W)^{−1/N}
when probabilities use base-2 logs. It is the inverse geometric mean probability: lower perplexity means the model was less “surprised” by the test tokens. For an n-gram, the history is truncated and boundary tokens/counting conventions must be specified.
Worked calculation. Suppose a three-token test sequence receives probabilities 0.5, 0.25, 0.5. Then P(W)=0.0625, so PP = 0.0625^{−1/3} ≈ 2.52. If a second model assigns 0.8, 0.2, 0.8, its product is 0.128, so its perplexity is lower: 0.128^{−1/3} ≈ 1.98. The second model is better on this test under the same tokenization and evaluation set.
Perplexity is meaningful for probability models and is a useful intrinsic metric, but not a universal measure of downstream quality. Comparisons require the same test text, token boundaries, vocabulary/OOV treatment, normalization, and probability convention. A word-level model and a subword model can have incomparable perplexities because N differs and units differ. Lower perplexity on one domain need not mean better translation, tagging, or human judgment elsewhere.
Exercise — reveal after answering
A model gives probability 0.01 to a 10-token sequence. Using natural logarithm form, write its perplexity.
Answer: PP = exp(−ln(0.01)/10) = 0.01^{−1/10} ≈ 1.585. The base changes the log formula, not the underlying geometric-mean interpretation.
Exam lens
Show both formulas, state lower is better, and show the geometric-mean derivation. List comparability conditions; this is where many otherwise correct answers lose marks.
Rapid revision checklist
- Define cross-entropy and perplexity.
- Compute from sequence probability.
- Explain lower-is-better.
- State tokenization/OOV comparability requirements.
Key takeaways
- Perplexity is inverse geometric mean probability.
- It evaluates predictive fit to a specified test distribution.
- It is not automatically comparable across tokenizations or tasks.