§ 2.8Module 2

Evaluating N-grams with Perplexity

On this page

Evaluating N-grams with Perplexity

Recall first

If a model assigns higher probability to the correct test sequence, should its perplexity be higher or lower? What must be held constant when comparing two models?

First principles

For a test sequence W = w₁…w_N, the cross-entropy in base 2 is

H(W) = −(1/N) Σᵢ log₂ P(wᵢ | history).

Perplexity is

PP(W) = 2^{H(W)} = P(W)^{−1/N}

when probabilities use base-2 logs. It is the inverse geometric mean probability: lower perplexity means the model was less “surprised” by the test tokens. For an n-gram, the history is truncated and boundary tokens/counting conventions must be specified.

Worked calculation. Suppose a three-token test sequence receives probabilities 0.5, 0.25, 0.5. Then P(W)=0.0625, so PP = 0.0625^{−1/3} ≈ 2.52. If a second model assigns 0.8, 0.2, 0.8, its product is 0.128, so its perplexity is lower: 0.128^{−1/3} ≈ 1.98. The second model is better on this test under the same tokenization and evaluation set.

Perplexity is meaningful for probability models and is a useful intrinsic metric, but not a universal measure of downstream quality. Comparisons require the same test text, token boundaries, vocabulary/OOV treatment, normalization, and probability convention. A word-level model and a subword model can have incomparable perplexities because N differs and units differ. Lower perplexity on one domain need not mean better translation, tagging, or human judgment elsewhere.

Exercise — reveal after answering

A model gives probability 0.01 to a 10-token sequence. Using natural logarithm form, write its perplexity.

Answer: PP = exp(−ln(0.01)/10) = 0.01^{−1/10} ≈ 1.585. The base changes the log formula, not the underlying geometric-mean interpretation.

Exam lens

Show both formulas, state lower is better, and show the geometric-mean derivation. List comparability conditions; this is where many otherwise correct answers lose marks.

Rapid revision checklist

Key takeaways

Sources