§ 2.6Module 2

N-gram Sensitivity to the Training Corpus

On this page

N-gram Sensitivity to the Training Corpus

Recall first

Would a trigram trained on newspaper text be reliable for chat messages? Predict two kinds of errors before reading.

First principles

An n-gram model estimates language from observed counts, so its behavior is a direct function of the training corpus. Sensitivity comes from:

Worked comparison. In a news corpus, the minister said may have high trigram probability. In a chat corpus, lol, emoji, abbreviations, and missing punctuation are frequent. The news model may assign low probability to perfectly normal chat sequences—not because they are ungrammatical, but because the training distribution differs.

Sensitivity is both a weakness and a feature: domain adaptation can improve a target task, while a general corpus provides broader coverage. Keep training, development, and test data separate; tuning on the test set makes reported perplexity optimistic. Evaluate by matched and mismatched domains, and inspect OOV and rare n-gram rates. Smoothing reduces zeroes but cannot invent accurate domain knowledge.

Exercise — reveal after answering

Two corpora produce different P(the | in). Is one necessarily wrong?

Answer: No. The estimate reflects corpus domain, tokenization, and sampling. Compare corpus conditions and use a target-matched evaluation before judging.

Exam lens

Mention corpus size, domain, genre, preprocessing, vocabulary, and time. A strong answer explains why counts change and names mitigation: interpolation, adaptation, representative sampling, or evaluation on target data.

Rapid revision checklist

Key takeaways

Sources