N-gram Sensitivity to the Training Corpus
On this page
N-gram Sensitivity to the Training Corpus
Recall first
Would a trigram trained on newspaper text be reliable for chat messages? Predict two kinds of errors before reading.
First principles
An n-gram model estimates language from observed counts, so its behavior is a direct function of the training corpus. Sensitivity comes from:
- Size: larger corpora reduce variance and expose more types, but do not guarantee relevant evidence.
- Domain/genre: news, biomedical text, code, and chat have different vocabulary, syntax, punctuation, and phrase frequencies.
- Sampling and balance: a corpus dominated by one author, topic, or language variety gives that distribution too much influence.
- Time: names, events, meanings, and spelling change; old counts can mis-rank current text.
- Preprocessing/tokenization: case folding, punctuation, stemming, and token boundaries change counts and the vocabulary.
- Vocabulary and OOV policy: an unseen word may be a genuine new type, misspelling, or segmentation failure.
Worked comparison. In a news corpus, the minister said may have high trigram probability. In a chat corpus, lol, emoji, abbreviations, and missing punctuation are frequent. The news model may assign low probability to perfectly normal chat sequences—not because they are ungrammatical, but because the training distribution differs.
Sensitivity is both a weakness and a feature: domain adaptation can improve a target task, while a general corpus provides broader coverage. Keep training, development, and test data separate; tuning on the test set makes reported perplexity optimistic. Evaluate by matched and mismatched domains, and inspect OOV and rare n-gram rates. Smoothing reduces zeroes but cannot invent accurate domain knowledge.
Exercise — reveal after answering
Two corpora produce different P(the | in). Is one necessarily wrong?
Answer: No. The estimate reflects corpus domain, tokenization, and sampling. Compare corpus conditions and use a target-matched evaluation before judging.
Exam lens
Mention corpus size, domain, genre, preprocessing, vocabulary, and time. A strong answer explains why counts change and names mitigation: interpolation, adaptation, representative sampling, or evaluation on target data.
Rapid revision checklist
- Explain count dependence.
- Distinguish domain shift from grammaticality.
- State why preprocessing changes the model.
- Protect the test set from tuning.
Key takeaways
- An n-gram model is a model of its corpus, not “English” in the abstract.
- Smoothing addresses sparsity, not domain mismatch.
- Corpus documentation and matched evaluation are essential.