Tokenization, Stemming, and Lemmatization
On this page
Tokenization, Stemming, and Lemmatization
Recall first
For Studies, studying, and studied. predict which of tokenization, stemming, and lemmatization changes the representation, and which one needs grammatical or dictionary knowledge.
First principles
Tokenization maps a character sequence to units used by a model: words, punctuation, subwords, or characters. It must decide boundaries around punctuation, contractions, numbers, URLs, emoji, and scripts. Tokenization is preprocessing; it does not claim what a token means.
Stemming applies heuristic affix removal to conflate related forms, often producing a non-word stem. For example, a Porter-style stemmer may reduce connected, connecting, and connection toward related strings, but outcomes are algorithm-dependent. It can over-stem unrelated words or under-stem related words.
Lemmatization maps an inflected form to a dictionary base form (lemma), normally using vocabulary and POS/morphological information: studies → study (verb), but studies may be a plural noun in another context. It is more linguistically informed and usually slower or more resource-dependent than stemming.
Worked trace. The studies were useful. → tokens [The, studies, were, useful, .]. A stemmer may output studi, while a lemmatizer using POS outputs study for studies and be for were. The surface token remains useful for reconstruction; normalization is a task choice, not an automatic improvement.
Trade-offs: normalization reduces vocabulary size and sparsity, but can erase distinctions needed for translation, sentiment, or morphology. Modern subword tokenizers address rare words differently and do not make classical concepts irrelevant.
Exercise — reveal after answering
Why is saw a warning against naïve lemmatization?
Answer: It can be the past tense of see or the noun saw; POS and context are needed. A fixed dictionary mapping can be wrong.
Exam lens
Write the distinction explicitly: tokenization = boundaries; stemming = heuristic truncation; lemmatization = lexicon/POS-aware base form. Never call a stemmer a complete morphological analyzer.
Rapid revision checklist
- Define token, stem, and lemma.
- Give one over-stemming example/risk.
- Explain why POS helps lemmatization.
- State vocabulary-reduction benefit and information-loss cost.
Key takeaways
- These operations solve different problems.
- Lemmas are valid linguistic words more often; stems need not be.
- Keep the original text when downstream tasks need exact wording.