§ 2.1Module 2

Tokenization, Stemming, and Lemmatization

On this page

Tokenization, Stemming, and Lemmatization

Recall first

For Studies, studying, and studied. predict which of tokenization, stemming, and lemmatization changes the representation, and which one needs grammatical or dictionary knowledge.

First principles

Tokenization maps a character sequence to units used by a model: words, punctuation, subwords, or characters. It must decide boundaries around punctuation, contractions, numbers, URLs, emoji, and scripts. Tokenization is preprocessing; it does not claim what a token means.

Stemming applies heuristic affix removal to conflate related forms, often producing a non-word stem. For example, a Porter-style stemmer may reduce connected, connecting, and connection toward related strings, but outcomes are algorithm-dependent. It can over-stem unrelated words or under-stem related words.

Lemmatization maps an inflected form to a dictionary base form (lemma), normally using vocabulary and POS/morphological information: studies → study (verb), but studies may be a plural noun in another context. It is more linguistically informed and usually slower or more resource-dependent than stemming.

Worked trace. The studies were useful. → tokens [The, studies, were, useful, .]. A stemmer may output studi, while a lemmatizer using POS outputs study for studies and be for were. The surface token remains useful for reconstruction; normalization is a task choice, not an automatic improvement.

Trade-offs: normalization reduces vocabulary size and sparsity, but can erase distinctions needed for translation, sentiment, or morphology. Modern subword tokenizers address rare words differently and do not make classical concepts irrelevant.

Exercise — reveal after answering

Why is saw a warning against naïve lemmatization?

Answer: It can be the past tense of see or the noun saw; POS and context are needed. A fixed dictionary mapping can be wrong.

Exam lens

Write the distinction explicitly: tokenization = boundaries; stemming = heuristic truncation; lemmatization = lexicon/POS-aware base form. Never call a stemmer a complete morphological analyzer.

Rapid revision checklist

Key takeaways

Sources