Unknown Words: Open and Closed Vocabulary Tasks
On this page
Unknown Words: Open and Closed Vocabulary Tasks
Recall first
What happens when a test sentence contains a word absent from the training vocabulary? Contrast a closed-vocabulary and an open-vocabulary response.
First principles
A vocabulary is the set of types a model recognizes. In a closed-vocabulary task, the vocabulary is fixed in advance; an unseen test word must be rejected, mapped to an unknown symbol, or treated as an error. In an open-vocabulary task, new words are expected, so the model needs a representation or mechanism for them.
An OOV (out-of-vocabulary) token is unseen relative to a chosen training vocabulary, not necessarily a new word in the language. OOVs arise from names, spelling errors, productive morphology, domain terms, code-mixing, tokenization differences, and finite training data.
A classic n-gram solution replaces rare training types with <UNK> before counting. At test time, unseen words map to <UNK>, giving them shared probability. The replacement must be performed consistently during training, and a threshold can map words with count ≤ k. The model then cannot distinguish two unseen words and may still assign poor probability if the unknown class is badly estimated.
Open-vocabulary alternatives include character n-grams, subword units, morphology-aware representations, copy mechanisms, and character-level models. They improve coverage but can produce bad segmentations or lose whole-word semantics. A vocabulary can also be “open” at the application level while using a finite subword inventory.
Worked trace. Training vocabulary {the, cat, sat, <UNK>} sees test the kitten sat. Map kitten → <UNK> and score P(the|<s>) P(<UNK>|the) P(sat|<UNK>) .... If kitten had appeared in a subword model as kit + ten, no whole-word OOV is necessary. Never compare perplexity across systems without checking tokenization and OOV policy.
Exercise — reveal after answering
Why is mapping only test-time unknown words to <UNK> incorrect?
Answer: The model never learned <UNK> counts during training, so its probability is undefined or uncalibrated. Training must replace rare types consistently before estimating counts.
Exam lens
Define open vs closed vocabulary, OOV, and <UNK>. Explain the training-time replacement rule and compare word-level with subword/character solutions.
Rapid revision checklist
- Define type, token, vocabulary, OOV.
- Contrast open and closed tasks.
- Explain
<UNK>training and test mapping. - State subword benefits and trade-offs.
Key takeaways
- “Unknown” is relative to a model vocabulary.
<UNK>controls sparsity by collapsing rare words, at the cost of identity.- Open-vocabulary systems still use finite internal units; subwords are a design choice.