§ 2.7Module 2

Unknown Words: Open and Closed Vocabulary Tasks

On this page

Unknown Words: Open and Closed Vocabulary Tasks

Recall first

What happens when a test sentence contains a word absent from the training vocabulary? Contrast a closed-vocabulary and an open-vocabulary response.

First principles

A vocabulary is the set of types a model recognizes. In a closed-vocabulary task, the vocabulary is fixed in advance; an unseen test word must be rejected, mapped to an unknown symbol, or treated as an error. In an open-vocabulary task, new words are expected, so the model needs a representation or mechanism for them.

An OOV (out-of-vocabulary) token is unseen relative to a chosen training vocabulary, not necessarily a new word in the language. OOVs arise from names, spelling errors, productive morphology, domain terms, code-mixing, tokenization differences, and finite training data.

A classic n-gram solution replaces rare training types with <UNK> before counting. At test time, unseen words map to <UNK>, giving them shared probability. The replacement must be performed consistently during training, and a threshold can map words with count ≤ k. The model then cannot distinguish two unseen words and may still assign poor probability if the unknown class is badly estimated.

Open-vocabulary alternatives include character n-grams, subword units, morphology-aware representations, copy mechanisms, and character-level models. They improve coverage but can produce bad segmentations or lose whole-word semantics. A vocabulary can also be “open” at the application level while using a finite subword inventory.

Worked trace. Training vocabulary {the, cat, sat, <UNK>} sees test the kitten sat. Map kitten → <UNK> and score P(the|<s>) P(<UNK>|the) P(sat|<UNK>) .... If kitten had appeared in a subword model as kit + ten, no whole-word OOV is necessary. Never compare perplexity across systems without checking tokenization and OOV policy.

Exercise — reveal after answering

Why is mapping only test-time unknown words to <UNK> incorrect?

Answer: The model never learned <UNK> counts during training, so its probability is undefined or uncalibrated. Training must replace rare types consistently before estimating counts.

Exam lens

Define open vs closed vocabulary, OOV, and <UNK>. Explain the training-time replacement rule and compare word-level with subword/character solutions.

Rapid revision checklist

Key takeaways

Sources