Morphological Models: Dictionaries, FSTs, and Porter Stemming
On this page
Morphological Models: Dictionaries, FSTs, and Porter Stemming
Recall first
How would a dictionary lookup, a finite-state analyzer, and a Porter stemmer treat cats differently? Which one can naturally return grammatical features?
First principles
A dictionary-lookup model stores whole forms and their analyses: cats → cat + N + PL. It is simple and accurate for listed words, but cannot generalize to unseen regular forms and grows with vocabulary.
A finite-state transducer (FST) maps an input string to an output string while reading left to right with finite memory. A morphological analyzer often maps surface form → lexical analysis, e.g. walked → walk+V+PAST; the inverse direction can generate surface forms from analyses. FSTs can encode lexicons, affix rules, and spelling alternations, and their composition lets separate components be combined. They are efficient and interpretable for regular morphology, but difficult or insufficient for phenomena requiring unbounded/context-rich computation, and their coverage depends on the lexicon and rules.
Worked FST trace. Imagine a toy transducer with lexical path cat : cat, followed by a plural transition +PL : s, and an end transition. Input lexical representation cat+PL outputs cats. Running the relation backwards analyzes cats as cat+PL if the path is available. An analysis may also emit tags on an analysis tape; epsilon transitions consume or output nothing. Real analyzers add alternation paths, such as city+PL → cities, so the mapping is a relation, not mere suffix deletion.
The Porter stemmer is a lexicon-free, rule-based suffix-stripping algorithm. It applies ordered rewrite steps based on a measure of vowel/consonant structure and conditions such as a minimum stem measure. It is useful for information retrieval because related spellings may share an index key, but it does not promise a valid lemma or grammatical analysis. Exact output depends on the algorithm/version; do not present it as an FST morphological analyzer even though many suffix rules can be compiled into finite-state machinery.
Model choice
Dictionary = high precision for known forms; FST = analyzable/generative and linguistically structured; Porter = cheap heuristic normalization. Hybrid systems commonly use a lexicon plus productive rules and an unknown-word fallback.
Exercise — reveal after answering
Why can studies defeat a suffix-only model?
Answer: Plural/third-person morphology involves y → ie spelling alternation, and the base may be study; a rule system must encode the alternation or a dictionary/FST path.
Exam lens
Draw the two FST tapes and explain analysis vs generation. State the Porter distinction explicitly: lexicon-free stemmer, not lemma generator or full morphological parser.
Rapid revision checklist
- Compare dictionary, FST, and Porter.
- Define transducer and two directions.
- Trace
cat+PL ↔ cats. - Mention coverage and finite-state limitations.
Key takeaways
- FSTs model regular mappings between surface forms and analyses.
- A dictionary memorizes; rules generalize; a stemmer approximates.
- Morphological output must say whether it is a stem, lemma, or feature analysis.