Rule-based, Stochastic, and Transformation-based Tagging
On this page
Rule-based, Stochastic, and Transformation-based Tagging
Recall first
How would a rule-based tagger and a stochastic tagger handle an unknown word after the? What does each need before it can decide?
First principles
A rule-based tagger uses a lexicon of possible tags plus hand-written contextual or morphological rules. For example, assign an unknown token after the a noun-like tag, then correct a token ending in -ly toward an adverb when the context permits. It is interpretable and can encode expert knowledge, but rules are laborious, conflict-prone, and brittle under domain change.
A stochastic/probabilistic tagger estimates probabilities from an annotated corpus. A simple bigram-HMM-style decision scores lexical emissions and tag transitions; other stochastic systems use richer features or neural models. It handles ambiguity by ranking sequences and can learn preferences, but depends on representative data and may fail on unseen configurations.
A transformation-based tagger (Brill-style) starts with a baseline assignment, often from lexical frequency, then learns ordered correction rules such as “change NN to VB if the previous tag is TO.” Rules are learned by choosing the transformation that fixes the most training errors, then applied in order. It combines a readable rule list with corpus evidence, but a greedy ordered rule list can miss globally best decisions.
Worked trace. to book: baseline may tag book/NN from its frequent noun use. A learned rule “NN → VB after TO” changes it to book/VB. In the book, the condition is absent, so NN remains. A stochastic sequence model instead compares whole tag sequences using learned transition/emission or feature scores.
Comparison
- Knowledge: rules vs estimated probabilities vs learned transformations.
- Strength: interpretability and precise exceptions vs data-driven coverage vs compact readable corrections.
- Risk: maintenance brittleness vs data sparsity/calibration vs order dependence.
“Stochastic” describes uncertainty/modeling, not one exact algorithm; HMM is one generative example, while MaxEnt/CRF are discriminative families covered next.
Exercise — reveal after answering
Why might a transformation rule learned on news fail in social media?
Answer: Its contextual patterns, vocabulary, punctuation, and baseline distributions may shift. The rule may still fire but no longer represent the target domain.
Exam lens
Compare all three along representation, training, interpretability, and failure mode. Mention that transformation-based tagging is sequential rule correction, not merely “rule-based.”
Rapid revision checklist
- Define lexicon + rules.
- Define stochastic estimation and sequence scoring.
- Describe baseline → learned transformations.
- State one trade-off for each family.
Key takeaways
- Rule-based systems encode expert constraints; stochastic systems estimate uncertainty.
- Transformation-based tagging learns interpretable corrections over a baseline.
- Model family and feature choices matter as much as the label “tagger.”