Maximum Entropy Models and Conditional Random Fields
On this page
Maximum Entropy Models and Conditional Random Fields
Recall first
What does a discriminative model condition on that a generative HMM models jointly? Why are overlapping features useful for POS tagging?
First principles
A maximum entropy (MaxEnt) classifier models the conditional distribution of a label given an observation:
P(y|x) = exp(Σ_k θ_k f_k(x,y)) / Z(x).
f_k can represent arbitrary overlapping evidence: current word, suffix, previous word, capitalization, and neighboring tags. Training chooses weights that fit labeled data under the conditional likelihood while, in practice, regularization controls overfitting. For independent token classification, each token’s label is predicted separately; a sequence extension may condition on previous predicted labels (often called a MEMM-style local model), which introduces exposure and normalization issues.
A Conditional Random Field (CRF) directly models the probability of an entire label sequence given the observation:
P(y|x) = exp(score(x,y)) / Σ_{y'} exp(score(x,y')).
For a linear-chain CRF, the score sums state/word features and adjacent-label transition features. Forward-backward computes the partition function and marginals; Viterbi finds the highest-scoring sequence. CRFs are discriminative and globally normalized over complete sequences, so they can use rich, overlapping input features while modeling label dependencies.
Worked contrast. In to book, a MaxEnt token classifier can use previous=to and word=book to prefer VB. A CRF can additionally score the whole sequence and transition TO → VB, making neighboring decisions consistent. A local label model normalizes choices at each position; a CRF normalizes complete sequences, avoiding the classic local-normalization label-bias problem (though it has more expensive sequence training/inference).
MaxEnt vs CRF
- Output: individual label vs jointly scored label sequence.
- Normalization: local
Z(x)per decision vs global sequence partition function. - Inference: classification vs dynamic programming for sequences.
- Strength: simple flexible features vs global consistency.
- Cost: CRF training/inference and feature engineering; modern neural encoders often replace manual features but retain the distinction between local and structured decoding.
Exercise — reveal after answering
If labels must obey an I-ORG-after-B-ORG constraint, which model naturally scores the transition? Why?
Answer: A linear-chain CRF naturally includes adjacent-label features/transitions and scores the sequence jointly; a standalone MaxEnt classifier needs an added constraint or post-processing.
Exam lens
Use the equations and say discriminative = model P(y|x). The decisive distinction is local independent/locally normalized decisions vs globally normalized sequence modeling, not “MaxEnt is old and CRF is new.”
Rapid revision checklist
- Write the MaxEnt softmax form.
- Define feature functions and regularization.
- Write CRF global normalization.
- Contrast local vs global sequence decisions.
Key takeaways
- Discriminative models focus on predicting labels from observed features.
- CRFs add structured global normalization and label-transition features.
- Flexible features do not remove the need for correct data, constraints, and evaluation.