§ 3.4Module 3

Maximum Entropy Models and Conditional Random Fields

On this page

Maximum Entropy Models and Conditional Random Fields

Recall first

What does a discriminative model condition on that a generative HMM models jointly? Why are overlapping features useful for POS tagging?

First principles

A maximum entropy (MaxEnt) classifier models the conditional distribution of a label given an observation:

P(y|x) = exp(Σ_k θ_k f_k(x,y)) / Z(x).

f_k can represent arbitrary overlapping evidence: current word, suffix, previous word, capitalization, and neighboring tags. Training chooses weights that fit labeled data under the conditional likelihood while, in practice, regularization controls overfitting. For independent token classification, each token’s label is predicted separately; a sequence extension may condition on previous predicted labels (often called a MEMM-style local model), which introduces exposure and normalization issues.

A Conditional Random Field (CRF) directly models the probability of an entire label sequence given the observation:

P(y|x) = exp(score(x,y)) / Σ_{y'} exp(score(x,y')).

For a linear-chain CRF, the score sums state/word features and adjacent-label transition features. Forward-backward computes the partition function and marginals; Viterbi finds the highest-scoring sequence. CRFs are discriminative and globally normalized over complete sequences, so they can use rich, overlapping input features while modeling label dependencies.

Worked contrast. In to book, a MaxEnt token classifier can use previous=to and word=book to prefer VB. A CRF can additionally score the whole sequence and transition TO → VB, making neighboring decisions consistent. A local label model normalizes choices at each position; a CRF normalizes complete sequences, avoiding the classic local-normalization label-bias problem (though it has more expensive sequence training/inference).

MaxEnt vs CRF

Exercise — reveal after answering

If labels must obey an I-ORG-after-B-ORG constraint, which model naturally scores the transition? Why?

Answer: A linear-chain CRF naturally includes adjacent-label features/transitions and scores the sequence jointly; a standalone MaxEnt classifier needs an added constraint or post-processing.

Exam lens

Use the equations and say discriminative = model P(y|x). The decisive distinction is local independent/locally normalized decisions vs globally normalized sequence modeling, not “MaxEnt is old and CRF is new.”

Rapid revision checklist

Key takeaways

Sources