Evaluating Parsers, Parser-based Language Models, and Regional Treebanks
On this page
Evaluating Parsers, Parser-based Language Models, and Regional Treebanks (Self-learning supplement)
Recall first
A parser returns a plausible tree, but how can we compare it with a gold tree? What does a parser-based language model add beyond an ordinary word n-gram?
First principles
For constituency parsing, labeled bracket precision is the fraction of predicted labeled spans that are correct, and labeled bracket recall is the fraction of gold labeled spans recovered:
P = correct predicted brackets / predicted brackets
R = correct predicted brackets / gold brackets, with F1 = 2PR/(P+R).
Punctuation and root conventions must be fixed. Exact match asks whether the whole tree is identical and is stricter. For dependency parses, common metrics are unlabeled attachment score (UAS: correct heads) and labeled attachment score (LAS: correct heads plus relation labels). Evaluation must use the same tokenization, annotation guidelines, and split; otherwise scores are not comparable. Human agreement gives a useful ceiling when annotation is ambiguous.
A parser-based language model assigns probability to structures as well as words, commonly factorizing a sentence through a grammar/parse process. A PCFG can define P(tree) and sum or maximize over trees; a grammar can therefore score structural plausibility, generate strings, or support syntactic surprisal. This is different from a surface n-gram: the parser model uses nonterminals/derivations and may capture hierarchical dependencies, but grammar independence assumptions and ambiguity affect probabilities. The model’s fit still needs held-out evaluation.
Worked metric example. If a predicted tree has 8 labeled brackets, gold has 10, and 6 labels/spans match, then P=6/8=.75, R=6/10=.60, and F1=2(.75)(.60)/1.35≈.667. A tree can have good bracket F1 but a wrong root attachment; inspect error types rather than relying on one number.
Regional-language treebanks (supplement)
A treebank is a corpus annotated with syntactic structures and guidelines. For a regional language, document script/Unicode, tokenization, morphology, word order, case marking, code-mixing, dialect, annotation scheme, and annotator agreement. Do not assume Penn Treebank labels or English parser behavior transfers unchanged. The syllabus’s NPTEL/IIT Bombay/IIIT-H links are starting references; specific corpus/tool coverage must be verified from the official release and evaluated on the target variety.
Exercise — reveal after answering
A parser predicts 9 brackets, gold has 12, and 6 match. Compute precision, recall, and F1.
Answer: P=6/9=.667, R=6/12=.5, F1=2(.667)(.5)/(1.167)≈.571.
Exam lens
Define bracket P/R/F1 and UAS/LAS. Explain exact match. For parser-based LMs, contrast word n-gram local history with grammar/derivation probability. For regional treebanks, emphasize annotation and comparability rather than naming unsupported tools.
Rapid revision checklist
- Compute constituent P/R/F1.
- Distinguish exact match, UAS, LAS.
- Explain grammar-based sentence probability.
- List regional treebank documentation needs.
Key takeaways
- Parser scores depend on representation and annotation conventions.
- One metric cannot explain structural errors.
- A treebank is both data and a specification of linguistic decisions.