Origin & History of NLP
On this page
Origin & History of NLP
Recall first
Without looking ahead, name three different ways a computer might represent or process the sentence “The bank raised its rate.” What changed when systems moved from hand-written rules to learned models?
First principles
Natural Language Processing (NLP) is the study and engineering of systems that analyze or generate human language. It sits between linguistics (what language is), computer science (how to compute), and statistics/AI (how to infer from uncertain evidence). Language is productive and context-sensitive: a finite vocabulary can form indefinitely many sentences, and the same form can have different meanings.
A useful historical map is a change in the dominant source of knowledge—not a sequence in which older ideas disappeared:
- Symbolic/rule-based era (roughly 1950s–1980s): grammars, dictionaries, logic, and hand-written translation rules made assumptions explicit. They were interpretable but expensive to author and brittle outside the covered cases.
- Statistical/corpus era (roughly 1990s–2000s): probabilities estimated from corpora handled variation better. HMM taggers, n-gram language models, statistical parsers, and phrase-based translation traded linguistic elegance for measurable robustness, but needed representative annotated or large raw data.
- Neural era (roughly 2010s): distributed representations and neural sequence models learned useful features rather than relying entirely on manual feature engineering. They improved generalization but became harder to interpret and data/compute hungry.
- Transformer and foundation-model era (late 2010s onward): self-attention and large-scale pretraining allow one model to transfer across tasks. They are powerful, but can hallucinate, inherit corpus bias, and perform unevenly across languages and domains. The current draft of Jurafsky & Martin presents this history as a continuing interaction of linguistic structure, statistical modeling, and engineering.
Worked comparison. A rule system might encode if “bank” is followed by “rate”, choose FINANCE; an n-gram uses local counts such as P(rate | raised its); a neural model can use a wider contextual representation. None automatically “understands” the sentence: each estimates or encodes useful regularities under assumptions.
Exercise — reveal after answering
Give one strength and one failure mode of symbolic and statistical NLP.
Answer: Symbolic systems are explicit and interpretable but brittle and costly to maintain. Statistical systems adapt to observed variation but depend on data quality and can fail on rare, shifted, or biased inputs.
Exam lens
- History is best written as rules → statistics → neural pretraining, with overlap and trade-offs, not as “rules became useless.”
- Connect each era to its knowledge source, typical models, strength, and limitation.
- The syllabus is older than current practice: mention transformers/foundation models as modern context, but answer named syllabus models directly.
Rapid revision checklist
- Define NLP and its interdisciplinary character.
- Contrast symbolic, statistical, neural, and transformer eras.
- State one trade-off for each.
- Explain why historical methods still coexist.
Key takeaways
- NLP methods change mainly in how linguistic knowledge is represented and learned.
- More data-driven models reduce manual rules, not the need for assumptions or evaluation.
- Modern systems should be judged by robustness, fairness, interpretability, and language coverage—not accuracy alone.