§ 1.1Module 1

Origin & History of NLP

On this page

Origin & History of NLP

Recall first

Without looking ahead, name three different ways a computer might represent or process the sentence “The bank raised its rate.” What changed when systems moved from hand-written rules to learned models?

First principles

Natural Language Processing (NLP) is the study and engineering of systems that analyze or generate human language. It sits between linguistics (what language is), computer science (how to compute), and statistics/AI (how to infer from uncertain evidence). Language is productive and context-sensitive: a finite vocabulary can form indefinitely many sentences, and the same form can have different meanings.

A useful historical map is a change in the dominant source of knowledge—not a sequence in which older ideas disappeared:

  1. Symbolic/rule-based era (roughly 1950s–1980s): grammars, dictionaries, logic, and hand-written translation rules made assumptions explicit. They were interpretable but expensive to author and brittle outside the covered cases.
  2. Statistical/corpus era (roughly 1990s–2000s): probabilities estimated from corpora handled variation better. HMM taggers, n-gram language models, statistical parsers, and phrase-based translation traded linguistic elegance for measurable robustness, but needed representative annotated or large raw data.
  3. Neural era (roughly 2010s): distributed representations and neural sequence models learned useful features rather than relying entirely on manual feature engineering. They improved generalization but became harder to interpret and data/compute hungry.
  4. Transformer and foundation-model era (late 2010s onward): self-attention and large-scale pretraining allow one model to transfer across tasks. They are powerful, but can hallucinate, inherit corpus bias, and perform unevenly across languages and domains. The current draft of Jurafsky & Martin presents this history as a continuing interaction of linguistic structure, statistical modeling, and engineering.

Worked comparison. A rule system might encode if “bank” is followed by “rate”, choose FINANCE; an n-gram uses local counts such as P(rate | raised its); a neural model can use a wider contextual representation. None automatically “understands” the sentence: each estimates or encodes useful regularities under assumptions.

Exercise — reveal after answering

Give one strength and one failure mode of symbolic and statistical NLP.

Answer: Symbolic systems are explicit and interpretable but brittle and costly to maintain. Statistical systems adapt to observed variation but depend on data quality and can fail on rare, shifted, or biased inputs.

Exam lens

Rapid revision checklist

Key takeaways

Sources