§ 1.3Module 1

Stages in NLP

On this page

Stages in NLP

Recall first

For Riya emailed Prof. Sen yesterday., list a plausible processing order from raw characters to an interpretation. Which stages can modern systems merge?

First principles

A traditional NLP pipeline decomposes language into levels. The boundaries are pedagogical, not physical:

  1. Input/script handling: decode Unicode, normalize where safe, identify language/script, and preserve information needed downstream.
  2. Sentence and word tokenization: find sentence and token boundaries; this is not always whitespace splitting.
  3. Normalization and preprocessing: case handling, punctuation policy, noise filtering, and optional stop-word treatment. Choices are task-dependent.
  4. Morphological analysis: split or analyze morphemes and infer features such as lemma, number, tense, or case.
  5. Lexical/POS analysis: assign lexical categories such as noun or verb.
  6. Syntactic analysis: build constituents or dependencies and identify relations.
  7. Semantic analysis: compose or infer meaning, entities, roles, and senses.
  8. Discourse/pragmatic analysis: connect sentences, resolve references, and use communicative context.
  9. Task output: translate, retrieve, summarize, answer, classify, or generate.

Critical distinction: preprocessing changes or organizes the input representation (for example, token boundaries or Unicode normalization); morphology analyzes the internal structure and grammatical information of words. Lowercasing “Dogs” is preprocessing. Identifying dog + PL is morphology. A stemmer is a heuristic normalization tool, not automatically a morphological analyzer.

Worked trace. Riya emailed Prof. Sen yesterday. → tokens [Riya, emailed, Prof., Sen, yesterday]; morphology may map emailed to lemma email and past tense; POS may assign proper noun, verb, proper noun abbreviation, proper noun, adverb; syntax builds NP VP NP/AdvP; semantics identifies an emailing event and participants; discourse links the event to yesterday. Each output is uncertain and can feed back to earlier decisions.

Pipelines are useful because they are modular and explainable, but error propagation is real: a bad token boundary can corrupt morphology, tagging, and parsing. Joint or neural systems learn several tasks together and may be more robust, but are harder to inspect and still require preprocessing, evaluation, and task-specific constraints.

Exercise — reveal after answering

Classify each operation: (a) split can't into tokens, (b) map walked to walk + PAST, (c) decide that bank means a financial institution.

Answer: (a) tokenization/preprocessing; (b) morphological analysis; (c) lexical-semantic/contextual interpretation.

Exam lens

Rapid revision checklist

Key takeaways

Sources