Stages in NLP
On this page
Stages in NLP
Recall first
For Riya emailed Prof. Sen yesterday., list a plausible processing order from raw characters to an interpretation. Which stages can modern systems merge?
First principles
A traditional NLP pipeline decomposes language into levels. The boundaries are pedagogical, not physical:
- Input/script handling: decode Unicode, normalize where safe, identify language/script, and preserve information needed downstream.
- Sentence and word tokenization: find sentence and token boundaries; this is not always whitespace splitting.
- Normalization and preprocessing: case handling, punctuation policy, noise filtering, and optional stop-word treatment. Choices are task-dependent.
- Morphological analysis: split or analyze morphemes and infer features such as lemma, number, tense, or case.
- Lexical/POS analysis: assign lexical categories such as noun or verb.
- Syntactic analysis: build constituents or dependencies and identify relations.
- Semantic analysis: compose or infer meaning, entities, roles, and senses.
- Discourse/pragmatic analysis: connect sentences, resolve references, and use communicative context.
- Task output: translate, retrieve, summarize, answer, classify, or generate.
Critical distinction: preprocessing changes or organizes the input representation (for example, token boundaries or Unicode normalization); morphology analyzes the internal structure and grammatical information of words. Lowercasing “Dogs” is preprocessing. Identifying dog + PL is morphology. A stemmer is a heuristic normalization tool, not automatically a morphological analyzer.
Worked trace. Riya emailed Prof. Sen yesterday. → tokens [Riya, emailed, Prof., Sen, yesterday]; morphology may map emailed to lemma email and past tense; POS may assign proper noun, verb, proper noun abbreviation, proper noun, adverb; syntax builds NP VP NP/AdvP; semantics identifies an emailing event and participants; discourse links the event to yesterday. Each output is uncertain and can feed back to earlier decisions.
Pipelines are useful because they are modular and explainable, but error propagation is real: a bad token boundary can corrupt morphology, tagging, and parsing. Joint or neural systems learn several tasks together and may be more robust, but are harder to inspect and still require preprocessing, evaluation, and task-specific constraints.
Exercise — reveal after answering
Classify each operation: (a) split can't into tokens, (b) map walked to walk + PAST, (c) decide that bank means a financial institution.
Answer: (a) tokenization/preprocessing; (b) morphological analysis; (c) lexical-semantic/contextual interpretation.
Exam lens
- Draw the pipeline and explain both the benefit (modularity) and cost (error propagation).
- Examiners often test preprocessing vs morphology: give the
Dogsexample. - Modern practice may be joint, but the stages remain analysis levels.
Rapid revision checklist
- Recall the pipeline order.
- Distinguish normalization, morphology, syntax, semantics, pragmatics.
- Explain error propagation.
- State why pipelines are still useful.
Key takeaways
- Stages move from form toward structure, meaning, and context.
- Preprocessing prepares representation; morphology explains word structure.
- No stage is perfectly deterministic, and modern models can combine stages.