§ 1.5Module 1

Challenges of NLP

On this page

Challenges of NLP

Recall first

Why can a system score highly on a benchmark yet fail for a new language variety? List two causes that are linguistic and two that are engineering or social.

First principles

NLP is difficult because language is an open-ended communication system, while a model sees finite, imperfect evidence. The main challenges are:

  1. Ambiguity and context: words, structures, references, intentions, and implied facts need context beyond a local sentence.
  2. Variation: spelling, dialect, slang, code-switching, speech disfluency, domain jargon, and genre alter surface form.
  3. Productivity and sparsity: new words and combinations appear after training; rare events make count-based estimates unreliable.
  4. Long-range dependence: agreement, coreference, discourse and topic may be separated by many tokens.
  5. Noisy input and segmentation: OCR, ASR errors, punctuation, Unicode, and script boundaries complicate tokenization.
  6. Multilingual and low-resource conditions: languages differ in morphology, word order, script, and available corpora, tagsets, and treebanks.
  7. World knowledge and grounding: textual patterns do not guarantee factual, physical, or causal understanding.
  8. Evaluation: one metric can hide subgroup errors, annotation disagreement, data leakage, and domain shift; generated text also has many acceptable answers.
  9. Bias, privacy, and safety: training data can encode stereotypes or personal information, and deployment can amplify unequal errors.

Worked diagnosis. A POS tagger trained on edited news sees Google the meeting in a chat message. Failure may be domain shift (new verb use), tokenization (emoji/hashtags), and unknown-word handling—not simply “the model cannot do grammar.” Diagnosis should separate input, data, model, and evaluation causes.

A robust workflow defines the task and population, builds representative train/dev/test splits, measures errors by language and subgroup, and inspects examples. More data is not a universal fix: biased or mismatched data can increase confident failure. Modern pretrained models reduce some data scarcity but do not remove domain shift or language inequality.

Exercise — reveal after answering

A sentiment classifier performs well on movie reviews and poorly on product reviews. Is this necessarily a model-capacity problem?

Answer: No. Domain vocabulary, label conventions, length, and distribution shift may differ. First compare data and error slices, then consider adaptation or a model change.

Exam lens

Organize answers as linguistic, data, computational, and ethical/evaluation challenges. Tie each challenge to a failure mechanism and mitigation; do not just list buzzwords.

Rapid revision checklist

Key takeaways

Sources