§ 1.4Module 1

Concept of Hadoop, Core Hadoop Components, and Hadoop Ecosystem

On this page

1.4 Hadoop: Components and Ecosystem

Recall first

Match each responsibility before reading: store blocks; manage cluster resources; run map/reduce computation; provide shared libraries. Which component is not itself a database?

First principles

Apache Hadoop is an ecosystem for distributed storage and processing of large datasets across clusters. Its core modules are:

The official project overview lists these four core modules (Apache Hadoop Modules). HDFS and MapReduce are often colocated so computation can be scheduled near data (MapReduce Tutorial).

The execution picture

  1. A client submits data to HDFS. The NameNode records namespace/block metadata; DataNodes store replicated blocks.
  2. A client submits a job and its configuration to YARN.
  3. YARN’s ResourceManager allocates containers; a NodeManager supervises work on each worker node; an application master coordinates one application. In MapReduce, the MRAppMaster coordinates map and reduce tasks (YARN Architecture).
  4. Map tasks read input splits, emit intermediate pairs, and reducers receive grouped partitions.
  5. Output is committed to the filesystem.

Do not confuse roles: HDFS stores; YARN allocates; MapReduce computes.

Ecosystem: tools around the core

Common ecosystem tools include:

The ecosystem changes across versions and vendors. In an exam, lead with the four core components, then explain tools by responsibility rather than reciting a memorized list.

Worked architecture trace

For a batch job reading click logs and producing daily counts:

click files -> HDFS blocks -> YARN allocates containers
             -> MapReduce maps local blocks
             -> shuffle/group by day or URL
             -> reducers write HDFS output

Hive could provide the SQL interface; HBase would be a better candidate if the application needs low-latency lookup/update by row key; neither replaces the core distinction between durable storage, resource scheduling, and computation.

Exercise — revealed answer

Exercise: A reducer is waiting for CPU, a file is missing, and a query is expressed in SQL. Which Hadoop-area responsibility is involved in each case?

Answer: YARN/resource scheduling handles CPU allocation; HDFS handles file blocks and replicas; Hive can provide the SQL-like query interface and translate it to execution work. MapReduce is the computation model, not the file namespace or SQL language.

Exam lens

Draw a three-layer diagram: HDFS (storage) → YARN (resources) → MapReduce (processing), with Common underneath/alongside as shared support. Add Hive/HBase as higher-level tools, and explicitly state that the ecosystem is broader than the core.

Rapid revision checklist

Key takeaways

  1. The exam-safe separation is HDFS stores, YARN schedules, MapReduce computes.
  2. Hadoop’s value comes from combining partitioning, locality, parallelism, and fault recovery.
  3. Ecosystem tools solve ingestion, querying, serving, coordination, or workflow problems around the core.

Sources