Concept of Hadoop, Core Hadoop Components, and Hadoop Ecosystem
On this page
1.4 Hadoop: Components and Ecosystem
Recall first
Match each responsibility before reading: store blocks; manage cluster resources; run map/reduce computation; provide shared libraries. Which component is not itself a database?
First principles
Apache Hadoop is an ecosystem for distributed storage and processing of large datasets across clusters. Its core modules are:
- Hadoop Common: shared libraries and utilities.
- HDFS: distributed file storage using blocks and replication.
- YARN: resource management and application scheduling.
- MapReduce: a batch processing model and execution framework.
The official project overview lists these four core modules (Apache Hadoop Modules). HDFS and MapReduce are often colocated so computation can be scheduled near data (MapReduce Tutorial).
The execution picture
- A client submits data to HDFS. The NameNode records namespace/block metadata; DataNodes store replicated blocks.
- A client submits a job and its configuration to YARN.
- YARN’s ResourceManager allocates containers; a NodeManager supervises work on each worker node; an application master coordinates one application. In MapReduce, the MRAppMaster coordinates map and reduce tasks (YARN Architecture).
- Map tasks read input splits, emit intermediate pairs, and reducers receive grouped partitions.
- Output is committed to the filesystem.
Do not confuse roles: HDFS stores; YARN allocates; MapReduce computes.
Ecosystem: tools around the core
Common ecosystem tools include:
- Hive: SQL-like data warehousing/query layer that compiles queries into execution jobs; it is not the storage layer itself.
- Pig: data-flow scripting, historically used for ETL.
- HBase: distributed wide-column database commonly backed by HDFS, designed for random reads/writes rather than only batch scans.
- Sqoop: historical bulk transfer between relational databases and Hadoop; use depends on distribution/version.
- Flume/Kafka: ingestion/event transport choices; they are not interchangeable with HDFS.
- Oozie: historical workflow scheduling.
- ZooKeeper: coordination service used by distributed systems.
- Spark: a separate distributed processing engine that can use Hadoop storage/resource infrastructure; it is not one of Hadoop’s four core modules.
The ecosystem changes across versions and vendors. In an exam, lead with the four core components, then explain tools by responsibility rather than reciting a memorized list.
Worked architecture trace
For a batch job reading click logs and producing daily counts:
click files -> HDFS blocks -> YARN allocates containers
-> MapReduce maps local blocks
-> shuffle/group by day or URL
-> reducers write HDFS output
Hive could provide the SQL interface; HBase would be a better candidate if the application needs low-latency lookup/update by row key; neither replaces the core distinction between durable storage, resource scheduling, and computation.
Exercise — revealed answer
Exercise: A reducer is waiting for CPU, a file is missing, and a query is expressed in SQL. Which Hadoop-area responsibility is involved in each case?
Answer: YARN/resource scheduling handles CPU allocation; HDFS handles file blocks and replicas; Hive can provide the SQL-like query interface and translate it to execution work. MapReduce is the computation model, not the file namespace or SQL language.
Exam lens
Draw a three-layer diagram: HDFS (storage) → YARN (resources) → MapReduce (processing), with Common underneath/alongside as shared support. Add Hive/HBase as higher-level tools, and explicitly state that the ecosystem is broader than the core.
Rapid revision checklist
- Define Hadoop as distributed storage/processing ecosystem.
- Name the four core modules and their jobs.
- Distinguish NameNode, DataNode, ResourceManager, NodeManager, and MRAppMaster.
- Explain data locality.
- Classify Hive, HBase, Sqoop, ZooKeeper, and Spark by role.
Key takeaways
- The exam-safe separation is HDFS stores, YARN schedules, MapReduce computes.
- Hadoop’s value comes from combining partitioning, locality, parallelism, and fault recovery.
- Ecosystem tools solve ingestion, querying, serving, coordination, or workflow problems around the core.
Sources
- Apache Hadoop Modules.
- Apache HDFS Design.
- Apache Hadoop MapReduce Tutorial.
- Apache YARN Architecture.
- Syllabus-aligned supplement: the ecosystem-tool descriptions extend the four core modules in the syllabus using official Hadoop documentation; tool availability and preferred engines vary by Hadoop distribution and release.