Physical Organization of Compute Nodes
On this page
2.1 Physical Organization of Compute Nodes
Recall first
Imagine 20 machines in two racks. Where should a computation run if its input block is on machine 7? What changes if machine 7 fails?
First principles
A distributed cluster is a set of ordinary computers connected by a network. Each machine has CPU, RAM, local disks, network interfaces, an operating system, and Hadoop daemons. Machines are grouped physically into racks; nodes in one rack usually have faster/cheaper communication than nodes across racks because cross-rack traffic crosses switches.
Hadoop’s important physical principle is shared-nothing operation: each worker owns its local resources and does not depend on a single shared disk or shared memory. Data is partitioned across workers; copies provide resilience. The HDFS design explicitly targets commodity hardware, large clusters, high-throughput streaming, and failure recovery (HDFS Design).
Roles
- Master/control services: maintain metadata, allocate resources, schedule work, and monitor health. In HDFS, the NameNode manages the namespace and block locations; in YARN, the ResourceManager schedules applications.
- Worker/data services: DataNodes store blocks and serve reads/writes; NodeManagers run containers/tasks on worker nodes.
- Client: submits jobs and reads/writes data. HDFS data does not flow through the NameNode; the NameNode supplies metadata and the client talks to DataNodes for file data.
Modern Hadoop can run high-availability master services, but the conceptual distinction remains: control plane versus data/compute plane.
Locality and rack awareness
Moving computation is often cheaper than moving a huge block. A scheduler therefore prefers:
- node-local: task and block on the same machine;
- rack-local: task on another machine in the block’s rack;
- off-rack: task elsewhere.
If no local slot is available, waiting forever for locality is worse than running remotely. Scheduling balances locality against queueing delay.
HDFS replica placement also uses rack awareness. For replication factor three, the default policy described by Apache places one replica locally or in the writer’s rack, one in a remote rack, and another on a different node in that remote rack. This reduces the chance that one rack failure destroys all copies while limiting write traffic (HDFS Design — replica placement). Exact placement is policy/configuration, not a promise that every deployment has three replicas.
Worked placement example
There are nodes A and B in rack R1 and C and D in rack R2. A block is written by a client on A with replication factor 3. A typical placement is A, C, D. A map task reading the block is best placed on A; if A is full, B is rack-local; if R1 is unavailable, C or D is remote but still correct because replicas exist.
The trade-off is visible: more rack diversity improves failure tolerance, but a write crossing racks consumes network bandwidth. Locality improves throughput but can cause uneven load if many jobs prefer the same hot nodes.
Exercise — revealed answer
Exercise: Why is “run every task on the same machine as its data” not a valid absolute rule?
Answer: A local machine may be busy or failed, and waiting can cost more than remote execution. Rack-local or remote execution is a correctness-preserving fallback because HDFS exposes replicated blocks; locality is an optimization, not a requirement.
Exam lens
Draw racks → nodes → disks/daemons. Label node-local, rack-local, and off-rack data locality. Then connect rack-aware replication to fault tolerance and network cost.
Rapid revision checklist
- Define node, rack, master/control plane, and worker/data plane.
- Explain shared-nothing partitioning.
- Order node-local, rack-local, and off-rack locality.
- Explain why the NameNode is not on the data path.
- State the availability-versus-write-bandwidth trade-off of rack-aware replicas.
Key takeaways
- Hadoop makes a large cluster useful by combining independent workers with control services.
- Data locality reduces network movement, but scheduling must still handle load and failure.
- Rack awareness protects against correlated failures at the cost of cross-rack traffic.
Sources
- Apache HDFS Design.
- Apache Hadoop Cluster Setup.
- Apache Hadoop Rack Awareness.
- Apache MapReduce Tutorial.
- Syllabus-aligned supplement: the physical rack/node diagram and locality hierarchy are an exam-oriented synthesis of the official HDFS and cluster-setup documentation.