Large-Scale File-System Organization
On this page
2.2 Large-Scale File-System Organization
Recall first
For an HDFS file, what is stored by the NameNode and what is stored by the DataNodes? Why split a file into blocks instead of storing the whole file on one machine?
First principles
HDFS presents a hierarchical file namespace but stores each file as a sequence of blocks. Blocks are distributed across DataNodes and replicated for fault tolerance. The NameNode stores namespace metadata and the mapping from files to block locations; DataNodes store block bytes and serve client I/O. User data does not pass through the NameNode (HDFS Design).
Block organization gives three benefits:
- a file can exceed one disk’s capacity;
- different blocks can be read in parallel;
- replicas allow recovery from node/disk failure.
HDFS favors large files and write-once-read-many access. It supports appends/truncates but not arbitrary in-place updates as a normal database would. This simpler coherency model is appropriate for batch workloads.
Metadata and data path
Create/write path
- The client asks the NameNode to create a file.
- The NameNode checks namespace permissions and chooses DataNodes for a block’s replicas.
- The client streams block data to the first DataNode; DataNodes pipeline the data to the other replicas.
- Acknowledgements travel back through the pipeline. The NameNode records the completed block and locations.
Read path
- The client asks the NameNode for block locations.
- The NameNode returns replica locations, not the bytes.
- The client chooses a nearby replica and reads directly from its DataNode.
- Checksums detect corruption; another replica can be tried.
Persisting namespace metadata
The NameNode keeps an in-memory namespace/block map and persists changes in the EditLog. A point-in-time namespace image, FsImage, is periodically combined with edits during a checkpoint. DataNodes send heartbeats and block reports; missing heartbeats signal possible failure, while a block report lists blocks stored by that node.
Worked trace
Let a 250 MB file use 128 MB blocks and replication factor 3. It becomes three logical blocks: 128 MB, 122 MB, and (depending on the exact block allocation) each block has three replicas. The NameNode records “file → ordered blocks → DataNode locations”; DataNodes store the bytes. A reader can fetch the first block from its closest healthy replica and the second from another node concurrently.
This is not three complete copies of the whole file in one location. It is three copies of each block, placed according to policy. Increasing replication improves availability and read opportunities but consumes storage and write bandwidth.
Exercise — revealed answer
Exercise: A DataNode has the bytes for a block, but the NameNode has lost the filename-to-block mapping. Can a normal HDFS client find the file by asking that DataNode alone?
Answer: Not through the normal namespace interface. The NameNode is the metadata authority that maps filenames to blocks. DataNodes report their blocks to the NameNode; they do not independently maintain the HDFS namespace.
Exam lens
Draw two separate paths: metadata path client ↔ NameNode and data path client ↔ DataNode. Label block, replica, heartbeat, block report, FsImage, and EditLog. This diagram earns more marks than a vague “HDFS distributes files.”
Rapid revision checklist
- Define block, replica, namespace, NameNode, and DataNode.
- Trace HDFS write and read paths.
- Explain why metadata is in memory and changes are logged.
- Explain heartbeats and block reports.
- State write-once-read-many and replication trade-offs.
Key takeaways
- HDFS separates metadata management from block-byte transfer.
- Blocks enable capacity, parallel reads, and per-block replication.
- EditLog/FsImage persist namespace state; heartbeats and block reports support recovery.
Sources
- Apache HDFS Design.
- Apache HDFS User Guide.
- Apache HDFS High Availability overview.
- Mining of Massive Datasets.
- Syllabus-aligned supplement: the write/read trace and metadata vocabulary consolidate details from the official HDFS design documentation for exam recall.