Hadoop Limitations
On this page
2.8 Hadoop Limitations
Recall first
Hadoop is strong at large, sequential batch processing. Why is it a poor default for a millisecond key lookup, a transaction updating ten related rows, or a tiny-file-heavy workload?
First principles
Hadoop is an excellent pattern for distributed storage and batch computation, not a universal database. Its limitations follow directly from its design goals: HDFS favors large files and streaming throughput; MapReduce materializes map output and performs a network shuffle; fault tolerance and generality add coordination overhead.
Main limitations and causes
- High latency: starting tasks, reading distributed storage, and shuffling data make classic MapReduce unsuitable for interactive/millisecond queries. HDFS explicitly prioritizes throughput over low latency (HDFS Design).
- Small files: many tiny files create namespace and block metadata for the NameNode, while each file adds task/input overhead. HDFS is tuned for large files, so aggregate small-file workloads waste metadata memory and scheduling effort.
- Batch and iterative cost: each MapReduce stage writes/interchanges intermediate data. Algorithms needing many iterations, such as graph ranking or machine learning, repeatedly pay startup and I/O costs.
- Shuffle bottleneck and skew: all values for a key must meet at a reducer. A hot key or uneven partitions creates a straggler while other workers sit idle.
- Limited transaction/update model: HDFS’s write-once-read-many model is not a replacement for row-level transactions, arbitrary in-place updates, or relational referential integrity.
- Complex operations: cluster installation, tuning, security, upgrades, schema/format management, monitoring, and data governance require specialist skills.
- Not a serving database: MapReduce output is batch-oriented; an online application usually needs a separate low-latency serving store and API.
- Network and storage overhead: replication improves resilience but consumes space and write bandwidth; shuffle moves data across the network.
These are limitations of the classic Hadoop/HDFS/MapReduce workload, not a claim that every current Hadoop ecosystem tool has identical behavior. Other engines and serving systems can use HDFS or YARN while addressing some limitations.
Worked diagnosis
Suppose a job repeatedly multiplies a matrix by a vector for 100 iterations. A straightforward MapReduce implementation broadcasts/reads the vector, shuffles partial products, writes the result, and starts another job 100 times. It is correct and fault-tolerant, but repeated materialization and scheduling can dominate arithmetic. A graph/iterative engine or an in-memory strategy may fit better—if its memory, failure, and cost requirements are acceptable.
Exercise — revealed answer
Exercise: Which limitation is exposed by 10 million 1 KB files: insufficient aggregate disk capacity or a metadata/overhead problem?
Answer: It is chiefly a metadata and operational-overhead problem. The NameNode must track many namespace/file/block objects, and processing them creates many splits/tasks. Combining small files into larger container files can help, though that introduces packaging and update trade-offs.
Exam lens
Do not merely list “slow.” Tie each limitation to a mechanism: batch startup/shuffle → latency; block/namespace metadata → small files; repeated stages → iterative cost; key partitioning → skew; write-once model → updates/transactions; replication → storage/network cost. End by saying Hadoop is workload-specific, not useless.
Rapid revision checklist
- Explain throughput versus latency.
- Explain why small files hurt HDFS.
- Connect MapReduce materialization to iterative cost.
- Explain shuffle skew and hot keys.
- Distinguish HDFS storage from a transactional/serving database.
- Name operational and replication costs.
Key takeaways
- Hadoop’s limitations are consequences of its scalable batch design.
- A workload should be judged by latency, update, transaction, iteration, skew, and file-size needs.
- The right response is often a hybrid architecture, not abandoning distributed storage entirely.
Sources
- Apache HDFS Design — goals and data organization.
- Apache NameNode filesystem limits and metadata.
- Apache Hadoop MapReduce Tutorial.
- Mining of Massive Datasets.
- Syllabus-aligned supplement: the limitation-to-mechanism table is a synthesis of the listed Hadoop/MMDS material and official docs; exact performance depends on workload, cluster, and configuration.