NoSQL Solution for Big Data
On this page
3.4 NoSQL Solution for Big Data
Recall first
For each workload, choose a likely family: exact session lookup; device/time-range telemetry; account relationship traversal; flexible product records. Then ask: where is the partition boundary and what happens when it fails?
First principles
Start with the problem type
“Big data problem” is too vague. Classify it by:
- access: point lookup, range scan, aggregation, or graph traversal;
- write/read shape: append-heavy, update-heavy, or mixed;
- latency: milliseconds, seconds, or batch;
- scale: data size, throughput, and growth;
- correctness: transaction and consistency requirements;
- failure geography: node, rack, or region;
- evolution: stable schema or changing fields.
Then choose a model and distribution strategy. NoSQL is a solution only when its trade-offs match those requirements.
Shared-nothing architecture
In a shared-nothing design, each node owns independent CPU, memory, and storage; data is partitioned across nodes, and replicas provide availability. Adding nodes can increase capacity and throughput, but only if partitioning spreads work evenly. A poor key creates a hot partition, and cross-partition queries require network coordination.
A practical design loop is:
- choose the query’s partition key;
- decide what is stored together and how it is replicated;
- design for failure and rebalancing;
- test skew, latency, consistency, and repair;
- add indexes/materialized views only when their write/storage cost is accepted.
Master-slave versus peer-to-peer
Master-slave (primary/replica)
A primary/master accepts or coordinates writes; replicas copy data and may serve reads depending on policy. This can simplify ordering and administration, but the master can become a bottleneck or failover point. MongoDB replica sets use a primary and secondary members with automatic election/failover concepts; replication is not the same as sharding (MongoDB replica sets).
Peer-to-peer / multi-primary
Nodes are more symmetric: any appropriate node can coordinate or accept a request, and data is partitioned/replicated among peers. Cassandra documents a peer-oriented, multi-primary architecture designed for scale-out and availability; it does not provide relational-style distributed joins or general full cross-partition transactions. Logged batches can provide limited atomicity for a bounded set of partitions, but add coordination and performance cost (Cassandra architecture). This improves distribution and avoids one permanent write master, but conflict resolution, consistency levels, repair, and operational reasoning become more complex.
Neither pattern is universally superior. Compare failure mode, write path, consistency, topology, and operational burden.
Choosing systems by problem
| Problem shape | Likely fit | Why / warning |
|---|---|---|
| Exact key lookup, ephemeral state | key-value | Simple and fast; poor rich queries |
| Flexible nested records | document | Embed/reference by access pattern |
| Huge sparse records, key/range access | HBase/Bigtable-style wide-column | Row-key design and region/partition balance matter |
| High-volume partitioned writes, global availability | Cassandra-style wide-column | Query by partition key; no distributed joins |
| Relationship/path queries | graph store | Natural traversal; distributed high-degree paths are costly |
| Audited multi-row transactions and joins | relational system may fit better | NoSQL is not a mandatory solution |
HBase’s official architecture uses a master/RegionServer split, with RegionServers managing table regions and HBase commonly relying on HDFS for durable storage (HBase architecture). Cassandra uses a partitioned wide-column model; MongoDB uses documents and separates sharding from replication. These are architectural distinctions, not brand labels.
Worked design
A global IoT platform appends readings keyed by (device_id, time), queries a device’s last hour, and tolerates a regional outage. A wide-column design can partition by device (possibly with time buckets to prevent unbounded partitions), order by timestamp, replicate across failure domains, and use a peer-to-peer topology for regional writes. It must handle hot devices, late events, retention, repair, and cross-device analytics—perhaps exporting data to a batch engine rather than forcing the operational store to perform global joins.
Exercise — revealed answer
Exercise: A graph traversal repeatedly crosses partitions and takes seconds, while key-based reads are fast. Is the database necessarily “slow”?
Answer: Not necessarily. The workload is mismatched with the partition boundary: local key reads fit, but cross-partition traversal requires coordination. Repartition/model the graph, use a graph-oriented system, or precompute the needed relation; measure before blaming the whole database.
Exam lens
Answer in this order: problem type → data model → partition key → replication/distribution → consistency → failure/rebalancing → trade-off. Compare master-slave and peer-to-peer explicitly and use one product example for each.
Rapid revision checklist
- Define shared-nothing and partition key.
- Explain master-slave versus peer-to-peer.
- Distinguish sharding, replication, and consistency.
- Match key-value/document/wide-column/graph to workloads.
- Explain hot partitions and cross-partition coordination.
- State when relational is still the right solution.
Key takeaways
- NoSQL solves specific big-data access patterns, not “big data” in the abstract.
- Shared-nothing scale depends on balanced partitioning; replication protects availability but adds cost.
- Master-slave simplifies write authority; peer-to-peer distributes authority and complexity.
Sources
- Apache Cassandra architecture.
- Apache HBase architecture.
- MongoDB sharding.
- MongoDB replica-set architectures.
- Neo4j graph database.
- Syllabus-aligned supplement: the system-selection table, failure-domain discussion, and IoT design are a synthesis of the syllabus’s requested distribution models and first-party architecture documentation; exact guarantees vary by configuration and release.