§ 3.4Module 3

NoSQL Solution for Big Data

On this page

3.4 NoSQL Solution for Big Data

Recall first

For each workload, choose a likely family: exact session lookup; device/time-range telemetry; account relationship traversal; flexible product records. Then ask: where is the partition boundary and what happens when it fails?

First principles

Start with the problem type

“Big data problem” is too vague. Classify it by:

Then choose a model and distribution strategy. NoSQL is a solution only when its trade-offs match those requirements.

Shared-nothing architecture

In a shared-nothing design, each node owns independent CPU, memory, and storage; data is partitioned across nodes, and replicas provide availability. Adding nodes can increase capacity and throughput, but only if partitioning spreads work evenly. A poor key creates a hot partition, and cross-partition queries require network coordination.

A practical design loop is:

  1. choose the query’s partition key;
  2. decide what is stored together and how it is replicated;
  3. design for failure and rebalancing;
  4. test skew, latency, consistency, and repair;
  5. add indexes/materialized views only when their write/storage cost is accepted.

Master-slave versus peer-to-peer

Master-slave (primary/replica)

A primary/master accepts or coordinates writes; replicas copy data and may serve reads depending on policy. This can simplify ordering and administration, but the master can become a bottleneck or failover point. MongoDB replica sets use a primary and secondary members with automatic election/failover concepts; replication is not the same as sharding (MongoDB replica sets).

Peer-to-peer / multi-primary

Nodes are more symmetric: any appropriate node can coordinate or accept a request, and data is partitioned/replicated among peers. Cassandra documents a peer-oriented, multi-primary architecture designed for scale-out and availability; it does not provide relational-style distributed joins or general full cross-partition transactions. Logged batches can provide limited atomicity for a bounded set of partitions, but add coordination and performance cost (Cassandra architecture). This improves distribution and avoids one permanent write master, but conflict resolution, consistency levels, repair, and operational reasoning become more complex.

Neither pattern is universally superior. Compare failure mode, write path, consistency, topology, and operational burden.

Choosing systems by problem

Problem shapeLikely fitWhy / warning
Exact key lookup, ephemeral statekey-valueSimple and fast; poor rich queries
Flexible nested recordsdocumentEmbed/reference by access pattern
Huge sparse records, key/range accessHBase/Bigtable-style wide-columnRow-key design and region/partition balance matter
High-volume partitioned writes, global availabilityCassandra-style wide-columnQuery by partition key; no distributed joins
Relationship/path queriesgraph storeNatural traversal; distributed high-degree paths are costly
Audited multi-row transactions and joinsrelational system may fit betterNoSQL is not a mandatory solution

HBase’s official architecture uses a master/RegionServer split, with RegionServers managing table regions and HBase commonly relying on HDFS for durable storage (HBase architecture). Cassandra uses a partitioned wide-column model; MongoDB uses documents and separates sharding from replication. These are architectural distinctions, not brand labels.

Worked design

A global IoT platform appends readings keyed by (device_id, time), queries a device’s last hour, and tolerates a regional outage. A wide-column design can partition by device (possibly with time buckets to prevent unbounded partitions), order by timestamp, replicate across failure domains, and use a peer-to-peer topology for regional writes. It must handle hot devices, late events, retention, repair, and cross-device analytics—perhaps exporting data to a batch engine rather than forcing the operational store to perform global joins.

Exercise — revealed answer

Exercise: A graph traversal repeatedly crosses partitions and takes seconds, while key-based reads are fast. Is the database necessarily “slow”?

Answer: Not necessarily. The workload is mismatched with the partition boundary: local key reads fit, but cross-partition traversal requires coordination. Repartition/model the graph, use a graph-oriented system, or precompute the needed relation; measure before blaming the whole database.

Exam lens

Answer in this order: problem type → data model → partition key → replication/distribution → consistency → failure/rebalancing → trade-off. Compare master-slave and peer-to-peer explicitly and use one product example for each.

Rapid revision checklist

Key takeaways

  1. NoSQL solves specific big-data access patterns, not “big data” in the abstract.
  2. Shared-nothing scale depends on balanced partitioning; replication protects availability but adds cost.
  3. Master-slave simplifies write authority; peer-to-peer distributes authority and complexity.

Sources