§ 3.2Module 3

NoSQL Data Architecture Patterns

On this page

3.2 NoSQL Data Architecture Patterns

Recall first

Choose a likely model before reading: session lookup by ID; product documents with varying attributes; time-series rows queried by device and time; friends-of-friends traversal.

First principles

The four patterns

1. Key-value stores

Data is addressed as key → opaque value.

2. Document stores

A key identifies a JSON/BSON-like document containing nested fields and arrays. MongoDB describes a flexible model in which documents in one collection need not have identical fields or types, and recommends storing data accessed together together (MongoDB data modeling).

3. Wide-column / column-family stores

Rows are partitioned by a row/partition key and contain sparse, grouped columns. A column family is a schema/physical grouping chosen for related access; it is not the same as a relational table column. Bigtable-style systems suit large sparse datasets and predictable row-key/range access. HBase stores values under row key → column family:qualifier, with versions/timestamps; its official guide shows put, get, and scan operations (HBase reference guide). Cassandra uses a partition key to locate data and clustering keys to order rows within a partition; performant queries supply the partition key (Cassandra architecture).

4. Graph stores

A graph stores nodes/vertices, relationships/edges, and properties. The relationship is a first-class object, making traversals such as “friends connected to this person” natural. Neo4j documents the property-graph model as nodes, relationships, and properties (Neo4j graph database).

Variations and distribution choices

Real products combine ideas. A document store may shard documents and replicate each shard. A wide-column store may use log-structured storage and tunable consistency. A graph store may be distributed or single-cluster. “Column-oriented” in analytics is not automatically the same as a NoSQL column-family database.

A useful comparison matrix:

PatternPrimary key/accessNatural operationMain risk
Key-valueexact keyget/putvalue opacity
Documentdocument ID + fieldswhole document/nested accessjoins/duplication
Wide-columnpartition key + rangepartition/range scanbad key/hot partition
Graphnode/edge identitytraversal/pathcross-partition traversal

Worked classification

The answer depends on query shape, not the data label alone.

Exercise — revealed answer

Exercise: A team chooses Cassandra but designs queries that omit the partition key and require joins across customers, orders, and products. What architectural mismatch is present?

Answer: The workload is relational/cross-partition while Cassandra is optimized for partition-key-oriented access and does not provide distributed joins or foreign keys. Redesign tables around known queries, or choose a system whose transaction/join model fits.

Exam lens

For each pattern write data model → access pattern → strength → limitation → example. Do not say “column store stores columns like a warehouse”; explain partition keys, column families, and sparse qualifiers.

Rapid revision checklist

Key takeaways

  1. NoSQL architecture starts with the access path and partition boundary.
  2. Document and wide-column systems are not interchangeable: nested document retrieval differs from partition/range retrieval.
  3. Graph stores optimize relationships, while their distributed traversal cost must be acknowledged.

Sources