NoSQL Data Architecture Patterns
On this page
3.2 NoSQL Data Architecture Patterns
Recall first
Choose a likely model before reading: session lookup by ID; product documents with varying attributes; time-series rows queried by device and time; friends-of-friends traversal.
First principles
The four patterns
1. Key-value stores
Data is addressed as key → opaque value.
- Strength: extremely simple, fast point lookup and horizontal partitioning by key.
- Trade-off: the server may know little about value structure; rich predicates and joins are limited.
- Use: sessions, carts, feature flags, caches, counters.
2. Document stores
A key identifies a JSON/BSON-like document containing nested fields and arrays. MongoDB describes a flexible model in which documents in one collection need not have identical fields or types, and recommends storing data accessed together together (MongoDB data modeling).
- Strength: application-shaped records, nested data, evolving fields, document-level retrieval.
- Trade-off: embedding can duplicate data or make documents large; referencing restores separation but may require multiple reads.
- Use: catalogs, profiles, content, event records.
3. Wide-column / column-family stores
Rows are partitioned by a row/partition key and contain sparse, grouped columns. A column family is a schema/physical grouping chosen for related access; it is not the same as a relational table column. Bigtable-style systems suit large sparse datasets and predictable row-key/range access. HBase stores values under row key → column family:qualifier, with versions/timestamps; its official guide shows put, get, and scan operations (HBase reference guide). Cassandra uses a partition key to locate data and clustering keys to order rows within a partition; performant queries supply the partition key (Cassandra architecture).
- Strength: high-throughput partitioned reads/writes and sparse, wide records.
- Trade-off: query patterns and row-key design are central; arbitrary joins and cross-partition queries are poor fits.
- Use: telemetry, messages by device/time, profiles with many qualifiers.
4. Graph stores
A graph stores nodes/vertices, relationships/edges, and properties. The relationship is a first-class object, making traversals such as “friends connected to this person” natural. Neo4j documents the property-graph model as nodes, relationships, and properties (Neo4j graph database).
- Strength: relationship traversal and path queries.
- Trade-off: partitioning highly connected graphs is difficult; global traversals can be expensive or coordination-heavy.
- Use: recommendations, fraud rings, dependency and social networks.
Variations and distribution choices
Real products combine ideas. A document store may shard documents and replicate each shard. A wide-column store may use log-structured storage and tunable consistency. A graph store may be distributed or single-cluster. “Column-oriented” in analytics is not automatically the same as a NoSQL column-family database.
A useful comparison matrix:
| Pattern | Primary key/access | Natural operation | Main risk |
|---|---|---|---|
| Key-value | exact key | get/put | value opacity |
| Document | document ID + fields | whole document/nested access | joins/duplication |
| Wide-column | partition key + range | partition/range scan | bad key/hot partition |
| Graph | node/edge identity | traversal/path | cross-partition traversal |
Worked classification
cart:U17 → {items...}: key-value or document.- Product with different optional attributes: document.
(device_id, timestamp)telemetry ordered by time: wide-column.- “Accounts sharing devices with a fraudulent account”: graph, or a carefully designed relational/search workflow if traversal is not dominant.
The answer depends on query shape, not the data label alone.
Exercise — revealed answer
Exercise: A team chooses Cassandra but designs queries that omit the partition key and require joins across customers, orders, and products. What architectural mismatch is present?
Answer: The workload is relational/cross-partition while Cassandra is optimized for partition-key-oriented access and does not provide distributed joins or foreign keys. Redesign tables around known queries, or choose a system whose transaction/join model fits.
Exam lens
For each pattern write data model → access pattern → strength → limitation → example. Do not say “column store stores columns like a warehouse”; explain partition keys, column families, and sparse qualifiers.
Rapid revision checklist
- Define key-value, document, wide-column, and graph models.
- Give one use and one trade-off for each.
- Explain embedding versus referencing in documents.
- Explain partition key and clustering/range access.
- Explain why graph relationships are first class.
Key takeaways
- NoSQL architecture starts with the access path and partition boundary.
- Document and wide-column systems are not interchangeable: nested document retrieval differs from partition/range retrieval.
- Graph stores optimize relationships, while their distributed traversal cost must be acknowledged.
Sources
- MongoDB data model.
- Apache Cassandra architecture.
- Apache HBase reference guide.
- Neo4j graph database concepts.
- Google Bigtable paper.
- Syllabus-aligned supplement: the comparison matrix and product examples supplement the listed NoSQL textbook with first-party documentation; product details and guarantees are version-sensitive.