NoSQL Case Study: MongoDB Product Catalog
On this page
3.3 NoSQL Case Study: MongoDB Product Catalog
Recall first
An e-commerce product page needs product details and its five newest reviews. What would you embed, what might you reference, and which access pattern decides?
First principles
Problem and requirements
Consider a catalog with products whose attributes differ by category: a phone has RAM and camera fields; a shoe has size and material. The application needs:
- product-page reads by product ID;
- category/search queries;
- updates to inventory and price;
- optional recent reviews;
- high availability and horizontal growth.
MongoDB stores BSON documents in collections. Its official data-model guidance emphasizes flexible document shape and the principle data accessed together should be stored together (MongoDB data modeling).
Model design
A product document can be shaped as:
{
"_id": "P42",
"category": "phone",
"name": "Aster",
"price": 499,
"specs": {"ram_gb": 8, "camera_mp": 50},
"recent_reviews": [
{"user_id": "U7", "rating": 5, "text": "..."}
]
}
Embed bounded, frequently read data such as a small set of recent reviews when one product page should return it in one document read. Reference or separate unbounded review history because continuously growing arrays make the product document and update pattern worse. This is a design choice, not a universal MongoDB rule.
For a large deployment, distribute data with a shard key aligned to common access. A product-ID lookup benefits from a key that routes that lookup to one shard, while a monotonically increasing key can create a write hotspot. MongoDB separates sharding for horizontal distribution from replica sets for redundancy/failover (MongoDB sharding; replica-set architecture). The exact shard-key choice must be tested against query distribution and cardinality.
Worked request trace
Request: GET /products/P42.
- The router uses the shard key to locate the owning shard.
- The shard reads the product document and returns its embedded bounded fields.
- If older reviews are referenced, the service performs a second targeted read or uses a read model.
- Replica-set members provide redundancy; a primary accepts writes under the configured replication protocol, while failover changes which member serves as primary.
Write: a price update targets P42, and the service must decide whether embedded derived values (such as a cached review average) are updated in the same operation or recomputed asynchronously. Duplication improves reads but creates consistency work.
Why not simply normalize everything?
Normalization reduces duplication and makes shared updates clean, but a product page may need joins/round trips. Embedding reduces read coordination at the cost of duplication, document growth, and update fan-out. MongoDB’s flexible schema accelerates category evolution, but application validation and indexes are still needed to protect correctness and performance.
Exercise — revealed answer
Exercise: Should an unlimited review history be embedded in one product document just because reviews are displayed with the product?
Answer: Not necessarily. A bounded recent-review array can fit the access pattern; unbounded history should normally be separate or otherwise bounded to avoid document growth and write/read inefficiency. The decision follows access frequency, size, update pattern, and consistency needs.
Exam lens
Present the case as requirements → document model → embedding/reference choice → shard-key question → replica-set role → trade-offs. Do not claim that “MongoDB has no schema”; say it has flexible schema with optional validation and application responsibility.
Rapid revision checklist
- Explain why product documents can have category-specific fields.
- Distinguish embedding and referencing.
- Explain sharding versus replica sets.
- State why shard-key choice follows access patterns.
- Name duplication, document growth, validation, and hotspot trade-offs.
Key takeaways
- MongoDB’s document model can match an application-shaped product record.
- Embed bounded data read together; reference data that grows independently or without the parent.
- Sharding distributes load/data; replica sets provide redundancy—different responsibilities.
Sources
- MongoDB data modeling.
- MongoDB embedding and references.
- MongoDB sharding.
- MongoDB replica-set architectures.
- Syllabus-aligned supplement: the e-commerce case and request trace apply MongoDB’s first-party modeling principles to the syllabus’s “NoSQL case study”; production choices require workload measurements and version-specific limits.