Traditional vs. Big Data Business Approach
On this page
1.2 Traditional vs. Big Data Business Approach
Recall first
A traditional data warehouse and a big-data platform both answer business questions. What changes: the question, the data, the schema, the scale, or the operating model? Give one advantage and one cost of each approach.
First principles
A traditional business-analytics approach usually starts with known, relatively stable business processes. Data is extracted from operational systems, cleaned and integrated, given a designed schema, and loaded into a warehouse. Analysts run governed reports and dashboards. This is often called schema-on-write: the structure and quality rules are enforced before storage.
A big-data approach is designed for large, diverse, rapidly arriving data and questions that may change. It commonly collects raw or lightly transformed data in a distributed store, then applies schema and computation for a particular use. This is often called schema-on-read. Hadoop’s design illustrates the infrastructure side: large files are spread over DataNodes and batch computation is scheduled near those blocks (HDFS Design; MapReduce Tutorial).
Comparison
| Dimension | Traditional approach | Big-data approach |
|---|---|---|
| Main question | Known KPIs and repeatable reports | Exploration, prediction, and changing questions |
| Data | Mostly structured, curated | Structured plus logs, text, media, sensor data |
| Schema | Designed before loading | Often interpreted per use; governance still required |
| Scaling | Scale up a database server or bounded cluster | Scale out by adding commodity/distributed nodes |
| Processing | SQL, joins, transactions, low-latency reports | Distributed batch/parallel processing; often also streaming systems |
| Strength | Consistency, governance, predictable BI | Breadth, scale, discovery, and new data products |
| Cost/risk | ETL delay and rigid change process | Data quality, privacy, skills, skew, and operations |
This is not a “replace the warehouse” rule. A business may keep a warehouse for audited finance and use a data lake or distributed platform for clickstream exploration. A useful architecture is selected by workload, not fashion.
Business mechanism: from data to decision
A big-data business approach is a loop:
- Instrument: capture events and context, not just final transactions.
- Ingest and retain: store enough raw history to support future questions, with access controls and retention limits.
- Prepare: validate, deduplicate, enrich, and document lineage.
- Analyse: use distributed batch or streaming computation.
- Act: update a recommendation, detect fraud, optimize a route, or inform a human decision.
- Measure: compare outcome and cost; remove data that does not create value.
The technical platform creates no value unless the loop changes a decision. The business trade-off is therefore option value versus governance cost: retaining diverse data may enable an unknown future use, but it increases storage, privacy, quality, and stewardship obligations.
Worked example: retailer
A traditional retailer loads nightly sales into a warehouse and reports revenue by store. A big-data extension collects click events, search terms, reviews, and inventory events. A batch job computes customer/product features; a serving database returns recommendations; the warehouse remains the source for audited revenue.
The design is hybrid because the workloads differ:
- revenue reporting needs stable definitions and reconciled transactions;
- recommendations tolerate a model update every few hours;
- click data is high-volume and semi-structured;
- inventory decisions may need fresher data than a nightly ETL cycle.
Exercise — revealed answer
Exercise: A bank must produce a legally auditable monthly statement and also detect unusual card activity within seconds. Should both be handled by one identical pipeline?
Answer: No. Keep a governed, transactionally reliable path for statements and add a low-latency event-analysis path for detection. They may share raw events and identity controls, but their latency, consistency, and audit requirements differ.
Exam lens
Use the contrast schema-on-write vs. schema-on-read, scale-up vs. scale-out, and known reporting vs. exploratory/predictive workloads. Always include a trade-off and say that big data complements rather than universally replaces relational systems.
Rapid revision checklist
- Define schema-on-write and schema-on-read.
- Compare scale-up and scale-out.
- Explain why a hybrid warehouse plus distributed platform is common.
- Connect data capture to a business decision and measured value.
- Name big-data costs: governance, quality, privacy, skills, and operations.
Key takeaways
- Traditional analytics optimizes governed, repeatable questions; big-data analytics expands scale, variety, and discovery.
- Schema-on-read is flexible, not an excuse for no schema or no governance.
- Architecture follows workload: audit, latency, consistency, and value determine the choice.
Sources
- NIST Big Data Interoperability Framework.
- Apache HDFS Design.
- Apache Hadoop MapReduce Tutorial.
- Mining of Massive Datasets.
- Syllabus-aligned supplement: the business-loop and warehouse/data-lake comparison are an exam-oriented synthesis of the listed big-data texts and the cited architecture sources, not a claim that one named textbook prescribes one business architecture.