§ 1.2Module 1

Traditional vs. Big Data Business Approach

On this page

1.2 Traditional vs. Big Data Business Approach

Recall first

A traditional data warehouse and a big-data platform both answer business questions. What changes: the question, the data, the schema, the scale, or the operating model? Give one advantage and one cost of each approach.

First principles

A traditional business-analytics approach usually starts with known, relatively stable business processes. Data is extracted from operational systems, cleaned and integrated, given a designed schema, and loaded into a warehouse. Analysts run governed reports and dashboards. This is often called schema-on-write: the structure and quality rules are enforced before storage.

A big-data approach is designed for large, diverse, rapidly arriving data and questions that may change. It commonly collects raw or lightly transformed data in a distributed store, then applies schema and computation for a particular use. This is often called schema-on-read. Hadoop’s design illustrates the infrastructure side: large files are spread over DataNodes and batch computation is scheduled near those blocks (HDFS Design; MapReduce Tutorial).

Comparison

DimensionTraditional approachBig-data approach
Main questionKnown KPIs and repeatable reportsExploration, prediction, and changing questions
DataMostly structured, curatedStructured plus logs, text, media, sensor data
SchemaDesigned before loadingOften interpreted per use; governance still required
ScalingScale up a database server or bounded clusterScale out by adding commodity/distributed nodes
ProcessingSQL, joins, transactions, low-latency reportsDistributed batch/parallel processing; often also streaming systems
StrengthConsistency, governance, predictable BIBreadth, scale, discovery, and new data products
Cost/riskETL delay and rigid change processData quality, privacy, skills, skew, and operations

This is not a “replace the warehouse” rule. A business may keep a warehouse for audited finance and use a data lake or distributed platform for clickstream exploration. A useful architecture is selected by workload, not fashion.

Business mechanism: from data to decision

A big-data business approach is a loop:

  1. Instrument: capture events and context, not just final transactions.
  2. Ingest and retain: store enough raw history to support future questions, with access controls and retention limits.
  3. Prepare: validate, deduplicate, enrich, and document lineage.
  4. Analyse: use distributed batch or streaming computation.
  5. Act: update a recommendation, detect fraud, optimize a route, or inform a human decision.
  6. Measure: compare outcome and cost; remove data that does not create value.

The technical platform creates no value unless the loop changes a decision. The business trade-off is therefore option value versus governance cost: retaining diverse data may enable an unknown future use, but it increases storage, privacy, quality, and stewardship obligations.

Worked example: retailer

A traditional retailer loads nightly sales into a warehouse and reports revenue by store. A big-data extension collects click events, search terms, reviews, and inventory events. A batch job computes customer/product features; a serving database returns recommendations; the warehouse remains the source for audited revenue.

The design is hybrid because the workloads differ:

Exercise — revealed answer

Exercise: A bank must produce a legally auditable monthly statement and also detect unusual card activity within seconds. Should both be handled by one identical pipeline?

Answer: No. Keep a governed, transactionally reliable path for statements and add a low-latency event-analysis path for detection. They may share raw events and identity controls, but their latency, consistency, and audit requirements differ.

Exam lens

Use the contrast schema-on-write vs. schema-on-read, scale-up vs. scale-out, and known reporting vs. exploratory/predictive workloads. Always include a trade-off and say that big data complements rather than universally replaces relational systems.

Rapid revision checklist

Key takeaways

  1. Traditional analytics optimizes governed, repeatable questions; big-data analytics expands scale, variety, and discovery.
  2. Schema-on-read is flexible, not an excuse for no schema or no governance.
  3. Architecture follows workload: audit, latency, consistency, and value determine the choice.

Sources