§ 1.1Module 1

Introduction to Big Data

On this page

1.1 Introduction to Big Data

Recall first

Without looking below, answer: What makes a dataset a big-data problem: its size alone, or the way it must be stored and processed? Name three characteristics and two data types. Commit to an answer before reading.

First principles

Big data is data whose scale, speed, diversity, or variability makes ordinary single-machine storage and analysis inadequate or uneconomic. It is therefore a systems problem, not a magic threshold such as “more than 1 TB.” NIST describes big data as extensive datasets characterized primarily by volume, velocity, variety, and/or variability that require a scalable architecture for efficient storage, manipulation, and analysis (NIST SP 1500-1r2).

The common characteristics are:

Types of big data

  1. Structured: fixed fields and a defined schema, such as sales rows (order_id, amount, date). Relational tools fit naturally.
  2. Semi-structured: self-describing records with irregular fields, such as JSON, XML, and event logs. Records can vary while retaining keys or tags.
  3. Unstructured: content without a regular tabular schema, such as free text, images, video, and audio. Analysis normally begins with feature extraction or indexing.

A single application usually mixes all three: an online store has structured payments, semi-structured click events, and unstructured reviews. “Type” describes representation; “big” describes the scale and processing challenge.

Mechanism: why distribution helps

A distributed data system splits data across machines and runs independent work in parallel. If 100 machines each scan one hundredth of a dataset, elapsed time can approach one hundredth of a single scan—subject to network, skew, coordination, and stragglers. The design must also tolerate machines failing, because a large cluster has many components. Hadoop’s HDFS design explicitly assumes hardware failure is normal and emphasizes high-throughput streaming access over low-latency interactive access (HDFS Design).

The central trade-off is:

Scale and resilience require partitioning and coordination; partitioning creates network, consistency, and operational costs.

Worked classification

Suppose a hospital stores:

The correct conclusion is not “use NoSQL” automatically. First identify queries, latency, correctness, retention, privacy, and partitioning needs; then choose storage and processing systems.

Exercise — revealed answer

Exercise: A 5 GB CSV file is processed once on a laptop. A 50 MB stream of events arrives every second and must be analysed across a year. Which is more clearly a big-data problem, and why?

Answer: The stream is more clearly big data: its annual volume is large, but more importantly its velocity and continuous ingestion prevent a one-shot laptop workflow. The 5 GB file may still be inconvenient, but size alone does not determine the category.

Exam lens

Write: “Big data is not merely large data; it is data whose volume, velocity, variety, or variability demands scalable architecture.” Then classify structured/semi-structured/unstructured examples. Distinguish a characteristic (what makes the problem hard) from a type (how the data is represented).

Rapid revision checklist

Key takeaways

  1. Big data is defined by the mismatch between data demands and conventional infrastructure.
  2. Data type and data scale are separate dimensions.
  3. Distribution enables parallelism and fault tolerance but adds network and coordination trade-offs.

Sources