Introduction to Big Data
On this page
1.1 Introduction to Big Data
Recall first
Without looking below, answer: What makes a dataset a big-data problem: its size alone, or the way it must be stored and processed? Name three characteristics and two data types. Commit to an answer before reading.
First principles
Big data is data whose scale, speed, diversity, or variability makes ordinary single-machine storage and analysis inadequate or uneconomic. It is therefore a systems problem, not a magic threshold such as “more than 1 TB.” NIST describes big data as extensive datasets characterized primarily by volume, velocity, variety, and/or variability that require a scalable architecture for efficient storage, manipulation, and analysis (NIST SP 1500-1r2).
The common characteristics are:
- Volume: amount of data and its growth rate. A web-scale click log may exceed one server’s disk or memory.
- Velocity: speed of generation, ingestion, and required processing. A fraud detector may need a decision while an event is still fresh.
- Variety: different structures and formats: tables, JSON, text, images, audio, graphs, and sensor records.
- Veracity/variability: uncertainty in quality and changing distributions, schemas, or meanings. These are useful engineering concerns, but the exact list of “Vs” varies by author; do not treat every expanded V-list as a universal definition.
- Value: the business or scientific decision extracted. Storage without a useful question is not an analytics solution.
Types of big data
- Structured: fixed fields and a defined schema, such as sales rows
(order_id, amount, date). Relational tools fit naturally. - Semi-structured: self-describing records with irregular fields, such as JSON, XML, and event logs. Records can vary while retaining keys or tags.
- Unstructured: content without a regular tabular schema, such as free text, images, video, and audio. Analysis normally begins with feature extraction or indexing.
A single application usually mixes all three: an online store has structured payments, semi-structured click events, and unstructured reviews. “Type” describes representation; “big” describes the scale and processing challenge.
Mechanism: why distribution helps
A distributed data system splits data across machines and runs independent work in parallel. If 100 machines each scan one hundredth of a dataset, elapsed time can approach one hundredth of a single scan—subject to network, skew, coordination, and stragglers. The design must also tolerate machines failing, because a large cluster has many components. Hadoop’s HDFS design explicitly assumes hardware failure is normal and emphasizes high-throughput streaming access over low-latency interactive access (HDFS Design).
The central trade-off is:
Scale and resilience require partitioning and coordination; partitioning creates network, consistency, and operational costs.
Worked classification
Suppose a hospital stores:
- a daily relational billing table: structured;
- JSON readings from bedside devices: semi-structured, with high velocity;
- radiology scans and doctor notes: unstructured;
- ten years of records: high volume and possible veracity issues.
The correct conclusion is not “use NoSQL” automatically. First identify queries, latency, correctness, retention, privacy, and partitioning needs; then choose storage and processing systems.
Exercise — revealed answer
Exercise: A 5 GB CSV file is processed once on a laptop. A 50 MB stream of events arrives every second and must be analysed across a year. Which is more clearly a big-data problem, and why?
Answer: The stream is more clearly big data: its annual volume is large, but more importantly its velocity and continuous ingestion prevent a one-shot laptop workflow. The 5 GB file may still be inconvenient, but size alone does not determine the category.
Exam lens
Write: “Big data is not merely large data; it is data whose volume, velocity, variety, or variability demands scalable architecture.” Then classify structured/semi-structured/unstructured examples. Distinguish a characteristic (what makes the problem hard) from a type (how the data is represented).
Rapid revision checklist
- Define big data as a scalable-systems challenge, not a fixed size.
- Explain volume, velocity, variety, and veracity/variability.
- Distinguish structured, semi-structured, and unstructured data.
- Explain why distributed processing introduces both speed and coordination costs.
- Say why “big” does not automatically imply NoSQL.
Key takeaways
- Big data is defined by the mismatch between data demands and conventional infrastructure.
- Data type and data scale are separate dimensions.
- Distribution enables parallelism and fault tolerance but adds network and coordination trade-offs.
Sources
- NIST Big Data Interoperability Framework, Volume 1.
- Mining of Massive Datasets — authors’ online book.
- Apache HDFS Design.
- Syllabus-aligned supplement: the V-list and the structured/semi-structured/unstructured teaching classification are used here to connect the listed big-data texts to the syllabus; the authoritative NIST definition is the controlling source.