Big data is not simply a lot of data. The guide gives four characteristics, and an answer that names all four is doing better than one that says “a very large amount”.
| Characteristic | What it means | Why it is hard |
|---|---|---|
| Volume | There is far more than a person could read | Storage and processing have to be spread across many machines |
| Variety | Text, images, video, sensor readings, all mixed | It does not fit neatly into rows and columns |
| Velocity | It arrives continuously and fast | Decisions may have to be made before it is all in |
| Veracity | Some of it is wrong, missing or biased | Volume does not fix errors — it hides them |
Veracity is the one students forget: It is also the one worth the most. More data makes a wrong pattern look more convincing, not less.
Free preview
This is the free notes preview
You're reading the free notes. Aimnova Pro unlocks the full study experience — and you can try it with your first topic free to keep:
- FlashcardsLock in vocabulary and key terms with spaced repetition.
- Practice questionsAnswer exam-style questions and get instant AI marking.
- Mock exams & past-paper vaultSit full mocks and see exactly how examiners award marks.
- Personalised study planA daily plan built around your exam date and weak areas.
Analytics means finding something useful in all of that. The guide names prediction, modelling, and understanding human behaviour past, present and future.
What organizations use it for
- Predicting what someone will do next \u2014 buy, click, leave, repay.
- Modelling what would happen under different conditions, without having to try them.
- Understanding behaviour at a scale no survey could reach.
- Personalising \u2014 showing each person something different.
Prediction is about groups: A model says people like you usually do X. Applied to you personally, that can be unfair even when it is accurate about the group \u2014 which is the core of most arguments about scoring systems.
Feeling unprepared for exams?
Get a clear study plan, practice with real questions, and know exactly where you stand before exam day. No more guessing.
One case where scale genuinely solved something that scale was the only route to, and one where it went the other way.
Real-world examples you can name
AlphaFold and the protein structure database — AlphaFold 2 published 2021; over 200 million predicted structures released July 2022
A machine-learning system that predicts a protein's three-dimensional shape from its amino-acid sequence, a problem that had resisted fifty years of work. The predictions were released openly rather than licensed.
Who it affected: Biologists worldwide, including laboratories that could never afford the equipment to determine structures experimentally.
YouTube's recommendation system — deep-learning system described 2016; borderline-content changes from January 2019
A ranking system that predicts what a viewer will watch next, which the company has said drives the majority of watch time. After criticism that it led viewers towards ever more extreme material, the company began demoting 'borderline' content rather than removing it.
Who it affected: Roughly two billion monthly viewers, and creators whose income depends on the ranking.
The pattern to notice: AlphaFold and the protein structure database (UK, with global use, AlphaFold 2 published 2021; over 200 million predicted structures released July 2022) used huge data to answer a question that has a checkable right answer.
A recommender uses huge data to predict a preference, where there is no right answer to check against — so it gets judged by what it does to people instead.
How this is tested — you have to name the characteristics of big data and then judge what analytics can and cannot settle. It comes up two ways:
Paper 1 — structured question
- Part a: identify characteristics of big data
- Part c: evaluate a decision made using predictive analytics
Paper 2 — source-based question
- Q1: read a figure showing the scale of a dataset
- Q4: synthesise sources on an organization's use of analytics
The trap: “more data means better answers”: More data improves a prediction only if the data is right and covers everyone. If it leans one way, more of it makes the wrong answer more confident.
Identify two characteristics of big data.
Model answer plan
See the mark-by-mark plan — for / against / judgement, with marking guidance — in study mode.
To what extent does having more data lead to better decisions?
Model answer plan
See the mark-by-mark plan — for / against / judgement, with marking guidance — in study mode.