Lecture 1 – Introduction to data science
Perspective (PNAS)Blei & Smyth (2017) — Science and data science
Blei, D. M., & Smyth, P. (2017). Science and data science. Proceedings of the National Academy of Sciences, 114(33), 8689–8692. https://doi.org/10.1073/pnas.1702076114 · Open paper
Data science is 'the child of statistics and computer science'. Its core is combining three ways of thinking (statistical, computational and human), worked through step by step together with experts in the field, to answer that field's scientific questions.
Question
What is data science, and why should scientists care about it? The authors look at data science from three perspectives (statistical, computational and human) and argue that combining all three well is what data science is really about.
Why it matters
Scientists in many fields now have huge amounts of data but cannot yet use it fully. Examples are genetic data linked to people's diseases, large archives of digitised texts for social scientists, and sky surveys in astronomy with hundreds of terabytes of images. The existing methods from statistics and computing are not built for these modern problems: very large datasets, data with very many variables, models of the world that are always somewhat wrong ('misspecified'), and finding out what causes what. The authors see this tension as the trigger for the new label 'data science'.
Approach
This is an opinion essay; it contains no data. It builds on Tukey (1962), who described 'data analysis' as much broader than mathematical statistics. The statistical perspective covers uncertainty (every dataset has some), complex data with links over time, space or between variables (handled for example with Bayesian models: models that state assumptions about the world as probabilities and update them with the data), data with thousands of variables per person (handled with regularisation, which keeps a model simple so it does not chase noise, and with machine learning such as deep learning for prediction), and causality (telling cause apart from correlation). The computational perspective covers how methods run as algorithms on a computer: optimisation (finding the best model settings by stepwise 'climbing' towards the best fit, for example the highest likelihood, i.e. how probable the data are under the model), sampling methods (the bootstrap: drawing many resamples from the data to estimate a confidence interval; MCMC (Markov chain Monte Carlo): a sampling method for Bayesian models) and distributed computing (spreading data and work over many computers). The human perspective is shown with a neuroscientist who films mice brains and behaviour and works with a data scientist.
Findings
- Data science is the 'child of statistics and computer science': it takes over their methods and blends, refocuses and develops them for modern scientific data, in the spirit of Tukey's broad 'data analysis'.
- Statistical perspective: all datasets involve uncertainty, and statistics is the foundation for reasoning about it. Three key areas are complex, structured data (e.g. Bayesian models), high-dimensional data, meaning very many variables per data point (regularisation; machine learning such as deep learning for prediction), and causality (correlation is not causation; drawing conclusions from observational data, where nothing was controlled by the researcher).
- Computational perspective: how methods are run as algorithms, and the trade-off between statistical accuracy and computer resources (time and memory). Examples are optimisation, sampling (the bootstrap for confidence intervals, MCMC for Bayesian models) and distributed computing.
- Human perspective: data science cannot be fully automated, because applying the tools needs human judgement and deep knowledge of the field. The data scientist works step by step together with the domain expert (the expert in the field, e.g. a coach or physician; data scientist and expert can be one person wearing two 'hats'). The work cycles through preprocessing, exploration, selection, transformation, analysis, interpretation and communication. Reproducibility (others can repeat the analysis and get the same result) and data provenance (a record of where the data came from and what was done to them) matter.
- Conclusion: data science is more than statistics plus computer science. It means weaving both into a larger framework, problem by problem; understanding the context of the data; taking responsibility for private and public data; and communicating clearly what a dataset can and cannot tell us.
Limitations
- Own inference: it is an opinion piece, so its claims are argued, not tested with data.
- Own inference: the examples come from genetics, social science, astronomy and neuroscience, not sport or health, so applying it to movement science is up to the reader.
- Own inference: it names challenges (causality, models that are always somewhat wrong, scale) but stays general and gives no concrete steps for solving them.
Link to the lectures
Slides 49–51 quote the paper's conclusion and frame the Venn diagram; slides 53–55 compare data science and statistics.
This is the course's definition of data science (Lecture 1 slides 49–51). It matches the Venn diagram of computer science, statistics and domain knowledge; the human perspective is the 'domain knowledge' circle. Its cycle (preprocessing → exploration → analysis → interpretation → communication) mirrors the data science lifecycle used from Lecture 3 onwards. Its statistical perspective (uncertainty, cause vs correlation) links to the data science vs statistics slides (Lecture 1 slides 53–55) and the spurious-correlation example in Lecture 4 (two things that rise together by chance). Machine learning for prediction and the bootstrap come back in Lectures 5–7.
Remember
- Opinion piece (PNAS 2017): data science is 'the child of statistics and computer science'.
- Three perspectives: statistical, computational, human. The core is combining all three.
- Statistical: uncertainty, complex/structured data, very many variables (high dimensionality), cause vs correlation.
- Computational: running methods as algorithms; trade-off between accuracy and computer resources (time, memory); optimisation, sampling (bootstrap, MCMC), distributed computing.
- Human: cannot be fully automated; needs domain knowledge and step-by-step collaboration with domain experts.
- Data science is a cycle: preprocessing → exploration → selection → transformation → analysis → interpretation → communication.
- Why the label arose: abundant data, but existing methods are not built for modern problems.
- Take-home: understand the data's context, take responsibility for private/public data, say clearly what the data can and cannot tell us.
Practice
Q1 · What data science is
According to Blei & Smyth (2017), what is the core of data science?
Show answer
Answer: C. The authors say each perspective is essential, but that combining all three is what data science is about. Option B contradicts their human perspective: data science cannot be fully automated.
Source: Paper p. 8689–8690
Q2 · Computational perspective
Which statement best describes what Blei & Smyth call the computational perspective of data science?
Show answer
Answer: A. Computational thinking is about how methods are implemented and the trade-off between accuracy and computer resources. Option B describes the statistical perspective and option C the human perspective.
Source: Paper p. 8690–8691
Q3 · Human perspective
Why, according to Blei & Smyth (2017), can data science not be fully automated?
Show answer
Answer: D. The human perspective says that understanding the field, choosing data, exploring, picking models and communicating results all need judgement and domain knowledge. Option B is false: the authors stress that models of the world are always somewhat wrong ('misspecified').
Source: Paper p. 8690–8691
Q4 · Why data science emerged
What do Blei & Smyth see as the trigger for the new label 'data science'?
Show answer
Answer: B. The authors describe a tension: scientists have abundant data, but classical methods cannot fully use it, and this tension gave rise to 'data science'. Option A is the opposite of their starting point, which is that data are abundant.
Source: Paper p. 8689–8690
Q5 · Three perspectives applied 3 points
A football club wants to use its players' GPS and injury data. Blei & Smyth (2017) say data science combines a statistical, a computational and a human perspective. State and motivate briefly what each perspective contributes in this project:
1) statistical;
2) computational;
3) human.
Show model answer
1) Statistical: model the data while taking uncertainty into account, and ask whether a link is causal or only a correlation (e.g. does high training load cause injury, or do both just rise together?). 2) Computational: run the analysis efficiently on large GPS data streams, balancing accuracy against computing time and memory (e.g. optimisation, resampling such as the bootstrap, spreading the work over computers). 3) Human: coaches and sport scientists bring domain knowledge to choose relevant data, interpret the results and communicate what the data can and cannot say. The data scientist works with them step by step; data science cannot be fully automated.
How the points are earned
- 1 pt Statistical: handling uncertainty and/or cause vs correlation, applied to the example
- 1 pt Computational: efficient algorithms, trade-off between accuracy and time/memory on large data
- 1 pt Human: domain knowledge of coach/expert to choose data, interpret and communicate; cannot be fully automated
Source: Paper p. 8690–8691; Lecture 1 Slides p. 49–51
How did your answer compare?