Home: all lectures
0 XP0 day streak

Data Science in Sport and Health · Papers

Course literature

12 papersquestion · why · approach · findings

What the exam asks about papers. The lecturers' study advice: a smaller number of questions come from the course literature. You do not need every detail — know each paper's question, why it matters, general approach and main finding. Example from the advice sheet: "Rommers et al. used preseason measurements… to predict whether a player would sustain an injury. What type of machine learning problem is this?" → supervised classification.

Lecture 1 – Introduction to data science

Perspective (PNAS)

Blei & Smyth (2017) — Science and data science

Blei, D. M., & Smyth, P. (2017). Science and data science. Proceedings of the National Academy of Sciences, 114(33), 8689–8692. https://doi.org/10.1073/pnas.1702076114 · Open paper

Data science is 'the child of statistics and computer science'. Its core is combining three ways of thinking (statistical, computational and human), worked through step by step together with experts in the field, to answer that field's scientific questions.

Question

What is data science, and why should scientists care about it? The authors look at data science from three perspectives (statistical, computational and human) and argue that combining all three well is what data science is really about.

Why it matters

Scientists in many fields now have huge amounts of data but cannot yet use it fully. Examples are genetic data linked to people's diseases, large archives of digitised texts for social scientists, and sky surveys in astronomy with hundreds of terabytes of images. The existing methods from statistics and computing are not built for these modern problems: very large datasets, data with very many variables, models of the world that are always somewhat wrong ('misspecified'), and finding out what causes what. The authors see this tension as the trigger for the new label 'data science'.

Approach

This is an opinion essay; it contains no data. It builds on Tukey (1962), who described 'data analysis' as much broader than mathematical statistics. The statistical perspective covers uncertainty (every dataset has some), complex data with links over time, space or between variables (handled for example with Bayesian models: models that state assumptions about the world as probabilities and update them with the data), data with thousands of variables per person (handled with regularisation, which keeps a model simple so it does not chase noise, and with machine learning such as deep learning for prediction), and causality (telling cause apart from correlation). The computational perspective covers how methods run as algorithms on a computer: optimisation (finding the best model settings by stepwise 'climbing' towards the best fit, for example the highest likelihood, i.e. how probable the data are under the model), sampling methods (the bootstrap: drawing many resamples from the data to estimate a confidence interval; MCMC (Markov chain Monte Carlo): a sampling method for Bayesian models) and distributed computing (spreading data and work over many computers). The human perspective is shown with a neuroscientist who films mice brains and behaviour and works with a data scientist.

Findings

  • Data science is the 'child of statistics and computer science': it takes over their methods and blends, refocuses and develops them for modern scientific data, in the spirit of Tukey's broad 'data analysis'.
  • Statistical perspective: all datasets involve uncertainty, and statistics is the foundation for reasoning about it. Three key areas are complex, structured data (e.g. Bayesian models), high-dimensional data, meaning very many variables per data point (regularisation; machine learning such as deep learning for prediction), and causality (correlation is not causation; drawing conclusions from observational data, where nothing was controlled by the researcher).
  • Computational perspective: how methods are run as algorithms, and the trade-off between statistical accuracy and computer resources (time and memory). Examples are optimisation, sampling (the bootstrap for confidence intervals, MCMC for Bayesian models) and distributed computing.
  • Human perspective: data science cannot be fully automated, because applying the tools needs human judgement and deep knowledge of the field. The data scientist works step by step together with the domain expert (the expert in the field, e.g. a coach or physician; data scientist and expert can be one person wearing two 'hats'). The work cycles through preprocessing, exploration, selection, transformation, analysis, interpretation and communication. Reproducibility (others can repeat the analysis and get the same result) and data provenance (a record of where the data came from and what was done to them) matter.
  • Conclusion: data science is more than statistics plus computer science. It means weaving both into a larger framework, problem by problem; understanding the context of the data; taking responsibility for private and public data; and communicating clearly what a dataset can and cannot tell us.

Limitations

  • Own inference: it is an opinion piece, so its claims are argued, not tested with data.
  • Own inference: the examples come from genetics, social science, astronomy and neuroscience, not sport or health, so applying it to movement science is up to the reader.
  • Own inference: it names challenges (causality, models that are always somewhat wrong, scale) but stays general and gives no concrete steps for solving them.

Link to the lectures

Slides 49–51 quote the paper's conclusion and frame the Venn diagram; slides 53–55 compare data science and statistics.

This is the course's definition of data science (Lecture 1 slides 49–51). It matches the Venn diagram of computer science, statistics and domain knowledge; the human perspective is the 'domain knowledge' circle. Its cycle (preprocessing → exploration → analysis → interpretation → communication) mirrors the data science lifecycle used from Lecture 3 onwards. Its statistical perspective (uncertainty, cause vs correlation) links to the data science vs statistics slides (Lecture 1 slides 53–55) and the spurious-correlation example in Lecture 4 (two things that rise together by chance). Machine learning for prediction and the bootstrap come back in Lectures 5–7.

Remember

  • Opinion piece (PNAS 2017): data science is 'the child of statistics and computer science'.
  • Three perspectives: statistical, computational, human. The core is combining all three.
  • Statistical: uncertainty, complex/structured data, very many variables (high dimensionality), cause vs correlation.
  • Computational: running methods as algorithms; trade-off between accuracy and computer resources (time, memory); optimisation, sampling (bootstrap, MCMC), distributed computing.
  • Human: cannot be fully automated; needs domain knowledge and step-by-step collaboration with domain experts.
  • Data science is a cycle: preprocessing → exploration → selection → transformation → analysis → interpretation → communication.
  • Why the label arose: abundant data, but existing methods are not built for modern problems.
  • Take-home: understand the data's context, take responsibility for private/public data, say clearly what the data can and cannot tell us.

Practice

Q1 · What data science is

According to Blei & Smyth (2017), what is the core of data science?

Show answer

Answer: C. The authors say each perspective is essential, but that combining all three is what data science is about. Option B contradicts their human perspective: data science cannot be fully automated.

Source: Paper p. 8689–8690

Q2 · Computational perspective

Which statement best describes what Blei & Smyth call the computational perspective of data science?

Show answer

Answer: A. Computational thinking is about how methods are implemented and the trade-off between accuracy and computer resources. Option B describes the statistical perspective and option C the human perspective.

Source: Paper p. 8690–8691

Q3 · Human perspective

Why, according to Blei & Smyth (2017), can data science not be fully automated?

Show answer

Answer: D. The human perspective says that understanding the field, choosing data, exploring, picking models and communicating results all need judgement and domain knowledge. Option B is false: the authors stress that models of the world are always somewhat wrong ('misspecified').

Source: Paper p. 8690–8691

Q4 · Why data science emerged

What do Blei & Smyth see as the trigger for the new label 'data science'?

Show answer

Answer: B. The authors describe a tension: scientists have abundant data, but classical methods cannot fully use it, and this tension gave rise to 'data science'. Option A is the opposite of their starting point, which is that data are abundant.

Source: Paper p. 8689–8690

Q5 · Three perspectives applied 3 points

A football club wants to use its players' GPS and injury data. Blei & Smyth (2017) say data science combines a statistical, a computational and a human perspective. State and motivate briefly what each perspective contributes in this project:
1) statistical;
2) computational;
3) human.

Show model answer

1) Statistical: model the data while taking uncertainty into account, and ask whether a link is causal or only a correlation (e.g. does high training load cause injury, or do both just rise together?). 2) Computational: run the analysis efficiently on large GPS data streams, balancing accuracy against computing time and memory (e.g. optimisation, resampling such as the bootstrap, spreading the work over computers). 3) Human: coaches and sport scientists bring domain knowledge to choose relevant data, interpret the results and communicate what the data can and cannot say. The data scientist works with them step by step; data science cannot be fully automated.

How the points are earned

  • 1 pt Statistical: handling uncertainty and/or cause vs correlation, applied to the example
  • 1 pt Computational: efficient algorithms, trade-off between accuracy and time/memory on large data
  • 1 pt Human: domain knowledge of coach/expert to choose data, interpret and communicate; cannot be fully automated

Source: Paper p. 8690–8691; Lecture 1 Slides p. 49–51

Lecture 1 – Introduction to data science

Perspective (non-technical overview; no new data)

Chmait & Westerbeek (2021) — Artificial Intelligence and Machine Learning in Sport Research: An Introduction for Non-data Scientists

Chmait, N., & Westerbeek, H. (2021). Artificial intelligence and machine learning in sport research: An introduction for non-data scientists. Frontiers in Sports and Active Living, 3, 682287. https://doi.org/10.3389/fspor.2021.682287 · Open paper

A plain-language overview for people in sport who are not data scientists. It explains how machine learning differs from classic rule-based analysis, shows supervised, unsupervised and reinforcement learning with sport examples, and discusses what AI means for the future of sport (privacy, data ownership, ethics).

Question

The authors want to give sport business professionals, coaches, policy makers and other non-technical readers a broad overview of the AI and machine learning methods used for sport performance and sport business problems. AI (artificial intelligence) means computer systems that do tasks that normally need human intelligence; machine learning (ML) is the part of AI where the computer learns patterns from data. The authors note that for many non-experts the link between AI and sport is still 'fuzzy' and the reasons for using ML are unclear. They also discuss how AI could shape sport in the coming years.

Why it matters

Sport is using AI more and more, driven by more computing power and more data. It started with Moneyball and SABRmetrics in baseball (using in-game statistics to find undervalued players) and now includes injury models, player tracking and ticket pricing. Decision-makers need to understand how data scientists think, so they can discuss the approach and method with them without needing the technical details.

Approach

No data are collected; it is a narrative essay. The authors first summarise earlier work on AI in sport: expert systems (programs built from rules written by experts), artificial neural networks and deep learning (models loosely inspired by the brain that learn complex patterns from many examples), evolutionary computation (methods that 'breed' better solutions, mimicking natural selection), Bayesian approaches (probability models updated with data), and the survey by Beal et al. (2019). They list four application areas. They then contrast classic analytics (rules + data → program → answers) with the ML way of working (data + known answers → the algorithm works out the rules itself, which are then tested on new data it has never seen). They write an ML prediction as f(w1·i1, …, wn·in) = y: each input i gets a weight w, and the model combines them into an output y. Next come three made-up sport examples: supervised injury prediction in Australian football, unsupervised fan grouping with K-means clustering (an algorithm that puts similar data points in the same group, where k is the number of groups), and reinforcement learning (Q-learning, SARSA: algorithms that learn by trial and error from rewards) for fantasy sport and game-playing AI, plus a short note on genetic/evolutionary algorithms. They end with views on the future of AI in sport.

Findings

  • Machine learning turns classic analytics around: instead of programming the rules, you give the algorithm data together with the known answers, it learns the rules, and those rules are checked by testing accuracy on new (unseen) data.
  • Supervised learning (the outcome is known in the past data): e.g. for each Australian football player you know whether he got injured and missed the next match, plus his match load, metres run, warm-up and tackles. The model learns from this, is tested on unseen data and is adjusted until accuracy is good enough (their example figure: 70%). Neural networks, decision trees or regression models can all do this.
  • Unsupervised learning (no outcome labels): the algorithm finds patterns nobody had noticed. Example: a football club groups its stadium visitors by gender, age, postcode or income with K-means. The groups come from the data, although the number of groups can be chosen in advance.
  • Reinforcement learning: an agent (the decision-making program) acts in a (simulated) environment, gets rewards or penalties, and learns a strategy that collects the most reward over time. Examples are picking fantasy sport teams and game AI (chess, Go, poker, StarCraft).
  • AI/ML in sport covers four areas: game activity and analytics, talent identification and recruitment, training and coaching, and fan- and business-focused applications.
  • Future: AI is unlikely to fully replace coaches and human experts. The main barriers to complete '360-degree' analyses of an athlete's value are that player, team and commercial data are kept private by their owners (proprietary) and privacy rules. Who owns the data, and ethics (e.g. a team's injury-prediction model versus a player's right to his own data), will become contested.

Limitations

  • It is an opinion piece, not a data study or systematic review. The authors say they do not try to cover the literature completely, so the examples are selective.
  • The supervised, unsupervised and reinforcement learning examples are made up, so no real model performance is reported (the 70% accuracy is only an illustration).
  • Own inference: it is deliberately simple and non-technical. It does not explain how to validate a model properly (e.g. cross-validation, overfitting), and its unsupervised example loosely says the algorithm will 'classify' visitors, although clustering forms groups without any labels.

Link to the lectures

Slide 52 shows the paper's four application areas, credited as 'Chmait & Westerpoort 2021'; its three types of machine learning return in Lecture 5 slides 8–16.

This paper is the course's plain-language map of the types of machine learning. The injury example is supervised learning (known outcome; a yes/no outcome makes it classification), the fan grouping is unsupervised clustering with K-means (Lecture 6), and the fantasy-sport and game example is reinforcement learning (Lecture 5 slides 14–15). Its point that learned rules must be tested on new, unseen data is the idea behind train/test splits and cross-validation (Lecture 5). Lecture 1 slide 52 uses its four application areas of data science in sport.

Remember

  • Opinion piece for people who are not data scientists; no new data.
  • Classic analytics: rules + data → answers. Machine learning: data + known answers → the algorithm learns the rules → tested on unseen data.
  • Supervised = the outcome (label) is known in the training data (e.g. injured yes/no) → predict it for new cases.
  • Unsupervised = no labels; the algorithm finds groups itself (e.g. K-means grouping of fans; k can be set in advance).
  • Reinforcement = an agent learns by trial and error to collect the most reward (Q-learning; games, fantasy sport).
  • Four application areas: game analytics, talent identification, training & coaching, fan & business.
  • Barriers: private (proprietary) data, privacy, data ownership. AI supports coaches rather than replacing them.

Practice

Q1 · Supervised learning

Chmait & Westerbeek (2021) describe a model that predicts muscle-strain injuries in Australian football players from match load, metres run, warm-up and tackles. What makes this a supervised learning problem?

Show answer

Answer: B. The key point is that the outcome (injury or not) is known in the past data used for training. Option A describes reinforcement learning, and option D is the opposite of what the authors describe: the model is tested on new, unseen data.

Source: Paper p. 4

Q2 · Unsupervised learning

A football club has data on its stadium visitors (age, gender, postcode, income) but no existing customer categories. It wants to discover groups of similar fans for its marketing. Which approach do Chmait & Westerbeek describe for this?

Show answer

Answer: C. Fan grouping is their unsupervised example: there are no labels, and K-means forms the groups from the data (the number of groups may be chosen beforehand). Supervised classification would need fans whose type is already known.

Source: Paper p. 4–5

Q3 · How ML differs

According to Chmait & Westerbeek (2021), how does machine learning differ from classic sports analytics?

Show answer

Answer: D. This reversal (data + answers → rules, then a test on unseen data) is the paper's central explanation of machine learning. Option A swaps the two, and option C contradicts the authors' point that learned rules must be tested on new data.

Source: Paper p. 3–4

Q4 · Barriers for AI in sport

Chmait & Westerbeek note that no study yet gives a '360-degree' analysis of an athlete's total value (sport performance plus business value such as ticket sales). What main obstacle do they name?

Show answer

Answer: A. The authors name private (proprietary) data, privacy rules and data ownership as the main challenges. Option D contradicts their description of clubs routinely recording visitor characteristics.

Source: Paper p. 6

Q5 · Three types of machine learning 3 points

Chmait & Westerbeek (2021) explain machine learning to sport professionals with three examples: injury prediction, grouping stadium visitors, and picking a fantasy-sport team. State and motivate briefly, for each example, which type of machine learning it is and what the model learns from:
1) injury prediction;
2) grouping visitors;
3) fantasy-sport team.

Show model answer

1) Supervised learning: the past data contain the inputs (load, metres run, tackles) and the known outcome (injured or not); the model learns the link and is tested on new, unseen players. 2) Unsupervised learning (clustering, e.g. K-means): there is no outcome label; the algorithm finds groups of similar visitors from their age, gender, postcode or income. 3) Reinforcement learning: an agent tries team choices, gets rewards or penalties (e.g. the fantasy points its team scores), and learns a strategy that collects the most reward over time.

How the points are earned

  • 1 pt Injury: supervised, because the outcome (injured or not) is known in the training data
  • 1 pt Visitors: unsupervised clustering (e.g. K-means), no labels, groups found from the data
  • 1 pt Fantasy team: reinforcement learning, learns by trial and error from rewards/penalties

Source: Paper p. 4–5

Lecture 3 – Data in sport and health

Original research (cross-sectional)

van der Zwaard et al. (2018) — Critical determinants of combined sprint and endurance performance: an integrative analysis from muscle fiber to the human body

van der Zwaard, S., van der Laarse, W. J., Weide, G., Bloemers, F. W., Hofmijster, M. J., Levels, K., Noordhof, D. A., de Koning, J. J., de Ruiter, C. J., & Jaspers, R. T. (2018). Critical determinants of combined sprint and endurance performance: an integrative analysis from muscle fiber to the human body. The FASEB Journal, 32(4), 2110–2123. https://doi.org/10.1096/fj.201700827R · Open paper

In 28 cyclists, the better a rider was at sprinting, the worse he tended to be at endurance (after correcting for body size). By relating performance to many measurements, from the whole body down to single muscle fibres, the authors found what explains sprint, endurance and the ability to combine both. They point to long muscle fibres and many capillaries as training targets for being good at both.

Question

Which body and muscle characteristics, measured from the whole body down to single muscle fibres, explain sprint performance, endurance performance and the ability to combine both? The underlying problem: large muscle fibres give power for sprinting, and fibres with a high oxidative capacity (the ability to use oxygen to make energy) give endurance, but large fibres tend to have a low oxidative capacity. So improving both at once is hard.

Why it matters

Many sports need both sprint and endurance. Earlier studies looked at only one of the two, and mostly at whole-body measures. Training for endurance and training for peak power interfere with each other, partly inside the muscle cells. Knowing the critical determinants (the factors that matter most) gives targets for training. The two performance tests plus a physiological profile could also help with talent identification and individual training plans.

Approach

This was a cross-sectional study: each cyclist was measured once, with no follow-up over time. It included 28 cyclists (14 road, 8 team pursuit, 6 track sprint), all at national to Olympic level except 4 amateur road cyclists. Sprint performance was the highest 1-second power in a Wingate test (a 30-second all-out sprint on a bike). Endurance performance was the average power in a 15-km time trial. Both were divided by lean body mass (body weight without fat) to the power 2/3, to remove the advantage that bigger riders have just from being bigger. The authors measured many possible determinants: whole-body oxygen uptake (VO2) during the time trial ('performance VO2'), maximal oxygen uptake (VO2max) and thresholds; gross efficiency (the share of the body's energy use that ends up as power on the pedals); blood values such as haemoglobin (Hb, the oxygen-carrying protein in red blood cells) and MCHC (how much haemoglobin is packed into the red blood cells); muscle oxygenation (how much oxygen the thigh muscle holds, measured with a light sensor on the skin, near-infrared spectroscopy); maximal knee-extension force; 3D ultrasound of the outer thigh muscle (vastus lateralis) for muscle volume, fascicle length (length of the fibre bundles), PCSA (physiological cross-sectional area: how thick the muscle is measured across its fibres) and fibre angle; and a muscle biopsy (a small piece of muscle taken with a needle), stained to measure fibre type (fast or slow), fibre size, oxidative capacity (with the enzyme SDH), capillaries (tiny blood vessels that bring oxygen) and myoglobin (the protein that stores oxygen in muscle). To score how well each rider combined sprint and endurance, they fitted a line through all riders' sprint and endurance scores with Deming regression (a straight-line fit that allows for measurement error in both variables, not only in the y-variable). Each rider's score was his perpendicular distance to that line (his residual): above the line means better at combining both than the group trade-off predicts. Links were tested with Pearson correlations (r: the strength of a straight-line relationship, from −1 to 1) and stepwise multiple regression: predictors were added one at a time and kept only if they significantly raised R² (the share of the differences between riders that the model explains, from 0 to 1; P < 0.05).

Findings

  • Sprint and endurance performance (corrected for body size) were inversely related: r = −0.66 (P < 0.001); per kg body mass r = −0.44. So being good at both is difficult.
  • Sprint: the percentage of fast fibres and the volume of the thigh muscle together explained 65% of the differences in sprint power (R² = 0.65).
  • Endurance: performance VO2, MCHC and muscle oxygenation together explained 92% of the differences in time-trial power (R² = 0.92). Performance VO2 itself was explained by fibre oxidative capacity, myoglobin × capillaries per fibre and, negatively, PCSA (R² = 0.67). So a thicker muscle (hypertrophy, growth in muscle size) works against oxygen use.
  • Combining both was explained by gross efficiency and performance VO2, and probably by muscle volume and fascicle length (P = 0.056 and P = 0.059). In the whole group, gross efficiency plus muscle volume explained 30%.
  • Inside the muscle, fibre size and fibre oxidative capacity were inversely related (r = −0.50, a curved, hyperbolic relationship). But riders who combined fairly large fibres with a high oxidative capacity had more capillaries.
  • Take-home: long fascicles (rather than a thick muscle) and many capillaries are suggested as training targets for improving sprint and endurance at the same time.

Limitations

  • Authors: the design is cross-sectional, so it shows associations, not what training does over time. The authors call for longitudinal training studies (following athletes over time); the training targets are hypotheses, not proven effects.
  • Authors: some possible determinants were not measured (e.g. how much blood the heart pumps, oxygen transfer in the lungs, nerve control of the muscle, muscle metabolites, tendon properties, glycogen), and the track sprinters' different muscle build may have left more variance unexplained.
  • Own inference: a small, mixed sample (28 riders; only 6 sprinters) and many candidate predictors in stepwise regression. No test on new riders (no train/test split or cross-validation) is reported, so the high R² values describe these 28 riders and may be too optimistic for others (overfitting).

Link to the lectures

The main case study, slides 14–30.

In Lecture 3 this is the main example of data acquisition: one question answered by combining many kinds of data from many levels (performance tests, oxygen uptake, blood, ultrasound images, muscle biopsies). The lecture shows the problem (fibre size vs oxidative capacity), the 28 cyclists and tests, the inverse sprint–endurance plot (r = −0.66, with the perpendicular residuals drawn as arrows), the explained-variance bar chart (Fig. 3), the 67% model for performance VO2 and the training targets. In data-science terms it is a regression problem with a known, numeric outcome (power corrected for body size), so it is supervised learning (learning from examples where the outcome is known). But it is used to explain differences in this group (R² = explained variance, Lecture 4), not to predict new athletes with a tested model. That would need a hold-out test set or cross-validation (Lecture 5): repeatedly fitting on part of the data and checking on the part left out, to catch overfitting (a model that fits noise in its own data and does worse on new data). Correcting for lean body mass^(2/3) and building a combined score from Deming residuals are examples of feature engineering (making new, more useful variables from raw data). Contrast with van der Zwaard et al. (2019) in Lecture 6, which groups cyclists without an outcome (k-means clustering, unsupervised).

Remember

  • Aim: find the critical determinants of combined sprint and endurance performance, from whole body down to muscle fibre.
  • 28 cyclists (track sprint, team pursuit, road); sprint = Wingate peak power; endurance = 15-km time-trial power; both corrected for body size (lean body mass^(2/3)).
  • Sprint and endurance are inversely related (r = −0.66): being good at both is difficult.
  • Combined performance = each rider's perpendicular distance (residual) to the Deming regression line (a line fit that allows for error in both variables).
  • Analysis: Pearson correlations + stepwise multiple regression (explained variance, R²).
  • Sprint ← % fast fibres + muscle volume (65%); endurance ← performance VO2 + MCHC + muscle oxygenation (92%); combined ← gross efficiency + performance VO2 (probably also muscle volume and fascicle length).
  • Take-home: long fascicles/fibres and capillaries are training targets; but the design is cross-sectional, so no proof of training effects.

Practice

Q1 · Main finding

Which statement best summarises the main findings of van der Zwaard et al. (2018) on sprint and endurance performance in cyclists?

Show answer

Answer: D. Sprint and endurance (corrected for body size) correlated at r = −0.66; fast fibres + muscle volume explained 65% of sprint and performance VO2 + MCHC + muscle oxygenation explained 92% of endurance. Option B is the tempting reverse: fibre size and oxidative capacity were inversely related (r = −0.50).

Source: Paper p. 2110 (abstract), pp. 2113–2116; Lecture 3 slides pp. 23, 27–29

Q2 · Combined performance score

Van der Zwaard et al. (2018) needed one number for how well each cyclist combined sprint and endurance. How did they get it?

Show answer

Answer: B. The score was the perpendicular residual to the Deming line: how far a cyclist lies above or below the group's sprint–endurance trade-off. Option A is tempting because the paper labels the score 'POpeak + POTT', but that label is notation for the residual-based score, not a sum.

Source: Paper pp. 2111–2112, Fig. 2 (p. 2116); Lecture 3 slides p. 23

Q3 · Research problem

What problem motivated the study of van der Zwaard et al. (2018)?

Show answer

Answer: A. The study starts from the trade-off: muscle size drives sprint, oxidative capacity drives endurance, and fibre size and oxidative capacity are inversely related, so combining both is difficult. Option D is tempting but wrong: the study explains differences between riders measured once; it does not forecast future results.

Source: Paper pp. 2110–2111 (Introduction); Lecture 3 slides pp. 15–19

Q4 · Limitation of the conclusion

Van der Zwaard et al. (2018) suggest fascicle length and capillaries as training targets. What is the main reason to treat this advice with caution?

Show answer

Answer: C. Each rider was measured once, so the study cannot show that training these properties improves performance; the authors call for longitudinal training studies. The other options are factually wrong: most riders were national to Olympic level, endurance was a 15-km time trial, and endurance R² was 0.92.

Source: Paper pp. 2120–2121 (Implications, Limitations)

Q5 · Explaining vs predicting 3 points

Van der Zwaard et al. (2018) explained 92% of the differences in endurance power (R² = 0.92) with stepwise multiple regression in 28 cyclists. The same 28 cyclists were used to build and to evaluate the model. State and motivate briefly:
1) whether this is supervised or unsupervised learning;
2) why R² = 0.92 may overestimate how well the model would predict new cyclists, and one way to check this.

Show model answer

1) Supervised: every cyclist has a measured outcome (endurance power), and the model learns to explain that known outcome from the predictors. Because the outcome is a number, it is a regression problem. 2) With few cyclists and many candidate predictors, stepwise selection can pick predictors that fit noise in these 28 riders (overfitting), so R² on the same data is too optimistic. Check it with a hold-out test set, k-fold cross-validation, or by testing the model on a new group of cyclists.

How the points are earned

  • 1 pt Supervised, motivated by a known/measured outcome (regression on a numeric outcome)
  • 1 pt Explains overfitting: few cyclists, many predictors, model fits noise of this sample, so in-sample R² is optimistic
  • 1 pt Names a way to check: hold-out test set, cross-validation or a new group of cyclists

Source: Paper p. 2113 (Statistical analysis), p. 2121 (Limitations); Lecture 5 slides pp. 43–44, 49, 60

Lecture 5 – Machine learning: regression

original research (original investigation)

Jaspers et al. (2018) — Relationships Between the External and Internal Training Load in Professional Soccer: What Can We Learn From Machine Learning?

Jaspers, A., Op De Beéck, T., Brink, M. S., Frencken, W. G. P., Staes, F., Davis, J. J., & Helsen, W. F. (2018). Relationships between the external and internal training load in professional soccer: What can we learn from machine learning? International Journal of Sports Physiology and Performance, 13(5), 625–630. https://doi.org/10.1123/ijspp.2017-0299 · Open paper

With two seasons of GPS and accelerometer data from a Dutch top-league soccer team, two machine learning models (LASSO and a neural network) predicted how hard players rated each training session (session RPE) better than a naive baseline that always guesses the average. In the main test (train on season 1, test on season 2) LASSO beat the neural network, braking movements (decelerations) turned out to be important, and one model for the whole group worked as well as or better than a separate model per player.

Question

Can machine learning predict a player's session RPE from many external load indicators measured with GPS and accelerometers? Session RPE (rating of perceived exertion) is the player's own score of how hard the session felt, on a 0–10 scale; it measures internal load (the body's response). External load indicators (ELIs) describe what the player physically did, such as distance, speed, accelerations and decelerations. The authors also asked which ELIs matter most for RPE in soccer, and whether a separate model per player beats one model for the whole group, i.e. whether players differ in meaningful ways.

Why it matters

Clubs track training load to improve fitness and lower injury risk. The same external load can feel very different to different players, so understanding how external load turns into internal load helps clubs manage training. Earlier studies used traditional statistics (correlation, multiple regression) on a few hand-picked ELIs. The only machine learning study was in Australian football, which has different physical demands, so its results may not hold for soccer.

Approach

38 professional outfield players from a team in the highest Dutch league were followed over two seasons (2014–15 and 2015–16). Only training sessions were used (no matches, recovery or rehabilitation sessions): 5917 sessions in total. External load came from GPS units (10 Hz, i.e. 10 measurements per second) and accelerometers (100 Hz; sensors that measure changes in speed and direction). From these the authors computed 67 ELIs on duration, distance, speed, accelerations and decelerations, PlayerLoad (an accelerometer score of total body movement) and repeated high-intensity efforts. The target (the value to predict) was session RPE on the modified Borg CR-10 scale (0–10), reported about 30 minutes after training. Three models were compared: an artificial neural network (ANN; a flexible model loosely inspired by the brain that can learn complex, non-linear patterns but is hard to interpret); LASSO (linear regression with a penalty that pushes many coefficients to exactly zero, so it keeps only the most useful ELIs); and a naive baseline that ignores all ELIs and always predicts the average RPE. Accuracy was measured on a separate test set (data the model did not learn from) with MAE (mean absolute error: the average size of the prediction error, in RPE points; lower is better). Experiment 1 trained group models on season 1 and tested them on season 2, so the test came later in time and many test players were new. Experiment 2 used the first 75% of each season to train and the last 25% to test, and compared group models (all players) with individual models (one per player). LASSO was also used to rank the ELIs by importance (how often each ELI was selected).

Findings

  • Both machine learning models beat the naive baseline. Experiment 1 MAE: LASSO 0.80, ANN 1.09, baseline 1.14 RPE points. LASSO cut the error by 29.8% compared with the baseline (a small effect, d = 0.44, where d is an effect size: the difference in standard-deviation units); the ANN improved only trivially (d = 0.06).
  • LASSO was more accurate than the ANN: its MAE was 26.6% lower than the ANN's in experiment 1.
  • The most important ELIs included high-intensity accelerations (> 3.5 m·s−2), repeated high-intensity efforts, distance covered while decelerating hard, running distance at 12–20 km·h−1, PlayerLoad and session duration. Decelerations were the new finding: braking is eccentric work (the muscle lengthens while it contracts), which can damage muscle.
  • Group models were as good as or better than individual models in both seasons (e.g. season 1 LASSO: group 0.79 vs individual 0.81 vs baseline 0.99; season 2: 0.85 vs 0.85 vs 1.11). This differs from the earlier Australian-football study. The authors explain it by far more data for a group model (> 2000 vs < 100 sessions) and a soccer squad that is more alike in load and characteristics.
  • Practical message: machine learning can help experts choose which ELIs to monitor, and a group model can predict RPE for new, transferred or youth players who have little data of their own.

Limitations

  • Paper: only training sessions were used; matches were excluded, and a season may have too few matches for machine learning.
  • Paper: only overall session RPE was predicted, and the inputs did not include other factors such as separate RPE for legs or breathing, wellness, recovery or psychological and social factors.
  • Paper: individual models had fewer than 100 sessions per player. With 100–150 sessions per player per season and frequent transfers, much more data per player is unrealistic.
  • Partly paper, partly own inference: the gain over the baseline was modest (small effect; MAE about 0.8 on a 0–10 scale), and all data came from one team with one GPS system, so comparing with or applying to other teams is hard.

Link to the lectures

Listed with Lecture 5 on Canvas as a worked example of supervised regression with LASSO, MAE, a naive baseline and a hold-out test set; the slides do not discuss the paper itself.

A clear example of supervised regression from Lecture 5: a numeric target (RPE) predicted from many predictors (ELIs), with the true RPE known for every session. It shows LASSO (its penalty can push coefficients to zero, which selects features automatically and copes with many or correlated features and small samples) against a flexible but hard-to-interpret neural network. It also shows evaluation with MAE in the target's own units, comparison with a naive baseline, and a hold-out test that respects time order (train on the earlier season, test on the later one). Group vs individual models shows the trade-off between having more data and personalising the model.

Remember

  • Supervised regression: predict session RPE (internal load: how hard it felt) from 67 GPS and accelerometer ELIs (external load: what the player did).
  • 38 professional soccer players, 2 seasons, training sessions only (5917 sessions).
  • Models: neural network (ANN), LASSO and a naive baseline that always predicts the average RPE. Metric: MAE (lower is better).
  • Both machine learning models beat the baseline; LASSO was best (train season 1 → test season 2 MAE: LASSO 0.80 vs ANN 1.09 vs baseline 1.14).
  • LASSO sets coefficients to zero, which selects ELIs and keeps the model easy to interpret. Decelerations (braking) were newly found to be important.
  • Group models ≥ individual models (more data, similar players), so a group model can be used for new players.
  • Test in time order: train on season 1, test on season 2; experiment 2 used the first 75% vs the last 25% of each season.

Practice

Q1 · Type of ML problem

Jaspers et al. (2018) used GPS and accelerometer measures to predict how hard each soccer player rated a training session on a 0–10 RPE scale. For every past session the reported RPE was known. How is this machine learning task best described?

Show answer

Answer: C. The reported RPE is the known outcome for each session, and the models predict a number that is scored with mean absolute error, so this is supervised regression. Option B is tempting, but the sessions were never put into categories: the models predicted the RPE value itself.

Source: Paper p. 1 (journal p. 625, abstract), p. 3 (p. 627)

Q2 · Why LASSO

Jaspers et al. compared LASSO with an artificial neural network. Why was LASSO an attractive choice?

Show answer

Answer: A. The paper describes LASSO as linear regression that biases many coefficients to 0, which selects features and makes the model easier to interpret and more robust. Option B describes the neural network, the flexible but hard-to-interpret option. Option D describes Ridge regression, and LASSO was still evaluated on a separate test set.

Source: Paper p. 3 (p. 627); Lecture 5 slides p. 58, 60

Q3 · Main result: MAE vs baseline

Trained on season 1 and tested on season 2, the LASSO model had a mean absolute error (MAE) of 0.80, while a naive baseline that always predicts the average RPE had an MAE of 1.14. What does this mean?

Show answer

Answer: D. MAE is the average distance between predicted and reported RPE, in RPE points, and lower is better; LASSO cut the baseline's error by 29.8%. Option A confuses MAE with R² (the share of differences explained), which is a proportion and can never be 114%.

Source: Paper p. 3 (p. 627), Table 2; Lecture 5 slides p. 41

Q4 · Why the study matters

Why did Jaspers et al. (2018) want to model the link between external load (what a player does) and internal load (how hard it feels)?

Show answer

Answer: B. The introduction stresses that equal external loads can cause different internal loads, that load monitoring aims at fitness and injury prevention, and that earlier work used traditional statistics on few, hand-selected measures. Option A is wrong: players did report RPE after every session, and RPE was the outcome the models learned to predict.

Source: Paper p. 1–2 (journal p. 625–626, Introduction)

Q5 · Testing the models 2 points

Jaspers et al. trained their models on season 1, tested them on season 2, and compared every model with a naive baseline that always predicts the average RPE. State and motivate briefly:
1) why testing on the later season makes the evaluation convincing;
2) why the comparison with the naive baseline is needed.

Show model answer

1) Season 2 is a separate test set that the model never learned from, and it lies later in time. This mirrors real use (build the model on past data, predict future sessions, even for new players) and avoids the too-optimistic error you get when you test on the training data. 2) The baseline uses no load data at all, so it shows how big the error is without any real model (a realistic upper limit for the MAE). Only a clearly lower MAE (LASSO 0.80 vs 1.14) shows that the load measures add real predictive information.

How the points are earned

  • 1 pt Separate, later test set: model judged on unseen (future) data, as in real use; avoids an optimistic error estimate
  • 1 pt Baseline shows the error without using load data; a model is only useful if its MAE is clearly lower

Source: Paper p. 3 (p. 627), Table 2

Lecture 6 – Classification and clustering

original research (proof-of-concept study)

Biswas et al. (2015) — Recognizing upper limb movements with wrist worn inertial sensors using k-means clustering classification

Biswas, D., Cranny, A., Gupta, N., Maharatna, K., Achner, J., Klemke, J., Jöbges, M., & Ortmann, S. (2015). Recognizing upper limb movements with wrist worn inertial sensors using k-means clustering classification. Human Movement Science, 40, 59–76. https://doi.org/10.1016/j.humov.2014.11.013 · Open paper

With one motion sensor on the wrist, the authors recognised three basic arm movements while people made a cup of tea. They used k-means clustering in a supervised way: groups were built from each person's labelled practice movements, and each new movement got the label of the nearest group. Accuracy was about 88% (accelerometer) and 83% (gyroscope) in healthy people and about 70% and 66% in stroke patients, better than two standard classifiers (LDA and SVM).

Question

Can one sensor on the wrist recognise three basic forearm movements while a person moves freely in daily life? The three movements are: (A) reaching for an object and bringing it back, (B) lifting a cup to the mouth, and (C) pouring or turning a key or lid (rotating the wrist). The long-term aim is a tool that counts how often a stroke (or cerebral palsy) patient uses the impaired arm for these movements outside the clinic.

Why it matters

If you can count how often the impaired arm makes these movements at home, you can follow rehabilitation from a distance: as arm function recovers, the movements should become more frequent. Earlier activity-recognition studies mostly detected whole-body activities and postures (sitting, walking, running), often in the lab. Fine arm movements in real life, with few sensors and little computing power, had received little attention.

Approach

Four healthy men and four stroke patients wore a Shimmer sensor on the back of the wrist. It is an inertial sensor: an accelerometer (measures acceleration in 3 directions, called axes) plus a gyroscope (measures rotation speed around 3 axes), recording 50 times per second (50 Hz). Training data: many repetitions of each movement under controlled conditions, labelled by the researchers (for example 480 trials per healthy person). Test data: the same person, on another day, doing a 20-step 'making a cup of tea' task freely; the researchers marked by hand where each movement started and ended. The signals were filtered to remove slow drift and fast noise. Then 10 features (summary numbers per movement) were computed for each axis, giving 30 per sensor. Examples: standard deviation (how much the signal varies), RMS (root mean square: the average size of the signal), entropy (how irregular it is), jerk (how quickly acceleration changes), number of peaks, kurtosis (how peaked) and skewness (how lopsided). The features were ranked by how well they separated the three movements and then added one at a time (sequential forward selection), keeping the set that worked best. The model was k-means clustering. k-means normally works without labels: it splits data into k groups (clusters), each with a centre (centroid); every point joins the nearest centre, the centres are recomputed, and this repeats until nothing changes. Here it was used in a supervised way (learning from examples with known answers): k was set to 3 because there were 3 known movements, each cluster got the label of the movement whose average was closest, and each new movement was labelled by its nearest cluster centre (a minimum-distance classifier). Distance was measured in two ways: Euclidean distance (the ordinary straight-line distance) or Mahalanobis distance (a distance that takes the cluster's shape into account, so a point counts as closer along a direction in which the cluster is widely spread). The feature sets were compared with 10 runs of 10-fold cross-validation on the training data (split the data into 10 parts, train on 9, test on the 10th, repeat). Each person got their own model per sensor (personalised models). The method was compared with two standard classifiers using the same features: LDA (linear discriminant analysis, which separates groups with straight boundaries) and SVM (support vector machine, which finds the boundary with the widest gap between groups).

Findings

  • Healthy people: overall accuracy (share of movements labelled correctly) 61–100%, mean 88%, with accelerometer data, and 60–94%, mean 83%, with gyroscope data.
  • Stroke patients: lower accuracy, 40–88% (mean 70%) with the accelerometer and 40–83% (mean 66%) with the gyroscope, because their movements were more variable and less repeatable. For example, one patient early in rehabilitation, tested on the non-dominant arm, reached only 40%.
  • The number and type of best features differed per person (from 2 to 30). Healthy people shared top features (for example the accelerometer's standard deviation and RMS on the y-axis); stroke patients had little overlap.
  • For stroke patients, Mahalanobis distance often worked better than Euclidean distance, because their clusters were stretched out (elongated) in some directions instead of round.
  • With the same features, LDA (mean accuracy 45–53%) and SVM (mean 50–68%) did worse than the clustering method. The accelerometer and gyroscope sometimes made up for each other's weak spots.

Limitations

  • Very small sample (4 healthy people, 4 stroke patients): this is a proof of concept. The authors plan a larger sample and other sensor positions.
  • The start and end of each test movement were marked by hand. Real-world use needs automatic detection of where movements start and end, which was not addressed (paper).
  • Personalised models need a long labelled training session for each person (up to 480 trials) and re-training as the patient recovers (partly paper, partly own inference).
  • Own inference: the LDA and SVM comparison used features chosen for the clustering method, and one setting (a 25% limit on cluster size) was chosen because it 'produced the best results'. Both choices may favour the authors' own method.

Link to the lectures

Reference literature on the Canvas Lecture 6 page, not on the slides; Canvas calls it 'Biswas et al 2018', but the paper is from 2015.

This paper shows that the line between supervised and unsupervised methods can blur: k-means, normally unsupervised, is used here as a supervised classifier with three classes, because labelled training data fix k = 3 and give each cluster its label. It shows how k-means works (centroids, repeated reassignment of points, a squared-distance cost) and its assumption of round clusters with Euclidean distance, versus Mahalanobis distance for stretched clusters. It also covers turning sensor signals into features and choosing among them (feature engineering and selection), 10-fold cross-validation on training data plus a separate, more realistic test set, and evaluation with a confusion matrix (table of true vs predicted movements), sensitivity per movement (share of that movement recognised correctly) and overall accuracy, compared against LDA and SVM.

Remember

  • Aim: recognise 3 arm movements (reach and retrieve, cup to mouth, pour or turn) with one wrist sensor, to monitor stroke rehabilitation at home.
  • k-means used in a supervised way: labelled training trials, k = 3, each cluster named after the nearest movement, and each new movement labelled by the nearest cluster centre.
  • Trained on controlled repetitions; tested on a free 'making a cup of tea' task on another day (more realistic, usually lower accuracy).
  • Features: 10 summary numbers × 3 axes per sensor, ranked and added one at a time, chosen with 10 × 10-fold cross-validation; one personalised model per person.
  • Accuracy: healthy about 88% (accelerometer) and 83% (gyroscope); stroke about 70% and 66%.
  • Mahalanobis distance helps for stroke patients, whose clusters are stretched and variable.
  • Beat LDA and SVM with the same features, but only 4 + 4 participants.

Practice

Q1 · Why the study matters

Why did Biswas et al. (2015) want to recognise arm movements with a wrist sensor during daily activities?

Show answer

Answer: B. The aim is remote monitoring of rehabilitation: these movements should become more frequent as arm function recovers. Option C is tempting but reversed: whole-body activities such as sitting and walking were what earlier studies already detected; fine arm movements were the gap.

Source: Paper p. 2–3 (journal pp. 60–61, Introduction)

Q2 · k-means used as a classifier

k-means normally finds groups in data without any labels (unsupervised learning). How did Biswas et al. use it?

Show answer

Answer: D. The authors knew the movement labels, so they fixed k = 3, labelled each cluster and classified test movements by the nearest centre; that is supervised use of k-means. Option A is how k-means is normally used, but here nothing was discovered without labels.

Source: Paper p. 4 (journal p. 62), p. 10 (p. 68)

Q3 · Euclidean vs Mahalanobis distance

Euclidean distance is the straight-line distance to a cluster centre. Mahalanobis distance also takes the cluster's shape (how widely it spreads in each direction) into account. For the stroke patients, Mahalanobis distance often worked better. Why?

Show answer

Answer: C. Variable movements give elongated clusters; Mahalanobis distance accounts for the spread in each direction, while Euclidean distance assumes round clusters. Option D is the opposite of what Mahalanobis distance does, and Euclidean distance works with any number of features.

Source: Paper p. 9 (journal p. 67), p. 14 (p. 72); Lecture 6 slides pp. 59–60, 68

Q4 · Main results

Which statement best summarises the main results of Biswas et al. (2015)?

Show answer

Answer: A. Healthy people averaged 88% and 83%, stroke patients 70% and 66%, and LDA and SVM averaged only about 45–68%. Option D overstates it: accuracy varied widely (one patient reached only 40%) and with 4 + 4 people the authors call it a proof of concept.

Source: Paper p. 2 (journal p. 60, abstract), pp. 11–15 (pp. 69–73), Tables 2–11

Q5 · Realistic test data 2 points

Biswas et al. trained each person's model on many repeated arm movements in a controlled setting. They tested it on another day, while the person freely made a cup of tea. State and motivate briefly:
1) why they tested on the tea-making task instead of on more controlled repetitions;
2) whether you expect the accuracy to be higher or lower than when testing on controlled repetitions.

Show model answer

1) The goal is to recognise movements in real daily life (monitoring rehabilitation at home). The free tea-making task on another day looks like real use, so it shows whether the model still works outside the conditions it learned from (generalisation). 2) Lower. Free movements vary more in posture, speed and context and differ more from the training trials, so they are harder to recognise. Testing on repetitions that resemble the training data would give an over-optimistic accuracy.

How the points are earned

  • 1 pt Motivates the realistic test: it resembles daily-life use and checks that the model generalises
  • 1 pt Expects lower accuracy because free movements vary more than the training trials (lab testing is over-optimistic)

Source: Paper p. 4 (journal p. 62), p. 5 (p. 63), p. 10 (p. 68)

Lecture 6 – Classification and clustering

original research (prospective cohort study over one season)

Rommers et al. (2020) — A Machine Learning Approach to Assess Injury Risk in Elite Youth Football Players

Rommers, N., Rössler, R., Verhagen, E., Vandecasteele, F., Verstockt, S., Vaeyens, R., Lenoir, M., D'Hondt, E., & Witvrouw, E. (2020). A machine learning approach to assess injury risk in elite youth football players. Medicine & Science in Sports & Exercise, 52(8), 1745–1751. https://doi.org/10.1249/MSS.0000000000002305 · Open paper

In 734 elite youth footballers, a machine-learning model (XGBoost) used simple preseason tests of body build, coordination and fitness to predict who would get injured during the season. On players it had not seen, 85% of the players it flagged really got injured, and it caught 85% of the players who did get injured (precision and recall both 85%; the paper calls this '85% accuracy'). A similar model told overuse from acute injuries a little less well (78%).

Question

Can a machine-learning model use simple preseason tests to predict which elite youth players will get injured in the coming season? The tests covered body measurements (anthropometry), how mature the player is, motor coordination, physical fitness and years of football experience. Second question: can a similar model tell whether the injury will be an overuse injury (no single event caused it) or an acute injury (one clear event, such as a kick or a fall)?

Why it matters

Elite youth football has a high injury risk, but clubs lack the time and money for extensive screening of every player. A model built on field tests that clubs already run would help them focus prevention on the players at highest risk. Earlier studies with traditional statistics found no link between single test results and injury. Injuries usually have many causes at once, so a method that combines many variables and how they interact might reveal risk profiles that a single variable cannot.

Approach

This prospective study (players were tested first and then followed forward in time) included 734 boys from under-10 to under-15 teams (mean age 11.7 years) at seven Belgian premier-league academies, during the 2017–18 season. In August, trained testers measured: anthropometry and maturity (height, sitting height, leg length, weight, body fat, and predicted age at peak height velocity (PHV): the age at which a child grows fastest, so a higher value means a later-maturing player); motor coordination (KTK3, a children's coordination test with three tasks: jumping sideways, moving sideways and balancing backwards, plus a dribbling test); and physical fitness (jumps, sprints, the agility t-test, sit-and-reach flexibility, the Yo-Yo IR1 endurance shuttle run, curl-ups). That made 29 preseason variables. Medical staff registered each player's first injury and whether it was overuse or acute; coaches recorded training and match time. The authors built two classifiers (models that sort cases into categories): injured vs not injured, and overuse vs acute. This is supervised learning: the model learns from examples where the true answer (injured or not) is known. They used XGBoost (extreme gradient boosting): it builds many small decision trees (flowcharts of yes/no splits on the predictors) one after another, each new tree fixing the errors of the earlier ones, and combines them. The data were randomly split into 80% training data and 20% test data that the model never saw while learning. On the training data, cross-validation (repeatedly training on part of the data and checking on the rest) and grid search (trying every combination from a list of model settings, the hyperparameters, and keeping the best) tuned the model. Performance was reported as precision (of the players the model flagged as 'will be injured', the share who really got injured), recall or sensitivity (of the players who really got injured, the share the model flagged) and F1 (one score combining precision and recall; it is only high if both are high). SHAP summary plots (SHapley Additive exPlanations: a method that shows how much each predictor pushed each player's prediction up or down) showed which variables drove the predictions.

Findings

  • 368 of 734 players (50%) had at least one injury. Of these first injuries, 173 were overuse and 195 acute (47% vs 53%).
  • Injury model: precision, recall and F1 of 84%, 83% and 83% on the training data, and 85%, 85% and 85% on the test data (147 players). So it worked just as well on players it had not seen.
  • Overuse vs acute model: 82%, 82% and 81% on the training data, and 78% for all three on the test data (74 injuries), a little lower than the injury model.
  • The five most important predictors of injury (from SHAP) were a higher predicted age at PHV, taller body height, longer legs, lower body-fat percentage and standing broad jump performance. These are mostly body-size and maturity measures, which the authors link to injury rates rising with age. Better sit-and-reach flexibility raised the predicted risk slightly.
  • For overuse vs acute: a lower predicted age at PHV, higher sitting height, a slower t-test and a lower moving-sideways (KTK3) score pointed towards overuse injuries.
  • Conclusion: combined in a machine-learning model, simple preseason tests can flag players at high risk and the likely injury type, so academies can spend their limited resources on those players.

Limitations

  • Authors: only each player's first injury was analysed, because test results may change after an injury. Repeated injuries are therefore ignored.
  • Authors: players were tested only once, before the season, although body build and fitness change during a season of growth and training. The authors suggest retesting every few months.
  • Authors: the model was tested on a random 20% of the same group in the same season. It still needs testing on a different group, for example players from other countries. Own inference: performance at new clubs or in new seasons may be lower.
  • Own inference: SHAP shows which links the model uses, not what causes injury. The top predictors mostly reflect age and maturity (older, taller players get injured more).

Link to the lectures

The lecture's worked example of a supervised classifier, slides pp. 37–54.

This is the standard supervised-classification example of Lecture 6: labelled data from a group followed over time, and a target with two classes (injured yes/no; then overuse vs acute). It follows the lecture workflow: a set of predictors (29 preseason variables), a tree-based ensemble method (XGBoost, boosting), an 80/20 split into training and test data with cross-validation and grid search on the training part, and evaluation with metrics from the confusion matrix (precision, recall, F1). Comparing training and test scores links to overfitting (a model that learns noise in its training data does clearly worse on new data) and the bias–variance trade-off; SHAP links to model interpretation and feature importance. About half the players got injured, so the classes are nearly balanced and the scores are not inflated by one large majority class. Paper wording note: the abstract calls the F1 score 'accuracy' ('85% accuracy (f1 score)'); in the course, accuracy and F1 are different metrics.

Remember

  • Aim: predict injury (yes/no) during one season from preseason tests in elite youth football; a second model separates overuse from acute injuries.
  • Supervised classification with two classes, using XGBoost: many small decision trees built one after another (boosting).
  • 734 players (under-10 to under-15) from 7 Belgian academies over 1 season; about 50% got injured.
  • Random 80/20 split into training and test data; cross-validation and grid search on the training data to tune the model.
  • Test data: 85% precision, recall and F1 for injury; 78% for overuse vs acute.
  • SHAP was used to interpret the model; the top predictors were mostly body-size and maturity measures (age at PHV, height, leg length, fat %) plus standing broad jump.
  • Limitations: only the first injury, tests only before the season, no test on an independent group.

Practice

Q1 · Why the study matters

What practical problem did Rommers et al. (2020) want to help youth football academies with?

Show answer

Answer: B. The authors start from the high injury risk in elite youth football and the clubs' limited time and money for thorough screening; a model on existing preseason tests lets them target prevention. GPS load data (A) were not part of this study: all predictors came from one preseason test day.

Source: Paper p. 1 (journal p. 1745), Introduction; p. 6 (p. 1750), Practical applications

Q2 · XGBoost

Rommers et al. built their injury model with XGBoost (extreme gradient boosting). Which description of this method is correct?

Show answer

Answer: D. The paper describes boosting as combining a set of weak learners (simple models) to improve prediction, and the slides call XGBoost a decision-tree-based ensemble. Option B is tempting because it mentions risk groups, but XGBoost learns from the injury labels, so it is supervised, not clustering.

Source: Paper p. 3 (journal p. 1747), p. 4 (p. 1748); Lecture 6 slides p. 44

Q3 · Training vs test performance

For the injured vs not-injured model, Rommers et al. report precision, recall and F1 of 84%, 83% and 83% on the training data (80% of players) and 85%, 85% and 85% on the test data (20% of players the model never saw). What does this comparison mainly tell you?

Show answer

Answer: A. Overfitting shows up as clearly better scores on training data than on test data; here the test scores are similar (even slightly higher), so the model generalised to new players from the same group. Option C is wrong: tuning was done with cross-validation and grid search on the training data only, and the test data were used once at the end.

Source: Paper p. 3 (journal p. 1747), p. 4 (p. 1748); Lecture 6 slides p. 9, 47

Q4 · Main predictors

Which preseason measures mattered most for predicting injury in Rommers et al. (2020)?

Show answer

Answer: C. The five most important predictors were predicted age at PHV, body height, leg length, body-fat percentage and standing broad jump; the authors link this to injury rates rising with age and maturity. Options A and B are impossible: all predictors came from the preseason test day, and neither GPS load nor injury history was measured.

Source: Paper p. 4 (journal p. 1748), p. 5 (p. 1749, Fig. 1A); Lecture 6 slides p. 47

Q5 · Why machine learning worked here 3 points

Two earlier studies with traditional statistics found no link between preseason test results and injury in youth football. Rommers et al. (2020) used similar tests in an XGBoost model and predicted injury in unseen players with 85% precision and recall. State and motivate briefly:
1) why a machine-learning model could succeed where the earlier studies did not;
2) one limitation that restricts how far you can trust the model in practice.

Show model answer

1) Injuries have many causes at once, so a single test result rarely has a clear link with injury. A tree-based ensemble such as XGBoost combines many variables (29 preseason measures) and their interactions, and is built to classify as accurately as possible instead of testing one risk factor at a time. So it can find risk profiles. 2) Any one: only the first injury per player was used; players were tested only once, before the season, while growth and training change their results; the model was tested on players from the same academies and season, so it still needs testing on an independent group (e.g. other countries).

How the points are earned

  • 1 pt Injuries are multifactorial: one variable alone is not enough to estimate risk
  • 1 pt The ML model combines many variables and their interactions (risk profile) instead of testing them one by one
  • 1 pt Names and explains one limitation (first injury only / tested only preseason / no independent validation)

Source: Paper p. 1 (journal p. 1745), p. 4 (p. 1748), p. 6 (p. 1750)

Lecture 6 – Classification and clustering

original research (observational, cross-sectional)

van der Zwaard et al. (2019) — Anthropometric Clusters of Competitive Cyclists and Their Sprint and Endurance Performance

van der Zwaard, S., de Ruiter, C. J., Jaspers, R. T., & de Koning, J. J. (2019). Anthropometric clusters of competitive cyclists and their sprint and endurance performance. Frontiers in Physiology, 10, 1276. https://doi.org/10.3389/fphys.2019.01276 · Open paper

The authors let a k-means clustering algorithm sort 24 competitive cyclists into groups by body build alone, without telling it their discipline. It found three groups: a muscular group that held all the sprinters and sprinted best, and a short and a tall slim-muscular group that mixed pursuit and road cyclists and had the best endurance. So the idea that cyclists pick the discipline that suits their build was only partly confirmed.

Question

If you group cyclists only by their own body measurements (anthropometry), using several aspects of build at once and without their discipline labels, do the groups match the disciplines they compete in (sprint, pursuit, road)? The authors also asked whether the groups differ in sprint and endurance performance, and how body build relates to performance.

Why it matters

Athletes are thought to specialise in the discipline that suits their body build. But body build is usually reported as the average of each discipline, and such group averages can hide individuals with a different build. Grouping athletes directly on their own body measures gives an unbiased test of whether build matches discipline, and helps to judge an athlete's performance in light of his build.

Approach

24 male cyclists (sprint, pursuit and road; national to Olympic level) visited the lab three times. Body measurements followed the ISAK protocol (an international standard for anthropometry). The clustering used nine variables, three for each aspect of build: body shape (the Heath–Carter somatotype: endomorphy = fatness, mesomorphy = muscularity, ectomorphy = slenderness), body size (height, weight, body surface area) and body composition (sum of eight skinfolds, body-fat %, skeletal-muscle-mass %). All variables were first turned into z-scores (value minus the mean, divided by the standard deviation), so each variable has the same scale. Then k-means clustering was run in R (Hartigan–Wong version). k-means is unsupervised learning: it gets no outcome or labels. It places k cluster centres, assigns each cyclist to the nearest centre by Euclidean (straight-line) distance, moves each centre to the mean of its cyclists, and repeats until nothing changes. The goal is to make the total squared distance between cyclists and their own cluster centre as small as possible. The number of clusters, k = 3, was chosen with the elbow criterion (plot the total within-cluster spread against k and pick the point where the curve bends), BIC (Bayesian information criterion, from the mclust package: a score that balances fit against complexity) and the NbClust package (many indices that each suggest a best k). The algorithm used 25 random starting points, and 1000 repeated runs checked that the result was stable. Afterwards, sprint performance (Wingate 1-s peak power, squat-jump power) and endurance performance (15-km time-trial power, VO2peak: the highest oxygen uptake in an all-out test), both per kg body mass, were compared between clusters (ANOVA or Kruskal–Wallis tests, which check whether groups differ) and correlated with body measures.

Findings

  • Three stable clusters: mesomorphic (muscular; n = 6), short meso-ectomorphic (muscular but slim; n = 9) and tall meso-ectomorphic (n = 9).
  • All six sprinters were in the mesomorphic cluster (heavier, larger girths, less lean). Pursuit and road cyclists were spread evenly over the two meso-ectomorphic clusters (4 pursuit + 5 road in each), which differed mainly in body size.
  • The mesomorphic cluster had higher sprint performance (Wingate peak power, squat-jump power; p < 0.05) and lower endurance performance (time-trial power, VO2peak; p < 0.001) than both meso-ectomorphic clusters.
  • Across all cyclists, better endurance went with a lean, slender build with small girths and a small frontal area (the body area facing the wind). Better sprint went with larger skinfolds and girths and a small frontal area per kg body mass.
  • The idea of separate clusters for sprint, pursuit and road was only partly confirmed. The clustering separated sprint-type from endurance-type cyclists, and short from tall endurance cyclists (matching the build of all-terrain vs flat-terrain road cyclists), but not pursuit from road cyclists.

Limitations

  • Authors: small sample (24 male cyclists; clusters of 6, 9 and 9) from only three disciplines. They call for a larger sample covering all cycling specialities.
  • Authors: k must be chosen in advance, and the result depends on which variables go in. They handled this with validity criteria (elbow, BIC, NbClust) and the same number of variables (three) for each aspect of build.
  • Own inference: the study is observational and measured each cyclist once (cross-sectional). It shows that build and discipline go together, but not whether athletes chose a discipline because of their build or their build changed through training.
  • Own inference: the clusters were formed and then compared in the same small data set. Repeated runs showed the algorithm gives the same answer, but no independent group of cyclists checked that the clusters reappear.

Link to the lectures

The lecture's worked example of k-means clustering, slides pp. 76–96.

A clear example of unsupervised learning from Lecture 6: the algorithm gets no outcome or labels, and the discipline labels are only used afterwards to interpret the clusters. Supervised learning, by contrast, learns from examples where the answer is known (such as each cyclist's discipline). It shows the k-means steps from the lecture: minimise the within-cluster sum of squared Euclidean distances, repeat 'assign to nearest centre' and 'update the centres', and use several random starts. It also shows the practical steps: z-scores so all variables count equally, choosing k with an elbow (scree) plot and NbClust, and checking the result with a clusplot (are the clusters round, of similar size and not overlapping?). The clusplot shows the clusters on two dimensions from a dimension reduction that keeps 85% of the variation, which links to PCA (principal component analysis: combining many variables into a few new ones that keep most of the information). Contrast with van der Zwaard et al. (2018) in Lecture 3, which explains measured performance with regression (supervised).

Remember

  • Unsupervised learning: k-means clustering of 24 competitive cyclists on 9 body measures (shape, size and composition).
  • Discipline labels (sprint, pursuit, road) were NOT used to make the clusters, only to check them afterwards.
  • All variables were turned into z-scores first, so each has the same scale and weight in the Euclidean distance.
  • k = 3 was chosen with the elbow plot, BIC (mclust) and NbClust; 25 random starts; the same clusters in 1000 repeated runs.
  • Result: the muscular (mesomorphic) cluster held all sprinters and sprinted best; the short and tall slim-muscular (meso-ectomorphic) clusters mixed pursuit and road cyclists and had the best endurance.
  • Specialisation by build only partly confirmed: sprint vs endurance types were separated, pursuit vs road cyclists were not.
  • k-means assumes round (spherical) clusters of similar size; the authors checked this with a clusplot.

Practice

Q1 · Research question

What did van der Zwaard et al. (2019) want to find out?

Show answer

Answer: B. The study grouped cyclists on anthropometry alone and then checked whether the groups matched sprint, pursuit and road, and how the groups performed. Option A is tempting, but predicting the discipline would be supervised classification; the authors deliberately kept the labels out. Option D describes the 2018 study by the same first author.

Source: Paper p. 1 (abstract), p. 2 (Introduction); Lecture 6 slides p. 77

Q2 · Standardising before k-means

Before running k-means, van der Zwaard et al. turned all body measures into z-scores. Why?

Show answer

Answer: D. k-means assigns each cyclist to the nearest cluster centre by Euclidean distance, so a variable on a large scale would otherwise outweigh the rest; the authors note that 'no single feature is more important than another'. Option B is tempting because the paper shows a 2-D plot, but that came from a separate dimension reduction for the clusplot.

Source: Paper p. 4, p. 9; Lecture 6 slides pp. 66–67

Q3 · Choosing k

k-means needs the number of clusters (k) to be set in advance. How did van der Zwaard et al. arrive at k = 3?

Show answer

Answer: A. The paper chose k with the elbow criterion, BIC (mclust) and NbClust. Option B is tempting because k happens to equal the number of disciplines, but using the labels to pick k (B or D) would undo the label-free approach, and k-means never chooses k itself.

Source: Paper p. 4, p. 9; Lecture 6 slides pp. 62–64, 82–83

Q4 · Main result

What was the main clustering result in van der Zwaard et al. (2019)?

Show answer

Answer: C. All six sprinters landed in the mesomorphic cluster, which sprinted better and had lower endurance. Option A is tempting because there were three clusters and three disciplines, but pursuit and road cyclists (4 + 5 in each) were split by body size, not by discipline, so specialisation by build was only partly confirmed.

Source: Paper p. 1 (abstract), pp. 4–5 (Table 1), p. 6, p. 7 (Fig. 3); Lecture 6 slides pp. 89–91, 95

Q5 · Why clustering, and what it showed 3 points

Van der Zwaard et al. (2019) knew each cyclist's discipline (sprint, pursuit or road) but still grouped the 24 cyclists with k-means clustering on body measures only. All sprinters ended up in one cluster; pursuit and road cyclists were mixed over two clusters that differed in body size. State and motivate briefly:
1) why the authors used unsupervised clustering instead of a supervised classifier;
2) what the result means for the idea that cyclists specialise in the discipline that suits their build.

Show model answer

1) A supervised classifier needs the discipline labels and learns to separate those given groups, so it cannot show whether such groups arise from body build by themselves. Clustering forms groups from body build alone, without labels, so comparing clusters with disciplines afterwards is an unbiased test of specialisation by build. 2) Only partly confirmed: build separates sprint-type from endurance-type cyclists, but not pursuit from road cyclists. The endurance cyclists split by size (short vs tall), which matches the build of all-terrain vs flat-terrain road cyclists.

How the points are earned

  • 1 pt Supervised learning needs labels / learns the predefined groups; clustering uses no labels
  • 1 pt Clusters formed on body build alone can then be compared with disciplines: an unbiased test
  • 1 pt Interprets the result: partly confirmed; sprint vs endurance separated, pursuit vs road not (split by size)

Source: Paper p. 1 (abstract), p. 9 (Data science), p. 10 (Conclusion); Lecture 6 slides pp. 77, 89–91, 95

Lecture 7 – Feature operations

original research

Mundt et al. (2020) — Prediction of lower limb joint angles and moments during gait using artificial neural networks

Mundt, M., Thomsen, W., Witter, T., Koeppe, A., David, S., Bamer, F., Potthast, W., & Markert, B. (2020). Prediction of lower limb joint angles and moments during gait using artificial neural networks. Medical & Biological Engineering & Computing, 58, 211–225. https://doi.org/10.1007/s11517-019-02061-3 · Open paper

The authors simulated wearable motion-sensor (IMU) signals from an existing lab database of walking trials and used them to train two kinds of neural network to predict hip, knee and ankle joint angles and joint moments in 3D. Both networks worked well, and the simpler feedforward network was slightly more accurate than the LSTM network.

Question

Can neural networks predict lower-body joint angles and joint moments during walking directly from IMU signals? An IMU (inertial measurement unit) is a small wearable sensor; here only its accelerometer (measures acceleration) and gyroscope (measures rotation speed) were used. Joint angles describe the motion (kinematics); joint moments are the turning forces around a joint (kinetics). The goal was to make IMU-based gait analysis independent of the magnetometer (a built-in compass) and of calibration poses. A second question: how does a standard feedforward network compare with a recurrent LSTM network on this task?

Why it matters

Gait analysis in the lab uses cameras and force plates (plates in the floor that measure the force of each step). This is expensive and only works in a small recording space. IMUs work anywhere, but they measure joint angles only indirectly, depend on calibration that easily goes wrong and on magnetometers that are easily disturbed by local changes in the magnetic field, and they cannot measure joint moments at all. If a network can turn raw IMU signals into gold-standard joint angles and moments (gold standard = the best available measurement), full gait analysis becomes possible in clinics and in the field.

Approach

The authors reused a database of walking trials on level ground: people without movement disorders plus patients with a knee replacement, walking at their own chosen speed (0.8–2.0 m/s). The trials had been recorded with optical motion capture (cameras tracking reflective markers on the skin) and force plates. Joint angles were computed from the markers. Joint moments were computed with inverse dynamics (working backwards from the measured motion and floor forces to the forces at each joint) in the AnyBody software, and scaled to body weight and height. These angles and moments were the targets (the known answers). From the marker paths the authors calculated what virtual IMUs on the body would have measured (acceleration and rotation speed, no magnetometer). This gave model inputs that match the gold-standard targets exactly. They trained two kinds of artificial neural network (a model made of layers of connected simple units that learns a flexible input-to-output function from examples): a feedforward network (FFNN), which sees the whole gait cycle or stance phase at once, and an LSTM (long short-term memory) network, a recurrent network that reads the signal step by step and keeps an internal memory of what came before. Training used standard settings: mean squared error as the quantity to minimise, the Adam training algorithm, early stopping (stop training when the error on validation data stops improving) and dropout (randomly switching off units during training); both reduce overfitting (learning noise in the training data so the model fails on new data). The data were split into training, validation and test sets by participant, so no person appeared in more than one set. They started from a baseline (IMU data, right steps only, trained 10 times) and tested changes: adding all or some body measurements, using both left and right steps, and data augmentation (creating extra training examples by randomly rotating the virtual sensor). A change was kept only if it beat the baseline by more than two standard deviations. Accuracy was reported as the correlation coefficient r (how closely the predicted curve follows the shape of the true curve) plus RMSE (root mean square error: typical size of the error; in degrees for angles) or nRMSE (RMSE divided by the range of the curve, in %, for moments). They reported both because r alone ignores an offset: a curve shifted up by 5° can still have r = 1.

Findings

  • Both networks predicted joint angles and moments well. The abstract reports mean correlations above 0.98 in the sagittal plane (forward-backward movement seen from the side) and above 0.80 in the smaller planes (frontal: sideways; transverse: rotation). The Discussion reports RMSE below 4° (angles) and nRMSE below 25% (moments) for all joints and planes. Caveat: a few single values are lower, for example knee frontal-plane angle r = 0.79 (FFNN) and 0.68 (LSTM) in Table 4.
  • The feedforward network beat the LSTM for both angles and moments in most joints and planes (baseline joint-angle RMSE 1.60° vs 2.15°; baseline moment nRMSE 12.2% vs 14.8%). The authors' explanation: the FFNN sees all time steps at once, while the LSTM can only carry information forward from earlier time steps.
  • Sagittal-plane motion was predicted best. For joint angles the transverse plane was worst; for joint moments the knee rotation moment and the frontal-plane ankle moment were worst.
  • For joint moments, using both left and right steps plus data augmentation (3 times as much data) improved the models. For joint angles, only adding some body measurements (height, weight, segment lengths, pelvis width) helped.
  • Conclusion: the FFNN suits analysis of fully recorded data afterwards; the LSTM is the more natural choice for real-time use and recordings of any length. A convolutional network with a small delay might combine both strengths.

Limitations

  • Authors: the IMU signals were simulated, so they lack the wobble of soft tissue (skin and muscle) that real body-worn sensors experience. The models were not tested on real sensor data.
  • Authors: all virtual sensors had exactly the same position and orientation. Whether the models still work when sensors are placed differently still needs testing.
  • Authors: outliers came from the manual split of the data (test patterns not present in training), the small share of knee-replacement patients, the weakness of camera systems for knee motion outside the sagittal plane, and errors in the inverse-dynamics calculations.
  • Own inference: only level walking at chosen speed was studied, so the models may not work for running, stairs or other tasks.

Link to the lectures

Neural-network part: example of a feedforward vs a recurrent (LSTM) network; also shows scaling, adding features and data augmentation.

Lecture 7 contrasts simple models with neural networks: networks fit complex functions and capture relations between inputs, but are harder to interpret, need more data and need more care with bias and variance (underfitting vs overfitting) (Slides p. 18). Mundt et al. show each point: they create more data by simulation and augmentation, limit overfitting with early stopping, dropout and a split by participant, and compare a non-recurrent network (FFNN) with a recurrent one (LSTM). The paper also shows feature operations from Part I of the lecture: joint moments are scaled to body weight and height, and body-measurement features are added and kept only if they improve accuracy. Note: the Lecture 7 slides do not discuss this paper directly; they give only a general neural-network slide (p. 18) and a video slide (p. 19).

Remember

  • Aim: predict 3D hip, knee and ankle joint angles and moments during walking from IMU data (acceleration + rotation speed only) with neural networks, without magnetometers or calibration poses.
  • IMU inputs were simulated from an existing camera-based lab database; targets came from camera kinematics and inverse dynamics (gold standard).
  • Supervised regression on time-series data: feedforward network (FFNN) vs recurrent LSTM.
  • The FFNN beat the LSTM in most joints and planes; both had r > 0.98 in the sagittal plane and > 0.80 in the other planes (abstract).
  • Training, validation and test sets were split by participant to prevent leakage; early stopping and dropout limited overfitting.
  • r alone ignores offsets, so RMSE (angles) and nRMSE (moments) were also reported.
  • Main limitation: simulated data lack soft-tissue wobble, so real IMU data still need testing.

Practice

Q1 · Type of ML problem

Mundt et al. (2020) trained neural networks to predict hip, knee and ankle joint angles and moments (continuous curves) from IMU signals. For every input, the true angles and moments from the lab were known. What type of machine learning problem is this?

Show answer

Answer: C. The outputs are continuous numbers (degrees, scaled moments) and every input has a known true value, so it is supervised regression. Classification would need categories (for example patient vs healthy), which the networks did not output.

Source: Paper p. 211 (PDF p. 1, Abstract); pp. 212–213 (PDF pp. 2–3, Sect. 1 and 2.1)

Q2 · Why simulate IMU data

Instead of recording new IMU data, Mundt et al. calculated what IMUs would have measured from an existing camera-based lab database. What was the main advantage?

Show answer

Answer: A. Camera-based lab data are plentiful and come with gold-standard targets, so simulation gives many perfectly matched input–target pairs. Option B is the reverse of the authors' own limitation: simulated signals lack the soft-tissue wobble of real sensors.

Source: Paper p. 211 (PDF p. 1, Introduction); p. 219 (PDF p. 9, Discussion)

Q3 · Main finding: FFNN vs LSTM

A feedforward network (FFNN) sees the whole step at once; an LSTM network reads the signal step by step with a memory of what came before. What did Mundt et al. (2020) find?

Show answer

Answer: B. The FFNN beat the LSTM for both angles and moments, with r > 0.98 in the sagittal plane. Option A sounds logical but is wrong: the authors argue the FFNN wins because it uses all time steps at once, while the LSTM only carries information from earlier steps.

Source: Paper p. 211 (PDF p. 1, Abstract); pp. 217, 219, 222 (PDF pp. 7, 9, 12; Tables 3–4, Discussion)

Q4 · Testing on new people

Mundt et al. wanted the reported accuracy to hold for people the network had never seen. How did they split their data?

Show answer

Answer: D. Splitting by person prevents leakage: the test set contains only new people. Option A is exactly the leakage they avoided: other steps of the same person in the training set would make the test results look too good.

Source: Paper p. 212 (PDF p. 2, Sect. 2.1)

Q5 · FFNN vs LSTM 2 points

Mundt et al. found the feedforward network (FFNN) more accurate than the LSTM network, yet they do not recommend the FFNN for every use. State and motivate briefly:
1) why the FFNN probably performed better;
2) in which situation the LSTM is the more natural choice.

Show model answer

1) The FFNN gets all time steps of the step (gait cycle or stance phase) at once, so it can use both earlier and later information. The LSTM only carries information forward from earlier time steps, which especially hurts the first time steps. 2) Real-time use (for example live feedback) and recordings of any length: the LSTM processes the signal step by step as it comes in. The FFNN suits data that are fully recorded and analysed afterwards.

How the points are earned

  • 1 pt FFNN sees all time steps at once (past and future); LSTM only uses earlier time steps
  • 1 pt LSTM fits real-time use / recordings of any length because it processes the signal step by step

Source: Paper pp. 219, 222 (PDF pp. 9, 12; Discussion and Conclusion)

Q6 · Why it matters and main limitation 2 points

A clinic wants to use the networks of Mundt et al. to analyse patients' walking with IMUs instead of in the gait lab. State and motivate briefly:
1) one advantage of IMU-based prediction over lab gait analysis;
2) one limitation of the study's data that must be solved before clinical use.

Show model answer

1) Lab gait analysis with cameras and force plates is expensive and limited to a small recording space. IMUs work outside the lab (clinic, field, daily life), and with the network they also give joint moments, which IMUs cannot measure directly. 2) The IMU signals were simulated from camera data, so they lack the soft-tissue wobble of real sensors, and all virtual sensors had the same position. The models must first be tested on real IMU data (and other sensor positions, tasks and patient groups).

How the points are earned

  • 1 pt Advantage: lab is expensive and limited to a small space; IMUs work outside the lab / gives joint moments that IMUs cannot measure
  • 1 pt Limitation: simulated data (no soft-tissue movement, same sensor position), so real IMU data must be tested

Source: Paper pp. 211–212 (PDF pp. 1–2, Introduction); pp. 219, 222 (PDF pp. 9, 12; Discussion)

Lecture 8 – Epidemiological data

Systematic literature review (structured review with content analysis)

Galetsi et al. (2020) — Big data analytics in health sector: Theoretical framework, techniques and prospects

Galetsi, P., Katsaliaki, K., & Kumar, S. (2020). Big data analytics in health sector: Theoretical framework, techniques and prospects. International Journal of Information Management, 50, 206–216. https://doi.org/10.1016/j.ijinfomgt.2019.05.003 · Open paper

A systematic review of 804 published papers (2000–2016) on big data analytics in healthcare. It sorts them by the type of health data used, the analysis techniques, the value created, the computer platforms that handle the data (mainly Hadoop) and the future directions the papers propose.

Question

How is big health data used and analysed to create value for patients, analysts and managers in healthcare? The authors use the resource-based view, a business theory that says an organisation gains value by combining its resources (here: data) with its capabilities (here: analysis). They map which data types are used, which analysis techniques are applied, which values are created, which platforms and tools handle big health data, and which future directions are proposed.

Why it matters

Healthcare produces huge amounts of data: electronic health records (EHR: the digital patient file), lab systems, medical images, home sensors, billing data and social media posts. Big data are defined by the 4 Vs: volume (very large amounts), velocity (data arrive fast, often in real time), variety (many formats: structured, unstructured and semi-structured) and veracity (whether the data are reliable). The number of papers on health analytics is growing fast, so doctors, policy makers and researchers need an organised overview that links technical progress to the value it creates.

Approach

A systematic literature review: a review that finds all papers on a topic with a fixed, repeatable search and fixed rules for which papers count, done in two steps (following Tranfield et al., 2003, and Dubey et al., 2017). Step 1: they searched two databases of scientific papers, Web of Science and Scopus (2000–2016), for 'business intelligence', 'analytics' and 'big data' combined with health, medical and clinical terms. They kept only English-language articles and reviews about both health and big data analytics: 6817 hits → 3241 after removing duplicates → 1877 after reading titles and abstracts → 804 included after reading the full texts. Step 2: content analysis, meaning they read each paper and coded (labelled) it, using the text-coding software NVivo, by data type, technique, value created and future direction. A paper could get several labels. Then they counted how often each label occurred and made cross-tables, for example technique by data type and technique by value.

Findings

  • Data types: clinical data (for example EHRs and medical images) dominate, in 69.9% of papers (562 of 804). Next come patient behavior and sentiment data (from wearable sensors and social media; 16.5%), administrative and cost data (7.3%) and pharmaceutical research data (4.7%).
  • Techniques: modelling (42.8%) and machine learning (40.7%) were used most, followed by data mining (finding patterns in large data sets; 24.9%), visualisation (19%) and statistics (16.4%). Machine learning (algorithms that learn patterns from data instead of following fixed rules) was the most-used technique across almost all values and data types, so the authors give it its own section.
  • Values created: the most frequent was better diagnosis for more personalised healthcare (35.6%), then supporting or replacing professionals' decisions with automated algorithms (25.6%) and new business models, products and services (24.5%). The values are grouped by who benefits: patients (P), analysts (A) or management (M).
  • Platforms: Apache Hadoop is the most-used computing platform. It stores data spread over many computers (HDFS) and processes them in parallel (MapReduce). A main difficulty is that most health data are unstructured (for example the free text of doctors' notes, not neat tables), which makes them hard to process by computer.
  • Future directions are mostly technological; the most frequent was developing the presented approach further (16.5%). For machine learning in particular they name: unsupervised learning (finding groups without known labels) to describe subtypes of complex diseases more precisely, automated risk prediction to guide care, reinforcement learning (learning by trial and reward) to support care providers, the tension between accuracy and interpretability (how easy it is to understand why a model gives its answer), and training clinicians to judge ML-based tools.

Limitations

  • Time lag (stated by the authors): the reviewed papers end in 2016, and technology moves faster than publication, so recent advances are missing.
  • One author sorted the papers into categories in NVivo (with advice from the others), and the categories overlap (a paper can be counted several times), so the coding is partly subjective.
  • Own inference: it only counts how often things occur; it does not judge the quality or effectiveness of the included studies. The search terms (business intelligence / analytics / big data, English only) may also have missed relevant work.

Link to the lectures

Slide 44 'Health Data' takes its types of health data from this paper; full reference on slide 76.

It supplies Lecture 8's list of health-data sources (slide 44). The slide calls the paper's 'patient behaviour and sentiment data' 'real-time patient: wearable sensors'. Several of its values match course machine-learning ideas: identifying a patient's care risk uses supervised models (models that learn from examples where the outcome is known), such as logistic regression and regression trees, while segmenting populations uses clustering (unsupervised: finding groups without known labels). It also shows the accuracy-versus-interpretability tension. Structured versus unstructured data and storing and processing data spread over many computers (Hadoop) connect to the data acquisition and storage topics of Lecture 3. Canvas labels the file 'L9', but among the lecture slides only Lecture 8 cites it.

Remember

  • Systematic literature review (no new data): 804 papers, 2000–2016, searched in Web of Science and Scopus.
  • Framework: resource-based view: data (resource) + analysis techniques (capability) → value.
  • Big data 4 Vs: volume, velocity, variety, veracity.
  • Data types: clinical (≈70%, e.g. EHR) ≫ patient behavior/sentiment (wearables, social media) > administrative/cost > pharmaceutical research.
  • Most-used techniques: modelling and machine learning, then data mining, visualisation, statistics.
  • Top value: better diagnosis → more personalised healthcare; then automated decision support.
  • Hadoop/MapReduce is the main platform; most health data are unstructured; the authors' stated limitation is the time lag.

Practice

Q1 · Study type

What kind of study is Galetsi et al. (2020) on big data analytics in healthcare?

Show answer

Answer: C. The authors searched Web of Science and Scopus (2000–2016), narrowed 6817 hits down to 804 papers and coded those papers. The 804 are papers, not patients (option A), and no numerical results were pooled (option B).

Source: Paper p. 207–208 (Methodology)

Q2 · Health data types

According to Galetsi et al. (2020), which type of health data was used in the large majority (about 70%) of the reviewed big-data studies?

Show answer

Answer: A. Clinical data appeared in 69.9% of papers (562 of 804), with electronic health records the most common. Wearable and social-media data (option B) came second, at only 16.5%.

Source: Paper p. 208–209, Table 1; Lecture 8 Slides p. 44

Q3 · ML type for segmenting populations

Galetsi et al. (2020) list 'offering customised actions by segmenting populations' as a value of big data analytics: patients are divided into new groups, without labels given in advance, to offer more targeted services. Which type of machine learning fits this task?

Show answer

Answer: D. Finding groups without labels given in advance is unsupervised clustering, which the paper names for this value. Supervised classification (A) is tempting, but it needs known group labels to learn from; the paper uses supervised models such as logistic regression for a different value, identifying a patient's care risk.

Source: Paper p. 210–211 (Table 3 and text)

Q4 · Main value created

Which value of big data analytics did Galetsi et al. (2020) find most often in the reviewed health studies?

Show answer

Answer: B. 'Better diagnosis for provision of more personalised healthcare' appeared in 35.6% of papers, the most of the 10 values. Protecting privacy (A) was the least frequent, at 5.1%.

Source: Paper p. 210, Table 3

Q5 · Approach and limitation 3 points

Galetsi et al. (2020) wanted an overview of how big data analytics creates value in healthcare. State and motivate briefly:
1) the general approach they used to collect and analyse the studies;
2) one limitation of this approach.

Show model answer

1) A systematic literature review. They searched two paper databases (Web of Science and Scopus, 2000–2016) with fixed search terms (big data / analytics / business intelligence + health, medical, clinical) and fixed inclusion rules (English articles and reviews on health and big data analytics), and screened 6817 hits down to 804 papers. They then coded each paper by data type, analysis technique, value created and future direction, and counted how often each occurred. 2) Any one of: time lag: only papers up to 2016, while technology moves fast (stated by the authors); the coding was done mainly by one author and the categories overlap, so it is partly subjective; it counts studies but does not judge their quality or effectiveness.

How the points are earned

  • 1 pt Systematic literature review: structured database search with fixed terms and inclusion rules, screening to 804 papers
  • 1 pt Papers coded/classified by data type, technique and value, then counted (content analysis)
  • 1 pt One limitation with explanation: time lag (up to 2016), subjective coding by one author, no quality appraisal

Source: Paper p. 207–208 (Methodology), p. 214 (limitation)

Lecture 9 – User-generated data

original research (4-page conference paper)

Altini & Amft (2018) — Estimating Running Performance Combining Non-invasive Physiological Measurements and Training Patterns in Free-Living

Altini, M., & Amft, O. (2018). Estimating running performance combining non-invasive physiological measurements and training patterns in free-living. In Proceedings of the 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 2845–2848. https://doi.org/10.1109/EMBC.2018.8512924 · Open paper

The authors used two years of everyday-life data from 2113 users of a running app to estimate each runner's best 10 km time. A multiple linear regression model got within about 2.6 minutes (about 4%), and the estimate improved each time they added a group of predictors: body data, morning heart data, amount of training, training intensity pattern and past performance.

Question

Can you estimate how fast someone runs 10 km using only data collected in daily life, without any lab tests? The data were morning heart measurements from the HRV4Training phone app, details the users typed in themselves, and running workouts copied in automatically from other apps. The authors also asked which groups of predictors matter most, and how many workouts you need for a good estimate.

Why it matters

A good estimate of running performance helps runners pace their races and fit their training plan to their ability, which may also lower injury risk. Earlier studies had three problems. They used small groups of similar people (for example only men). They needed lab measures such as VO2max (the most oxygen the body can use per minute) or lactate threshold, which ordinary runners cannot measure. And they often did not test their models on new people. Data from apps and wearables offer much larger, more realistic samples.

Approach

Users of the HRV4Training app agreed to share their data. Each morning they measured resting heart rate (HR) and heart rate variability (HRV: how much the time between heartbeats varies; higher HRV usually means a more rested body) with the phone camera, which had been checked against an ECG (the clinical heart monitor), or with a chest strap. Their running app (Strava or TrainingPeaks) was linked to HRV4Training through an API (a connection that lets one program fetch data from another), so every workout with GPS and heart rate came in automatically. In 2016–2017, 2113 users qualified: they trained with a heart-rate monitor, had linked workouts, had run at least one 10 km and had at least one month of morning measurements. The outcome (the 'true answer' the model learns from) was each user's fastest 10 km found in their own workouts. The predictors came from the 3 months before that run and were added in groups, one step at a time: (1) anthropometrics: body data such as BMI (weight relative to height), age and gender; (2) resting physiology: morning HR and rMSSD (the standard HRV number); (3) training volume and speed: how far and how fast they ran; (4) training physiology: speed divided by heart rate during runs, a fitness marker; (5) training polarization: how much of the training was clearly easy or clearly hard instead of moderate, measured as the share of workouts more than 5% faster or slower than the user's own average, and the share at moderate heart rate; (6) previous performance. The authors measured polarization relative to each user's own average, not with maximum, minimum or range, because single wrong readings from everyday sensors would distort those. The model was multiple linear regression (predicts a number as a weighted sum of the predictors). They chose it because its coefficients (the weights) are easy to read: each weight shows how strongly a predictor goes with a faster or slower time, and in which direction. This is supervised learning: the model learns from examples where the true answer is known. They tested it with 10-fold cross-validation: split the users into 10 parts, train on 9 parts and test on the left-out part, and repeat 10 times. Only users with at least 20 workouts were used to train, but all users were tested. Accuracy was reported as RMSE (root mean square error: the typical size of the error, here in minutes), mean percentage error, r (correlation between estimated and true times) and R² (the share of the differences between runners that the model explains, from 0 to 1).

Findings

  • The error went down with every group of predictors added: RMSE 6.27 min with body data only, 6.07 with morning HR and HRV, 4.04 with training volume and speed, 3.96 with training physiology, 3.64 with training polarization and 2.68 min with previous performance (Table II). The biggest single drop came from adding training volume and speed.
  • The best model (all predictors, including previous performance) had an RMSE of 2.6 min in the abstract (2.68 min in Table II), about 4% error, r = 0.93 and R² = 0.87. That is 58% better than using body data alone.
  • About 15 workouts were enough to compute useful training predictors and get the error close to its minimum (Fig. 3).
  • Direction of the coefficients: younger age, lower BMI, lower resting HR, higher HRV, longer and faster workouts, more workouts clearly faster or slower than one's own average, and less time at moderate heart rate all went with a faster 10 km. In short, more polarized training (mostly easy plus some hard, little moderate) went with better performance.
  • These links agree with earlier small lab studies, but here they were found without any lab testing. So recreational runners could use such a model to estimate their performance and adjust their training.

Limitations

  • Authors: only variables that are easy to get in daily life were used. Running technique (biomechanics) and running power could be added, and the results should be checked in a lab study.
  • Authors: the model estimates performance at one moment. Whether it can follow changes in performance over time is future work.
  • Own inference: the data are observational (nobody was told how to train), so the links show association, not cause. For example, polarized training going with faster times does not prove that changing to polarized training makes you faster.
  • Own inference: the users chose themselves (app users who own a heart-rate monitor; 1891 of the 2113 were men), and the 'true' outcome was the fastest 10 km found in training data, not an official race or a standard test. Both limit how far the results apply to other runners and how trustworthy the outcome is.

Link to the lectures

Guest lecture by Marco Altini; his example of the reference-data problem, Slides pp. 99–113.

In Lecture 9, Marco Altini uses this paper to show the reference-data problem of user-generated data: the classic study brings a few people to the lab for a treadmill time trial; here you instead take years of workouts from apps like Strava, use each person's best 10 km as the outcome, build predictors from their training and estimate performance (Slides pp. 99–113). It also shows regression from Lecture 5 (multiple linear regression, RMSE, R², cross-validation) and feature engineering from Lecture 7 (turning raw workouts into predictors such as volume, speed-to-HR ratio and polarization). Measuring predictors relative to each user's own average links to the lecture's points on noisy data and quality control (Slides pp. 81–95). Slide vs paper: Slide 113 says 'N = 2100, RMSE = 2 minutes (4%)'; the paper reports 2113 users and an RMSE of 2.6 min (abstract) or 2.68 min (Table II).

Remember

  • Aim: estimate a runner's best 10 km time from everyday app and wearable data, without lab tests.
  • Data: HRV4Training app plus Strava/TrainingPeaks workouts fetched through an API; 2113 users over 2 years. Outcome = each user's fastest 10 km in their own workouts.
  • Method: supervised regression with multiple linear regression (chosen because its weights are easy to interpret), tested with 10-fold cross-validation.
  • Accuracy improved as predictor groups were added: body data alone worst (RMSE ≈ 6.3 min); all predictors including past performance best (RMSE ≈ 2.6–2.7 min, about 4%, R² = 0.87).
  • Adding training volume and speed gave the biggest single improvement (RMSE 6.07 → 4.04 min).
  • About 15 workouts were enough for useful predictors.
  • More polarized training (fewer moderate-intensity workouts) went with a faster 10 km — an association, not proof of cause.

Practice

Q1 · Type of ML problem

Altini & Amft (2018) used body data, morning heart measurements and training data to estimate each user's best 10 km running time in minutes. For every user the true best time was known. What type of machine learning problem is this?

Show answer

Answer: A. The true outcome is known for every user (supervised), and it is a number on a continuous scale (minutes), so it is regression; the authors used multiple linear regression. Classification would only apply if runners were sorted into groups such as 'fast' and 'slow'.

Source: Paper PDF p. 2 (Abstract); PDF p. 4 (Sect. III-B)

Q2 · Reference data

A big problem with user-generated data is the missing reference outcome: the verified 'true answer' a model learns from. How did Altini & Amft get each user's true running performance?

Show answer

Answer: C. The fastest 10 km in each user's own workout data served as the outcome. The lecture contrasts this with option A, the classic lab approach, which works for a few people but not for thousands; lab measures (D) are exactly what the authors wanted to avoid.

Source: Paper PDF p. 3 (Sect. II-B); Slides pp. 100–103

Q3 · Main finding

Altini & Amft added groups of predictors to their model one at a time. What did they find?

Show answer

Answer: D. The RMSE fell from 6.27 min (body data only) to 2.68 min (all predictors), with the biggest drop when training volume and speed were added. Morning heart data (A) helped only a little (6.07 min). Exam note: Slide 113 rounds the best error to 'RMSE = 2 minutes (4%)'; the paper reports 2.6 min (abstract) and 2.68 min (Table II), so 'under 3 minutes, about 4%' fits both.

Source: Paper PDF p. 4 (Table II, Sect. IV-B); PDF p. 5 (Fig. 4); Slides pp. 108–111, 113

Q4 · Training polarization

Polarized training means most workouts are clearly easy or clearly hard, with little time at moderate intensity. What did Altini & Amft find about training pattern and 10 km time?

Show answer

Answer: B. More time at moderate heart rate went with slower times, while more clearly fast or slow workouts, longer runs and higher HRV went with faster times. Option A states the opposite of the finding. The data are observational, so this is an association, not proof that polarized training makes you faster.

Source: Paper PDF p. 5 (Sect. IV-C)

Q5 · Working with app data 3 points

Altini & Amft estimated runners' 10 km times from two years of app and wearable data from 2113 users, instead of testing a small group in the lab. State and motivate briefly:
1) one challenge of such user-generated data and how the authors dealt with it;
2) why they chose multiple linear regression as their model.

Show model answer

1) Any one challenge plus how it was handled. No reference outcome: nobody ran a lab test, so the authors used each user's fastest 10 km found in their own synced workouts. Or: noisy everyday sensor data: they measured training intensity relative to each user's own average instead of using maximum, minimum or range, which single wrong readings distort. Or: some users have few workouts: they trained only on users with at least 20 workouts and showed that about 15 workouts are enough. 2) Multiple linear regression has easy-to-read coefficients: each weight shows how strongly and in which direction a predictor goes with 10 km time (for example, polarized training goes with faster times). The authors wanted this insight, not just a prediction.

How the points are earned

  • 1 pt Names a challenge of user-generated data (no reference outcome, noisy data, too few workouts per user)
  • 1 pt Explains how the authors handled it (fastest 10 km from own workouts / deviation from own average / at least 20 workouts to train, 15 enough)
  • 1 pt Motivates the model choice: coefficients are easy to interpret (direction and size of each predictor's effect)

Source: Paper PDF pp. 3–4 (Sect. II-B, III-A to III-C, Fig. 3); Slides pp. 97–98, 112

Lecture 9 – User-generated data

perspective / magazine column (no new data)

Altini & Dunne (2021) — What's Next For Wearable Sensing?

Altini, M., & Dunne, L. (2021). What's next for wearable sensing? IEEE Pervasive Computing, 20(4), 87–92. https://doi.org/10.1109/MPRV.2021.3108377 · Open paper

Two editors of a wearable-computing magazine argue that wearables (sensors worn on the body, such as watches and rings) have moved from counting behavior, like steps, to tracking changes in the body's own signals, like resting heart rate. They give three steps for turning wearable data into new knowledge, list the main open problems (accurate data, where on the body a device can sit, battery power) and expect new sensors, devices that also act, and machine learning that forecasts each person's response.

Question

Where does wearable sensing stand after about ten years of fast commercial growth, and what comes next? The authors look at what wearables measure now, the shift from tracking behavior to tracking health, large-scale research with wearable data, the main open problems and the likely next steps.

Why it matters

Wearables are now mass-market products: the Apple Watch has sold more than 100 million units (up from 30 million in 2017) and Google bought Fitbit. They now make observational studies (studies that only watch people, without changing anything) with thousands of participants possible, including tracking infections during COVID-19. At the same time, people started to question how accurate and how biased (systematically wrong for some groups) the sensors are. Researchers therefore need a clear view of what wearable data can and cannot support.

Approach

This is a perspective column, not a study with new data: the authors summarize recent developments and a few chosen studies. They describe what is measured now: the standard phone sensors (GPS for location, accelerometers and gyroscopes for movement, the camera for heart measurements), blood pressure from PPG (photoplethysmography: an optical sensor that shines light into the skin and measures changes in blood flow) on the wrist or ear, earables (in-ear devices such as AirPods), and glucose measured in interstitial fluid (the fluid between cells under the skin) instead of blood. Measuring glucose from sweat has so far failed. They describe the shift from tracking behavior to tracking physiology (the body's own signals), and a three-step route to new knowledge, explained with a Fitbit heart-rate example and with the COVID-19 recovery data of Radin et al. (2021). Then they discuss the open problems (accurate data, wearability and access to the body, energy use) and look ahead (new sensors, devices that act as well as measure, and forecasting with machine learning).

Findings

  • The focus has moved from absolute measures of behavior (steps, energy expenditure) to relative changes in physiology: how a person's resting heart rate (HR), heart rate variability (HRV: how much the time between heartbeats varies) or blood glucose change compared with their own normal level, in response to stressors such as sickness, alcohol, exercise, diet and the menstrual cycle. Activity and location are still useful as context.
  • Three steps to new knowledge: (1) the device must measure accurately, shown by validation (comparing it with a reference system) and certification (official approval as a medical device, such as FDA approval or a CE mark); (2) it must reproduce known results from smaller clinical studies, for example the short-term heart-rate response to exercise or sickness; (3) only then can use at large scale reveal new relationships, for example COVID-19, long COVID and the return to one's normal physiology (Radin et al.).
  • Possibly the biggest challenge is the noise of everyday, unsupervised settings: sensors can be worn wrongly or fail (for example optical HR during exercise), and devices rarely report signal quality or how confident they are in a value. Skin color, motion artefacts (errors caused by movement) and sensor position also reduce accuracy: a watch sits on top of the wrist, but the large blood vessels run underneath, so the sensor must rely on tiny vessels (capillaries).
  • Other challenges: a device in one spot (the wrist) can measure only a limited set of things, while sensors built into clothing (e-textiles) are hard to develop, test and manufacture. Energy use and battery power also remain problems.
  • Next steps: new types of sensors, with more attention to bias and inclusion; closed-loop systems that act as well as measure (for example a glucose monitor that steers an insulin pump, or vagal nerve stimulation, a mild electrical stimulation at the ear meant to reduce stress); and machine-learning models that forecast an individual's response (for example tomorrow's HRV if you skip a glass of wine today) instead of only describing past changes.

Limitations

  • Own inference: this is an opinion piece, not a systematic review (a review that searches all studies on a topic in a fixed, repeatable way). The authors chose the examples themselves and analysed no new data.
  • Own inference: the author note says the first author founded HRV4Training and advises the wearable company Oura, so the view comes partly from the consumer-wearable industry; weigh the claims with that in mind.
  • Own inference: the forecasting vision is a hope, not a result. The authors themselves say that prediction from wearable data is still limited and that consumer estimates are not always accurate.

Link to the lectures

Its 'three steps to knowledge discovery' are the lecture's three key steps; its data-accuracy challenge matches the lecture's points on noisy wearable data.

Lecture 9 (Marco Altini's guest lecture) is built on the same ideas. User-generated data allow research at a larger scale and in real-life settings, but need three key steps: validate the technology ('garbage in, garbage out': bad data give bad results), use it to confirm what is known from the lab, then discover new relationships or build new products (Slides pp. 57–59). These are the paper's 'three steps to knowledge discovery'. The paper's accurate-data challenge matches the lecture's point that wearable data are very noisy and usually come without a signal-quality measure (Slides p. 81). Its COVID-19 recovery example (Fig. 2, reprinted from Radin et al.) links to Radin et al. (2021) and to the lecture's infection and sickness examples (Slides pp. 21, 69, 115–117). The slides do not show this paper itself.

Remember

  • A perspective column in IEEE Pervasive Computing (2021), not a study with new data.
  • Shift: from tracking behavior (steps, energy expenditure; absolute measures) to tracking health (relative changes in resting HR, HRV and glucose in response to stressors).
  • Three steps: validate accuracy → reproduce known lab or clinical results → use at large scale to discover new things.
  • Possibly the biggest challenge: the noise of everyday, unsupervised settings; devices rarely report signal quality.
  • Wrist optical HR is error-prone: the sensor sits on top of the wrist, but the large blood vessels run underneath, so it relies on capillaries.
  • Future: closed-loop devices that measure and act (e.g., glucose monitor + insulin pump) and machine learning that forecasts each person's response.

Practice

Q1 · From behavior to health

According to Altini & Dunne (2021), how has the focus of consumer wearables changed in recent years?

Show answer

Answer: B. The section 'From Monitoring Behavior to Monitoring Health' describes exactly this shift; behavior and location are still used, but as context for the body-signal changes. Option A states the shift in the wrong direction.

Source: Paper p. 88 (PDF p. 2)

Q2 · Three steps to knowledge discovery

Altini & Dunne describe three steps that must happen before wearable data can produce new knowledge. Which order is correct?

Show answer

Answer: D. The paper requires accurate technology first, then reproduction of known responses (for example the heart-rate change after exercise or during sickness), and only then large-scale discovery (for example COVID-19 recovery). Option C has the right ingredients but checks accuracy too late: without an accurate device you cannot trust the reproduced results.

Source: Paper pp. 88–89 (PDF pp. 2–3); Slides pp. 57–59

Q3 · Main challenge: data accuracy

What do Altini & Dunne name as possibly the biggest challenge for consumer wearables?

Show answer

Answer: A. The 'Accurate Data' section opens with this claim and adds that devices rarely report signal-quality problems, even though the sensor has that information. Option C contradicts the paper, which says the main smartphone sensors have stayed the same for years.

Source: Paper p. 90 (PDF p. 4); p. 87 (PDF p. 1)

Q4 · Look ahead: forecasting

Altini & Dunne say most current wearable insights only describe what has already happened. What role do they expect for machine learning?

Show answer

Answer: C. The authors say prediction from wearable data is still limited and expect machine-learning models that forecast individual responses, with examples such as glucose after a meal or next-day HRV. Battery power (D) is discussed as a separate hardware problem, not as the role of machine learning.

Source: Paper p. 91 (PDF p. 5)

Q5 · Three steps applied 3 points

A company launches a new wrist sensor and wants to use data from 100,000 users to discover new links between heart rate and illness. Altini & Dunne (2021) say two steps must come first. State and motivate briefly:
1) what these two steps are;
2) why skipping them is risky.

Show model answer

1) Step 1: validate the device: show it measures heart rate accurately compared with a reference system (and, for medical use, get certification). Step 2: show it reproduces known results from smaller lab or clinical studies, for example the heart-rate rise after exercise or during the first days of sickness. Only then (step 3) use the large data set to discover new relationships. 2) 'Garbage in, garbage out': everyday wearable data are noisy and can be biased, so without steps 1–2 the company could report confident but false findings, and it could not tell a real new effect from a device error.

How the points are earned

  • 1 pt Step 1: validate the device's accuracy against a reference system
  • 1 pt Step 2: reproduce known results from smaller lab or clinical studies (e.g., response to exercise or sickness)
  • 1 pt Risk: noisy or biased data give false findings; cannot tell a real effect from a device error (garbage in, garbage out)

Source: Paper pp. 88–90 (PDF pp. 2–4); Slides pp. 57–59

Lecture 9 – User-generated data

research letter (observational cohort study)

Radin et al. (2021) — Assessment of Prolonged Physiological and Behavioral Changes Associated With COVID-19 Infection

Radin, J. M., Quer, G., Ramos, E., Baca-Motes, K., Gadaleta, M., Topol, E. J., & Steinhubl, S. R. (2021). Assessment of prolonged physiological and behavioral changes associated with COVID-19 infection. JAMA Network Open, 4(7), e2115959. https://doi.org/10.1001/jamanetworkopen.2021.15959 · Open paper

In the DETECT app study, wearable data showed that people with COVID-19 took much longer than people with similar symptoms who tested negative to get back to their own normal levels. Their resting heart rate took on average 79 days, and in a small group it was still raised after more than 133 days.

Question

After COVID-19, how long does it take for resting heart rate (a body signal) and for sleep and daily steps (behavior) to return to the person's own pre-illness level, called their baseline? How much does this recovery differ between people? The comparison group was people with similar symptoms who tested negative for COVID-19.

Why it matters

Long-term problems after COVID-19, such as autonomic dysfunction (problems with the part of the nervous system that automatically controls heart rate and other body functions) and heart damage, had been reported for up to 6 months but had not been measured in numbers. Wearables (sensors worn on the body, such as smartwatches and fitness trackers) record body signals and behavior continuously, starting while people are still healthy. So they capture a person's healthy baseline, the illness and the recovery, which a normal clinical study rarely does.

Approach

DETECT is a remote, app-based cohort study (a study that follows the same group of people over time) that enrolled 37,146 adults across the US between March 2020 and January 2021 and collected their wearable data. It is observational: the researchers only watched, they did not change anything. This analysis used the 875 participants who reported symptoms of an acute respiratory illness (a sudden infection of the airways) and had a COVID-19 swab test: 234 tested positive and 641 tested negative (the comparison group). The swab test was the reference: the verified answer to 'who had COVID-19'. For each person, daily resting heart rate (RHR) was expressed as the difference from their own baseline: daily RHR minus their average RHR before illness. Sleep and daily step count were also followed, from 7 days before to 133 days after symptoms started. The COVID-positive people were then split into three groups by how far their RHR stayed above baseline 28–56 days after symptoms started (less than 1, 1–5, or more than 5 beats per minute, bpm). Symptoms in these groups were compared with χ² tests (tests whether how often something occurs differs between groups) and age with ANOVA (a test that compares the means of three or more groups). The analysis is descriptive: average curves with 95% confidence intervals (the range that very likely contains the true average). There is no prediction model.

Findings

  • People with COVID-19 took longer than symptomatic people who tested negative to return to their baseline for resting heart rate, sleep and activity.
  • The difference was largest for RHR: first a short drop below baseline (bradycardia: a slower heart rate than normal), then a long rise above baseline (tachycardia: a faster heart rate than normal) that on average did not return to baseline until 79 days after symptoms started. Step count and sleep returned sooner, after 32 and 24 days.
  • A small group of 32 COVID-positive people (13.7%) kept an RHR more than 5 bpm above baseline that had not returned to normal after more than 133 days.
  • In the first days of illness this group reported more cough (84.4%), body ache (62.5%) and shortness of breath (28.1%) than the other groups. This suggests that early symptoms and a larger first rise in RHR may go together with a longer recovery.
  • Overall, COVID-19 affected the body for about 2–3 months on average, with large differences in recovery. This may reflect problems in the autonomic nervous system or ongoing inflammation.

Limitations

  • Authors: symptoms were only recorded in the first (acute) phase of the illness, so the long-term changes in RHR, sleep and activity could not be compared with long-term symptoms (long COVID).
  • Authors: larger samples and more complete reports from participants are needed to explain why recovery differs between people.
  • Own inference: the sample is observational and self-selected (US volunteers who own a wearable, about 71% women), so the results may not apply to everyone, and associations do not prove cause.
  • Own inference: the curves are group averages; single people's recovery paths differ a lot (the lecture's warning: 'These are group averages. What about the individual?').

Link to the lectures

Example of wearable data showing sickness and recovery, which is hard to study in the lab; swab tests are the reference outcome.

In Lecture 9, user-generated wearable data are presented as a way to study outcomes that are hard to test in the lab, such as sickness; detecting an infection with the Oura ring is one of the opening examples (Slides pp. 21, 69, 73). Radin et al. show this in practice: a large remote group, each person compared with their own baseline, and swab-test results as the reference data the lecture calls essential ('When did we get infected? How do we determine it? Symptoms, tests', Slides pp. 115–120). Altini & Dunne (2021) reprint Radin's figure as their example of large-scale discovery. Note: the slides do not show this paper. The lecture's pandemic slides (pp. 124–128) show a different analysis (resting HR of HRV4Training users during the 2020 lockdown), and 'L10 - fitbit_study.pdf' (Natarajan et al., 2020) is a separate paper that the slides do not use.

Remember

  • Research letter; observational cohort study (DETECT app study, US); descriptive analysis, not a prediction model.
  • Compared 234 COVID-positive with 641 COVID-negative people who all had symptoms and a swab test.
  • Outcome = difference from each person's own baseline (daily RHR − baseline RHR), plus sleep and step count.
  • RHR took longest to return to normal: on average 79 days (steps 32 days, sleep 24 days).
  • RHR pattern: a short drop (bradycardia), then a long rise (tachycardia).
  • 13.7% of COVID-positive people had RHR more than 5 bpm above baseline for over 133 days; they had more cough, body ache and shortness of breath early on.
  • Main limitation: symptoms were only recorded in the acute phase.

Practice

Q1 · Study design

Which description best fits the approach of Radin et al. (2021)?

Show answer

Answer: D. The DETECT participants were followed over time, and their RHR, sleep and steps were expressed as differences from their own baseline; the results are descriptive average curves. Option B describes earlier detection studies that the paper cites, not this paper, which built no prediction model.

Source: Radin et al., Introduction and Methods sections

Q2 · Why wearables

Why were wearables especially suited to Radin et al.'s question about recovery from COVID-19?

Show answer

Answer: A. The introduction says wearables can track a person from when they are healthy, through infection, and back to baseline. Option B is wrong: who had COVID-19 was decided by swab tests, which served as the reference data.

Source: Radin et al., Introduction and Methods sections

Q3 · Main finding

What did Radin et al. find for the people who had COVID-19?

Show answer

Answer: C. RHR first dropped briefly, then stayed raised for a long time and returned to baseline only after 79 days on average; steps and sleep were back after 32 and 24 days. Option D describes only the short first drop, not the long rise that followed.

Source: Radin et al., Results section

Q4 · Slow-recovery group

Radin et al. found a group of COVID-positive people whose resting heart rate stayed more than 5 beats per minute above baseline for over 133 days. What was true of this group?

Show answer

Answer: B. 32 people (13.7%) were in this group, and early in the illness they reported more cough (84.4%), body ache (62.5%) and shortness of breath (28.1%). Option D states the opposite. The authors present this as an association, not a proven cause.

Source: Radin et al., Results and Discussion sections

Q5 · Limitations of wearable cohort data 2 points

Radin et al. followed the wearable data of 234 COVID-positive and 641 COVID-negative app users from a week before their symptoms until 133 days after. State and motivate briefly:
1) the limitation the authors themselves mention;
2) one other limitation of using these user-generated data.

Show model answer

1) Symptoms were only recorded in the first (acute) phase of the illness. So the long-lasting changes in resting heart rate, sleep and activity could not be linked to long-term (long COVID) symptoms: we know the body stays changed, but not whether people still feel ill. 2) Any one of: the participants chose to join and own a wearable (US volunteers, about 71% women), so the results may not apply to other people; the study is observational, so links (such as more early symptoms with a longer recovery) do not prove cause; the curves are group averages that hide large differences between people, so they cannot predict one person's recovery.

How the points are earned

  • 1 pt Names the authors' limitation: symptoms only recorded in the acute phase, so long-term changes cannot be linked to long-term symptoms
  • 1 pt Names and explains another limitation: self-selected sample / observational so not causal / group averages hide individuals

Source: Radin et al., Results and Discussion sections; Slides pp. 115–120, 129–130