Data Science in Sport and Health · Practice
Question bank
150 exam-style questions, in addition to the two on each lecture page. Tap an option to check it; open questions have a model answer to compare against. The course-paper questions are on the course literature page.
Lecture 1 · Introduction to data science
12 questions · lecture page
Q1 · Disciplines of data science (Venn diagram)
The lecture's Venn diagram places data science at the overlap of computer science/IT, math and statistics, and domain/business knowledge. A sports physiologist has strong statistics skills and deep domain knowledge but hardly any programming or IT skills. In which overlap of the diagram does this profile fall?
Show answer
Answer: C. Statistics plus domain knowledge without computer science is what the diagram calls traditional research. Machine learning is computer science plus statistics without domain knowledge, and only the overlap of all three circles is data science.
Source: Slides p. 51 and 58
Q2 · Data, information, knowledge, wisdom
The lecture's data-to-wisdom pyramid (from knowledge discovery in databases, KDD) describes one football moment at four levels. At which level does the statement 'I am pretty close to the goal' sit, and why?
Show answer
Answer: B. Knowledge is information understood in context: the player realises what the position means for him. The tempting option, information, is the organised description 'The right winger has the ball at the corner of the penalty area'; 'Ball, player11, x=84, y=81' is raw data and 'I better shoot on goal to score!' is wisdom (a decision).
Source: Slides p. 46–47
Q3 · Typical data science questions
A sports physician has three requests: (1) 'How many weeks will this athlete need to return to play?' (2) 'Is today's resting heart rate of this athlete weird compared with her usual values?' (3) 'Which of three training programmes should this athlete follow next month?' According to the lecture's list of typical data science questions, which approaches match requests 1, 2 and 3?
Show answer
Answer: D. 'How many weeks?' asks for a number (regression), 'Is this weird?' is anomaly detection and 'Which option should be taken?' is recommendation. Classification would need a known category (e.g. injured yes/no), and clustering finds groups that nobody defined in advance.
Source: Slides p. 66
Q4 · Structured vs unstructured data
Which statement about structured and unstructured data matches the lecture?
Show answer
Answer: A. The lecture describes unstructured data (video, images, audio, e-mails) as not fitting rows and columns, about 80% of enterprise data (all data an organisation stores) and needing more storage. Option C swaps the numbers: less than 50% of structured data is used in decisions and less than 1% of unstructured data is analysed; B gives the 80% to the wrong kind, and D is wrong because unstructured data have no data model.
Source: Slides p. 38, 40–41
Q5 · Data science vs statistics
The lecture contrasts statistics with data science. Which set of characteristics does the lecture attribute to data science?
Show answer
Answer: C. The slide calls data science an emerging field that starts with data (data mining: hypothesis generation) and works best with lots of data, many rows (long) and many columns (wide). Options A and B describe statistics; D is wrong because the slide says both fields need care and rigour.
Source: Slides p. 54
Q6 · Blei & Smyth (2017)
Blei & Smyth (2017) argue that data science is more than the combination of statistics and computer science. Which of the following is NOT one of the additional requirements named in their quote on the slide?
Show answer
Answer: D. The quote names three extra duties: understand the context of data, take responsibility for private and public data, and communicate what a dataset can and cannot tell us. Collecting everything without a question is not in the quote and clashes with the lecture's advice to start from the problem.
Source: Slides p. 49–50
Q7 · Data science lifecycle: order of steps
Which sequence follows the order of the data science lifecycle shown in the lecture?
Show answer
Answer: B. The lifecycle runs: identify the problem → acquisition → preparation → exploration → feature engineering → modelling → visualization → present & communicate → deployment & maintenance. The tempting option A puts exploration before preparation, but in the lecture's cycle the data are cleaned before they are explored.
Source: Slides p. 64–65
Q8 · R as a programming language
According to the lecture, which of the following is a disadvantage of R?
Show answer
Answer: A. The slide lists two minuses: slow computation with large datasets, and many packages that do the same thing ('many ways to Rome'). The other options contradict its pluses: R is free and open source, has a large community, is strong in visualizations and statistics, and is used in academia and many industries.
Source: Slides p. 24
Q9 · Data science applications in sport
The lecture (Chmait & Westerbeek, 2021) groups examples of data science in sports research into four areas. Which pairing of example and area is correct as listed on the slide?
Show answer
Answer: B. Training & coaching lists team formation efficacy, training optimization and player injury modelling. Ticket pricing is fan & business, event classification is game analytics, and player recruitment is talent identification. Exam note: the slide misspells the authors as 'Chmait & Westerpoort'; the paper is Chmait & Westerbeek (2021).
Source: Slides p. 52
Q10 · RStudio environment
You have imported a data frame in RStudio and want to see which objects you have created so far in this session. Which of the four RStudio panes shown in the lecture is meant for this?
Show answer
Answer: C. The workspace & history pane lists the objects you have created (and the commands you ran). The console runs commands and shows their output, but it does not keep an overview of your objects.
Source: Slides p. 27
Q11 · Hypothesis generation vs hypothesis testing 3 points
A club analyst explores three seasons of existing GPS and injury data. She finds that players who did more high-speed running in the week before a match had fewer hamstring injuries. State and motivate briefly:
1) Compare hypothesis generation and hypothesis testing: what does each start from, and which field does the lecture link to each?
2) In which phase does the analyst's finding sit, and what is needed before a scientific journal would accept it?
Show model answer
1) Hypothesis generation starts from an existing dataset: you explore it and come up with a possible relationship. The lecture links it to data science (data mining); a coach is already pretty happy here. Hypothesis testing starts from a stated hypothesis and checks it on newly collected data. The lecture links it to statistics; a journal is only happy here. 2) Hypothesis generation, because she found the pattern by exploring data she already had. Next, state the hypothesis in advance, collect new data (e.g. next season or another team) and test it with statistics, because a pattern found by exploring can be chance.
How the points are earned
- 1 pt Generation starts from existing data (exploring) and is linked to data science
- 1 pt Testing starts from a stated hypothesis plus new data and is linked to statistics
- 1 pt The finding is hypothesis generation; it must still be tested on newly collected data
Source: Slides p. 54–55
How did your answer compare?
Q12 · Step 1: identify the problem 3 points
A hockey coach says: 'We have lots of wearable data, can you do something with AI?' State and motivate briefly:
1) Name two things you do in the first lifecycle step (identify the problem) before acquiring any data.
2) Turn the coach's request into one specific question with a clear target variable, and choose the matching approach (regression, classification, clustering, anomaly detection or recommendation).
Show model answer
1) Any two of: talk to domain experts (coach, physio, players); keep asking 'why?' to find the real need; decide which questions you ask and how answering them helps the team's goal; specify the target variable. 2) Example: 'Will a player get injured in the next four weeks (yes/no)?' Target variable: injury yes/no in the next four weeks. Approach: classification, because the answer is a known category. Other well-defined questions also count, e.g. 'How tired will a player feel tomorrow on a 1–10 scale?' → regression, because the answer is a number.
How the points are earned
- 1 pt Two correct actions for step 1 (0.5 each)
- 1 pt One specific question with a clearly named target variable
- 1 pt Approach matches the question, with a reason (number → regression, category → classification)
Source: Slides p. 66–67
How did your answer compare?
Lecture 2 · Programming in R
13 questions · lecture page
Q1 · Vectors and coercion
What does the following code print?
x <- c(12, "15", TRUE)
typeof(x)Show answer
Answer: B. A vector holds one type only, so R converts (coerces) every element to the most flexible type present; one text value turns everything into text: "12" "15" "TRUE". Only a list keeps mixed types, and typeof() says "double" only when all elements are ordinary numbers.
Source: Slides p. 9–10, 16–18; DataCamp Introduction to R (vectors). Output verified in R 4.6.1.
Q2 · Variables, assignment and packages
Which statement about working in R is correct according to the lecture?
Show answer
Answer: D. install.packages("lubridate") downloads and installs a package once per computer; library(lubridate) loads it in every session so that ymd() works. R is case-sensitive, names cannot start with a digit (2nd_test gives an error), and ls() lists only object names; typeof() gives the type.
Source: Slides p. 13, 15–16, 57
Q3 · Matrices
What does the last line return?
m <- matrix(1:6, nrow = 2)
m[1, 3]Show answer
Answer: C. matrix() fills column by column unless byrow = TRUE, so the columns are (1, 2), (3, 4) and (5, 6), and row 1, column 3 holds 5. The answer 3 would only be right with byrow = TRUE; six values in two rows give three columns, so there is no error.
Source: Slides p. 19, 28; DataCamp Introduction to R (matrices). Output verified in R 4.6.1.
Q4 · Factors
What does the following code print?
event <- factor(c("sprint", "road", "track", "road", "sprint"))
as.integer(event)Show answer
Answer: A. Factor levels are sorted alphabetically by default (road, sprint, track), and each value is stored as the number of its level: sprint = 2, road = 1, track = 3. The codes do not follow the order in which the values first appear (that would give 1 2 3 2 1).
Source: Slides p. 17, 20; DataCamp Introduction to R (factors). Output verified in R 4.6.1.
Q5 · Lists
Given the list below, which expression returns the single number 61?
athlete <- list(name = "Sanne", hr = c(58, 61, 64), injured = FALSE)Show answer
Answer: C. Double brackets take the content out of a slot (the vector 58 61 64), and [2] then gives 61. Single brackets return a smaller list, so athlete["hr"][2] does not give a number; a list has no rows and columns, so athlete[2, 2] gives an error; athlete[[2]][3] gives 64.
Source: Slides p. 17–18, 20; DataCamp Introduction to R (lists). Output verified in R 4.6.1.
Q6 · Data frames and logical operators
What does the last line return?
d <- data.frame(name = c("Ann", "Bo", "Cas", "Dirk"),
vo2 = c(48, 55, 61, 52),
mass = c(60, 72, 80, 68))
d[d$vo2 > 50 & d$mass < 75, "name"]Show answer
Answer: B. & keeps rows where both conditions are TRUE: vo2 > 50 holds for Bo, Cas and Dirk, mass < 75 for Ann, Bo and Dirk, so only Bo and Dirk pass, and their "name" values are returned. The first option ignores the mass condition; the third is what | (or) would give.
Source: Slides p. 21, 29, 48–49; DataCamp Intermediate R (relational and logical operators). Output verified in R 4.6.1.
Q7 · Conditional statements
A coach labels a session by its rating of perceived exertion (rpe, an effort score from 0 to 10). What is the value of zone after running this code?
rpe <- 8
if (rpe >= 9) {
zone <- "maximal"
} else if (rpe >= 5) {
zone <- "hard"
} else if (rpe >= 7) {
zone <- "very hard"
} else {
zone <- "easy"
}
zoneShow answer
Answer: D. R checks the conditions from top to bottom and runs only the first TRUE branch: 8 >= 9 is FALSE, 8 >= 5 is TRUE, so zone becomes "hard" and the rest is skipped. The "very hard" branch can never run, because any value of 7 or more already passed rpe >= 5; that is why the highest threshold should be tested first.
Source: Slides p. 50–52; DataCamp Intermediate R (conditionals). Output verified in R 4.6.1.
Q8 · For-loops
What does the last line print?
hr <- c(150, 172, 185, 168, 190, 176)
count <- 0
for (h in hr) {
if (h >= 170 & h < 185) {
count <- count + 1
}
}
countShow answer
Answer: A. count goes up only when 170 ≤ h < 185, which is true for 172 and 176. 185 fails h < 185 (with h <= 185 the answer would be 3), and 6 is the number of rounds of the loop, not the number of TRUE conditions.
Source: Slides p. 49, 53–56; DataCamp Intermediate R (loops). Output verified in R 4.6.1.
Q9 · Writing functions
What do the two function calls return?
abs_vo2 <- function(rel, mass = 70) {
rel * mass / 1000
}
abs_vo2(c(50, NA, 60))
abs_vo2(mass = 80, rel = 50)Show answer
Answer: C. mass has a default of 70, so the first call works element by element: 50 × 70 / 1000 = 3.5, NA stays NA (nothing removes it), 60 × 70 / 1000 = 4.2. A default is only used when the argument is left out; in the second call mass = 80 is given by name, so 50 × 80 / 1000 = 4.
Source: Slides p. 43–46; DataCamp Intermediate R (functions). Output verified in R 4.6.1.
Q10 · Missing values and na.rm
What does the last line return?
vo2 <- c(40, 50, NA, 60)
mean(vo2, na.rm = TRUE)Show answer
Answer: D. na.rm = TRUE leaves the missing value out and averages the three known values: (40 + 50 + 60) / 3 = 150 / 3 = 50. NA is what mean(vo2) gives without na.rm = TRUE, and 37.5 wrongly divides by 4, still counting the removed NA.
Source: Slides p. 36–38. Output verified in R 4.6.1.
Q11 · Reading str() output 3 points
You import a CSV file (a plain-text table) with 500 m speed-skating times as the data frame races and run str(races); the output is shown below. State and motivate briefly:
1) How many observations (rows) and variables (columns) does races have?
2) Why does mean(races$time) return NA, and which line of R code fixes the problem?
str(races)
#> 'data.frame': 4 obs. of 3 variables:
#> $ skater: chr "Ann" "Bo" "Cas" "Dirk"
#> $ season: int 2021 2021 2022 2022
#> $ time : chr "35.40" "36.07" "34.98" "35.62"Show model answer
1) 4 observations (rows) and 3 variables (columns): '4 obs. of 3 variables'. 2) The time column is chr (character, i.e. text): the numbers were read in as text, and mean() of text returns NA with a warning. Convert it first with races$time <- as.numeric(races$time); then mean(races$time) works.
How the points are earned
- 1 pt 4 observations (rows) and 3 variables (columns)
- 1 pt time is stored as character (chr, text), so mean() cannot average it
- 1 pt Fix: races$time <- as.numeric(races$time)
Source: Slides p. 23, 29; Recording 22:34. Output verified in R 4.6.1.
How did your answer compare?
Q12 · Writing a function with if / else 3 points
A coach wants every heart-rate value labelled automatically. State and motivate briefly:
1) Write an R function hr_zone(hr) that returns "high" for 180 beats per minute (bpm) or more, "moderate" for 150–179 bpm and "low" below 150 bpm.
2) What does hr_zone(180) return, and why must the 180 test come before the 150 test?
Show model answer
1) hr_zone <- function(hr) { if (hr >= 180) { "high" } else if (hr >= 150) { "moderate" } else { "low" } } 2) "high", because 180 >= 180 is TRUE. R runs only the first branch whose condition is TRUE, so if hr >= 150 were tested first, 180 and above would already get "moderate" and the "high" branch would never run.
How the points are earned
- 1 pt Function defined as hr_zone <- function(hr) { ... }
- 1 pt if / else if / else chain with the highest threshold first and the right labels
- 1 pt hr_zone(180) gives "high"; only the first TRUE branch runs, so the highest threshold goes first
Source: Slides p. 43–46, 50–52. Output verified in R 4.6.1.
How did your answer compare?
Q13 · Storing results with a for loop 2 points
The function hr_zone(hr) returns "high" for 180 beats per minute (bpm) or more, "moderate" for 150–179 bpm and "low" below 150 bpm. You have session_hr <- c(132, 155, 181, 149). State and motivate briefly:
1) Write a for loop that stores the zone of every value of session_hr in a new vector zones.
2) Give the contents of zones after the loop.
Show model answer
1) zones <- character(length(session_hr)) for (i in seq_along(session_hr)) { zones[i] <- hr_zone(session_hr[i]) } 2) "low" "moderate" "high" "low": 132 is below 150, 155 lies in 150–179, 181 is 180 or more, and 149 is below 150.
How the points are earned
- 0.5 pt for loop over the positions (or values) of session_hr
- 1 pt Empty result vector made first and filled one slot per round: zones[i] <- hr_zone(session_hr[i])
- 0.5 pt Result: "low" "moderate" "high" "low"
Source: Slides p. 53–56. Output verified in R 4.6.1.
How did your answer compare?
Lecture 3 · Data in sport and health
12 questions · lecture page
Q1 · Structured vs unstructured data in sport and health
A rehabilitation centre films patients walking, runs markerless motion-capture software that writes the hip, knee and ankle positions of every video frame into a table, and lets therapists add a free-text note after each session. How should these three data sources be classified, following the lecture?
Show answer
Answer: A. Raw video is unstructured, but once software has written the joint positions into a table you are working with structured numbers ('it's not a video anymore'); open free-text answers are unstructured. Deriving a table from a video does not make the table unstructured (C).
Source: Slides p. 6; Recording 05:00–10:00
Q2 · Interpreting the Fitbit graph
The Fitbit slide plots average resting heart rate (y-axis) against BMI (x-axis) for men and women. Which conclusion is best supported by the graph?
Show answer
Answer: B. The women's points lie above the men's at every BMI, and both curves are lowest around BMI 20–24 and rise towards both lower and higher BMI (a U-shape); men at BMI 45 (about 70 bpm) sit above women at BMI 20 (about 65 bpm). The data are averages over different users, so they show a pattern, not what happens when one person changes their BMI (D).
Source: Slides p. 10; Recording 25:00–30:00 and 35:00–45:00
Q3 · Value of data: activity-based insurance discount
The lecture discussed an insurer (ASR) that gives discounts and gifts when members share activity-tracker data and reach personal activity goals. What did the lecturer report about the effect of such a scheme?
Show answer
Answer: C. According to the lecturer, the scheme mainly rewarded people who were already active, did not make inactive people active, and people who lost the discount experienced that as a punishment. Commercial value of activity data does not automatically mean behaviour change (A).
Source: Slides p. 9; Recording 15:00–25:00
Q4 · Restricted range in regression
In the van der Zwaard et al. (2018) cycling study, pennation angle (the angle at which the thigh-muscle fibres pull, measured with ultrasound) showed no relationship with sprint power in the elite cyclists. According to the lecture, how should you interpret this?
Show answer
Answer: D. This is restricted range: if a predictor hardly varies, no relationship can appear, so 'no significant relationship' does not prove 'no relationship' (the same logic as VO₂max in the 10 best marathon runners); always check each variable's spread first. Option A draws exactly the wrong conclusion, and Deming regression allows error in both variables rather than ignoring it (B).
Source: Slides p. 22, 27; Recording 60:00–70:00
Q5 · Determinants of combined sprint and endurance performance
In the van der Zwaard et al. (2018) results shown in the lecture, which combination of muscle characteristics together explained 67% of the differences between riders in performance VO₂ (R² = 0.67), a key driver of combined sprint and endurance performance?
Show answer
Answer: C. The slide's highlight box shows oxidative capacity ↑, capillaries × myoglobin ↑ and PCSA ↓ explaining 67% of performance VO₂. Option A is the best pair for sprint peak power (65%), not for performance VO₂. Exam note: in the recording the lecturer says this 'explains combined performance by 67%'; strictly, the 67% refers to performance VO₂.
Source: Slides p. 27–28, 30; Paper p. 2110 (abstract) and p. 2115
Q6 · Excel and CSV limits
A sports scientist stores heart rate once per second for 25 athletes during a 2-hour training, with one row per athlete per second: 25 × 7,200 = 180,000 rows and 3 columns. Which statement about saving this in a single sheet or file is correct, using the lecture's limits?
Show answer
Answer: A. The slide gives .xls 65,536 rows and 256 columns, .xlsx 1,048,576 rows and 16,384 columns, and csv 'unlimited'; 180,000 rows is too many for .xls but fits in .xlsx. 256 is the .xls column limit, not the .xlsx one (D).
Source: Slides p. 34
Q7 · Application Programming Interface (API)
Which statement about an API matches the lecture?
Show answer
Answer: B. The slides say the API is not the database or the server but the code that governs the access point(s) to the database, used by sending HTTP requests without a frontend, with data returned as e.g. JSON or XML. Option C describes web scraping, and the lecturer stressed that the quality of public data is often unknown (D).
Source: Slides p. 38, 41–43; Recording 75:00–90:00
Q8 · Speed-skating API example
Seasonal bests of speed skaters were retrieved through the SpeedskatingResults.com API and shown in a radar (spider-web) chart per distance, as the seasonal best relative to the world record on the race day. In this chart, what do the direction and the size of the filled shape show?
Show answer
Answer: D. The slides state that the direction of the shape shows specialisation and its size shows performance level; symbols show track type and a circle marks a world record. Option C swaps the two; because the rim is the world record (1.0), a bigger shape means faster times.
Source: Slides p. 45, 47, 49–51; Recording 80:00–85:00
Q9 · Extract–transform–load (ETL)
A club combines training files from three GPS-watch brands into its data warehouse. Removing duplicate sessions, converting all speeds to m/s and flagging heart rates above a plausible limit belong to which ETL step, and where does that usually happen?
Show answer
Answer: B. Transform covers cleaning, deduplication, format revision and threshold validation checks, usually in a staging area where the data are tested before loading. Extract only copies the raw data out of the sources into staging (A), and load writes the converted data into the warehouse.
Source: Slides p. 57–61
Q10 · Data warehouse vs data lake
Which characteristic belongs to a data lake rather than a data warehouse, according to the lecture's comparison?
Show answer
Answer: C. The comparison table gives the data lake raw structured, semi-structured and unstructured data, schema-on-read, low-cost storage, high agility and data scientists as users. Options A and B describe a data warehouse, and D describes a data mart.
Source: Slides p. 62, 64–65; Recording 90:00–100:00
Q11 · Deming regression and the sprint–endurance trade-off 4 points
In van der Zwaard et al. (2018), sprint power (Wingate test) and endurance power (15-km time trial) of 28 cyclists were negatively correlated (r = −0.66, p < 0.001). The trade-off line was fitted with Deming regression instead of ordinary least squares. State and motivate briefly:
1) How does Deming regression differ from ordinary least squares, and why does it suit this study? (2 points)
2) Interpret r = −0.66, including the share of variance that sprint and endurance power have in common. (2 points)
Show model answer
1) Ordinary least squares makes the vertical gaps between the points and the line as small as possible and treats the x variable as measured without error. Deming regression allows measurement error in both variables and (with equal errors) uses the perpendicular, shortest distance to the line. That suits this study because sprint and endurance power are both noisy test results and neither is the natural predictor; each rider's perpendicular distance to the line then gives a combined sprint-plus-endurance score. 2) r = −0.66 is a fairly strong negative relationship: riders with more sprint power tend to have less endurance power, so combining both is hard. Squared, r² = 0.66 × 0.66 ≈ 0.44, so about 44% of the differences are shared (not 66%). p < 0.001 makes luck unlikely, but a correlation does not prove cause and effect. Exam note: in the recording the lecturer said '36 to 40 percent'; the correct value is 0.66² ≈ 0.44.
How the points are earned
- 1 pt OLS minimises the vertical gaps and treats x as error-free
- 1 pt Deming allows error in both variables (perpendicular distance); suits because both tests are noisy and neither is the predictor, and the perpendicular distance is a combined score
- 1 pt Negative, fairly strong: more sprint power goes with less endurance power
- 1 pt r² ≈ 0.44, so about 44% shared variance (not 66%)
Source: Slides p. 21, 23, 25; Paper p. 2111, 2113 and Fig. 2 (p. 2116); Recording 45:00–55:00
How did your answer compare?
Q12 · Choosing a storage solution 4 points
A national skating federation has three data needs: (a) raw sensor files, race videos and spreadsheets whose future use is still unclear; (b) cleaned results and training data from many sources that management wants to combine for reports; (c) a medical staff that only needs summarised injury and load data. State and motivate briefly:
1) Which storage solution (data lake, data warehouse or data mart) fits each need? (3 points)
2) Name one drawback of a data lake that the lecturer mentioned. (1 point)
Show model answer
1) (a) Data lake: it keeps structured and unstructured data raw, in their original format, without planning the analysis first (schema-on-read). (b) Data warehouse: one central store that combines cleaned (ETL-processed) data from many sources across the organisation for reports and decisions (schema-on-write). (c) Data mart: a smaller, focused slice of the warehouse with summarised data for one department or purpose. 2) A lake is easy to fill but hard to reuse: the data stay raw and unclean, so every later analysis needs a lot of preparation (the lecturer's 'database for lazy people'). She advised building data marts so that others can reuse the data.
How the points are earned
- 1 pt (a) Data lake, because data are kept raw in their original format for analyses not yet planned
- 1 pt (b) Data warehouse, because it combines cleaned data from many sources across the organisation for reporting
- 1 pt (c) Data mart, because it is a focused, summarised slice of the warehouse for one unit
- 1 pt Drawback: low effort to store but high effort to reuse (chaotic, unclean data)
Source: Slides p. 55–56, 62–65; Recording 90:00–100:00
How did your answer compare?
Lecture 4 · Preparation, exploration and visualization
14 questions · lecture page
Q1 · Measurement scales
Which statement is meaningful, given the measurement scale of the variable?
Show answer
Answer: C. Training duration is a ratio variable (0 minutes means no time at all), so "twice as long" makes sense. Celsius is interval: 0 °C is not "no temperature", so "twice as warm" is meaningless; medal codes stay ordinal and shirt numbers nominal, so arithmetic on those codes means nothing.
Source: Slides p. 7–9
Q2 · Data quality dimensions
Which pairing of a data problem with a DOMA data-quality dimension is correct?
Show answer
Answer: B. Validity asks whether a value follows the rules for format, type and range, and "1m80" is text where a number in cm is required. The other pairs are mixed up: a double import is uniqueness, the 28 kg footballer is accuracy (the value does not match the real person), and male + pregnant is consistency.
Source: Slides p. 10
Q3 · FAIR data
You deposit your thesis dataset in a trusted repository. It has a persistent identifier (DOI) and rich, machine-readable metadata, and it can be downloaded after logging in. You did not add a usage licence or any information on how, when and by whom the data were collected. Which FAIR principle is least well met?
Show answer
Answer: D. Reusable means the data have a clear usage licence and information on where they came from and how they were collected (provenance); both are missing. The DOI and rich metadata cover Findable, the trusted repository with a login covers Accessible (accessible does not have to mean open to everyone), and machine-readable metadata supports Interoperable.
Source: Slides p. 11–12; Recording 00:10–00:15
Q4 · dplyr: filter with | and NA
How many rows does this pipeline return?
olympics_small <- data.frame(
Name = c("A", "B", "C", "D", "E"),
Year = c(2008, 2008, 2012, 2016, 2012),
Medal = c("Gold", NA, "Gold", "Silver", NA)
)
olympics_small %>%
filter(Year == 2008 | Medal == "Gold") %>%
nrow()Show answer
Answer: A. | keeps a row when at least one condition is TRUE: A (both), B (Year is 2008, so its missing medal does not matter) and C (Gold). D fails both; for E the year is not 2008 and "is NA equal to Gold?" has no answer, so the condition is NA and filter() drops the row. Output verified in R 4.6.1.
Source: Slides p. 24–25
Q5 · Duplicates and distinct()
In the Olympics data each row is one athlete in one event at one Games (271,116 rows; 135,571 different athlete IDs). The lecturer ran the code below. Why does distinct() remove 84,039 rows?
df1 <- data_olympics %>%
select(ID, Name, Age) %>%
rename(Athlete_code = ID)
df1 %>% nrow() # 271116
df1 %>% distinct() %>% nrow() # 187077Show answer
Answer: B. Whether rows count as duplicates depends on which columns you keep: once the event column is dropped, one athlete's events at the same Games look identical. 187,077 is not the number of athletes (135,571 IDs), because the same athlete at different ages keeps separate rows. Output verified in R 4.6.1.
Source: Slides p. 29, 39; Recording 01:10–01:15
Q6 · Joining data frames
You want a table with all four athletes and their countermovement-jump (CMJ) result, with NA where no test exists. Which call gives this?
athletes <- data.frame(ID = c(1, 2, 3, 4),
Sport = c("Rowing", "Judo", "Rowing", "Cycling"))
tests <- data.frame(Athlete_code = c(2, 4, 5),
CMJ_cm = c(41, 38, 45))Show answer
Answer: D. left_join() keeps every row of the first table and fills in NA where no test exists; by = c("ID" = "Athlete_code") matches key columns that have different names. The inner join keeps only IDs 2 and 4, the right join keeps IDs 2, 4 and 5, and by = "ID" gives an error because tests has no ID column (verified in R 4.6.1).
Source: Slides p. 26–28
Q7 · EDA: suspicious values
summary() of the Olympics data shows a minimum weight of 25 kg, a maximum height of 226 cm and a maximum age of 97 years. What is the most appropriate next step?
Show answer
Answer: C. EDA means checking odd values in context before deciding. In the lecture all three were real: a 14-year-old gymnast of 25 kg, Yao Ming at 226 cm and a 97-year-old in the art competitions. Deleting them automatically would remove real athletes; replacing them with the mean would invent data.
Source: Slides p. 34–35, 40–42, 48; Recording 01:20–01:25
Q8 · Boxplot reading
summary() of Olympic athletes' age gives 1st Qu. = 21, Median = 24, 3rd Qu. = 28 and Max. = 97. Under the boxplot rule shown in the lecture, which statement is correct?
Show answer
Answer: A. IQR = Q3 − Q1 = 28 − 21 = 7, and the upper limit is Q3 + 1.5 × IQR = 28 + 10.5 = 38.5; older ages are drawn as separate dots to check, not to delete. The whisker stops at the oldest age inside that limit (38), not at the maximum of 97, and the box runs from Q1 to Q3 (the middle 50%).
Source: Slides p. 40, 48; Recording 01:25
Q9 · Correlation and explained variance
In the lecture's corrplot of Olympic athletes, height and weight correlate at r = 0.8, and age and height at r = 0.14. Which statement is correct?
Show answer
Answer: D. The share of variation explained is r², and 0.8 × 0.8 = 0.64, so 64%, not 80%. A correlation alone never proves cause, and a small r only means a weak straight-line pattern (a curved relation is still possible). Exam note: slide p. 53 in the PDF shows the Nicolas Cage chart (r = 0.666) next to "0.947091² = 89.7%" because two animation steps were printed on top of each other; 89.7% belongs to the cheese/bedsheet chart (r = 0.947) and 44.4% to the Nicolas Cage chart.
Source: Slides p. 52–53
Q10 · ggplot2: facets and layers
Using the mpg dataset, which code produces a separate scatterplot panel of displ against hwy for each vehicle class?
Show answer
Answer: B. facet_wrap(~ class) makes one small panel per class (7 panels), and ggplot2 layers are added with +. Colour inside aes() keeps all classes in one panel, the %>% version gives an error, and group or theme_minimal() do not create panels.
Source: Slides p. 61–65, 68–70
Q11 · Missing values in the Olympics data 3 points
You added a BMI column (calculated from Height and Weight) to the Olympics data and printed the percentage of missing values (NA) per column, shown below. State and motivate briefly:
1) Why is Medal missing in 85% of the rows, and is this a data-quality problem?
2) How would you handle the missing Medal values in R?
3) Why does BMI have more missing values than Height or Weight?
> round(colSums(is.na(data_olympics)) / nrow(data_olympics) * 100, 1)
ID Name Sex Age Height Weight Team NOC Games Year Season
0.0 0.0 0.0 3.5 22.2 23.2 0.0 0.0 0.0 0.0 0.0
City Sport Event Medal BMI
0.0 0.0 0.0 85.3 23.7Show model answer
1) Most entries do not win a medal (only three medals per event), so NA here means "no medal". The missing value carries information; it is not lost data, so it is not a real completeness problem. 2) Recode it instead of deleting or guessing: mutate(Medal = ifelse(is.na(Medal), "No medal", Medal)). Deleting those rows would throw away 85% of the data. 3) BMI is calculated from both height and weight, so it is NA whenever either one is missing. It therefore has at least as many NAs as the more incomplete of the two (23.7% vs 23.2% and 22.2%).
How the points are earned
- 1 pt Medal NA means "no medal" (only 3 medals per event): information, not a real data-quality failure
- 1 pt Recode NA as a category such as "No medal" (mutate + ifelse(is.na(Medal), ...)), do not delete or impute
- 1 pt BMI needs both height and weight, so it is NA when either is missing
Source: Slides p. 10, 30–31, 54; Recording 01:10–01:15. Output verified in R 4.6.1.
How did your answer compare?
Q12 · dplyr: count and average per group 3 points
The data frame data_volley has one training session per row, with the player's Position ("setter" or "hitter") and the session RPE (rating of perceived exertion, 0–10); some RPE values are missing. Using dplyr, write the lines of code that give, for each position:
1) the number of sessions;
2) the mean RPE, skipping missing values.
data_volley <- data.frame(
Player = c("Anna", "Anna", "Bo", "Cas", "Cas"),
Position = c("setter", "setter", "hitter", "hitter", "hitter"),
RPE = c(6, 4, 7, NA, 8)
)Show model answer
data_volley %>% group_by(Position) %>% summarise(sessions = n(), mean_RPE = mean(RPE, na.rm = TRUE)) Result for these rows: hitter 3 sessions, mean RPE 7.5 (the NA is skipped, so (7 + 8) / 2); setter 2 sessions, mean RPE 5. n() counts rows, including the session with the missing RPE.
How the points are earned
- 1 pt Pipe the data into group_by(Position)
- 1 pt summarise() with n() to count the sessions per group
- 1 pt mean(RPE, na.rm = TRUE) inside summarise()
Source: Slides p. 22–25 (pipe, group_by + summarise, Exercise 2); Recording 54:19. Output verified in R 4.6.1.
How did your answer compare?
Q13 · dplyr: new feature and top-5 list 4 points
The data frame data_olympics has one row per athlete per event, with Height (in cm), Weight (in kg) and Sport; some heights and weights are missing. Using dplyr, write the lines of code that:
1) calculate BMI = Weight / (height in m)²;
2) compute the mean BMI per sport;
3) show the 5 sports with the highest mean BMI.
Show model answer
data_olympics %>% mutate(BMI = Weight / (Height / 100)^2) %>% group_by(Sport) %>% summarise(BMI = mean(BMI, na.rm = TRUE)) %>% arrange(desc(BMI)) %>% head(5) Height / 100 turns cm into m; without na.rm = TRUE every sport with a missing height or weight would get NA. On the course file this gives Weightlifting (27.7), Tug-Of-War (27.5), Bobsleigh (27.0), Rugby (25.7) and Baseball (25.7).
How the points are earned
- 1 pt mutate() creates BMI with height converted to metres: Weight / (Height / 100)^2
- 0.5 pt group_by(Sport)
- 1 pt summarise() with mean(BMI, na.rm = TRUE)
- 1 pt arrange(desc(BMI)) sorts from highest to lowest
- 0.5 pt head(5) keeps the top 5
Source: Slides p. 20–25 (Exercise 1: BMI with mutate and arrange; Exercise 2: group_by + summarise). Output verified in R 4.6.1.
How did your answer compare?
Q14 · Boxplots and sample size 2 points
The lecture shows boxplots of female Olympic athletes' age for every Games. Later Games give stable boxes in the early twenties, but the 1904 box has a median of 55 years. State and motivate briefly:
1) Give the most likely reason why the 1904 box does not show a real change in the age of female athletes.
2) Name one thing you would calculate or plot to check this.
Show model answer
1) Very few women competed in the early Games: in 1904 there were only 16 women's entries, all in archery, and only 13 had a known age. A box drawn from so few values is unstable, and one sport with older competitors decides it. 2) Count the female entries per Games, for example data_olympics %>% filter(Sex == "F") %>% group_by(Year) %>% summarise(n = n()), or plot the raw points (or the sports) for each year next to the boxes.
How the points are earned
- 1 pt Very few female entries in the early Games (small sample, all archery in 1904), so the box is unstable
- 1 pt Check: count the entries per year (e.g. group_by(Year) + n()) or show the raw points or sports per year
Source: Slides p. 44, 47–50; Recording 01:25–01:30. Counts computed from the course's Olympics file in R 4.6.1.
How did your answer compare?
Lecture 5 · Machine learning and regression
18 questions · lecture page
Q1 · Supervised task typecatch-up notes
You predict next week's fatigue score (a number) from training duration and workload, using past weeks in which the real fatigue score is known. Which task is this?
Show answer
Answer: B. The target is a number and its true values are known for training, so this is supervised regression. Classification predicts a category, clustering has no target at all, and reinforcement learning learns by trial and error from rewards and penalties.
Source: Slides p. 7, 11, 27
Q2 · Feature engineering: choosing an aggregatecatch-up notes
One faulty heart-rate reading can make a seven-day minimum unusually low. Why might a percentile (e.g. the 25th) of the same seven days be a more useful aggregate feature?
Show answer
Answer: A. The minimum is decided by one single value, so one faulty reading can drive the whole feature, whereas the 25th percentile reflects more of the week. That is a reason, not a guarantee: no aggregate function is always best (B) or removes measurement error (C).
Source: Slides p. 31
Q3 · Naive baseline comparisoncatch-up notes
On new athletes, a model predicting recovery time performs no better than always predicting the training-set mean. What is the strongest conclusion?
Show answer
Answer: C. Always predicting the mean is a naive baseline, a model without predictors; matching it means the features and model have shown no added value. It does not prove that no other model could work (A), and on its own it does not show overfitting (B). Exam note: the naive baseline is not on the Lecture 5 slides but appears in both reference exams.
Source: Slides p. 39–40
Q4 · AI, machine learning, deep learning and data science
Which statement matches how the lecture positions AI, machine learning (ML), deep learning and data science?
Show answer
Answer: D. The nested circles show deep learning inside ML inside AI, and the slide states that neither ML nor AI is a subset of data science, and data science is a subset of neither. They overlap, but data science also covers data preparation, statistics, visualisation and communication without any AI, so A is wrong.
Source: Slides p. 3, 24
Q5 · Dummy variables
A linear regression with an intercept predicts power output from cycling discipline (sprint, pursuit or road). Why use two dummy columns (road coded 0,
0) instead of three one-hot columns?
Show answer
Answer: A. A regression with an intercept already acts as if every row had a constant column of 1s, and road + pursuit + sprint = 1 in every row duplicates it, so the model cannot split the effect (the dummy variable trap). Road riders stay in the data as the reference group, coded 0, 0 (D is wrong). Exam note: the slide names the trap without explaining it; this duplication is the standard reason.
Source: Slides p. 33–35
Q6 · MAE and MSE calculation
A model predicts 400-m times (s) for five athletes. What are the MAE and MSE?
actual <- c(50, 52, 48, 55, 60)
predicted <- c(48, 54, 47, 56, 54)Show answer
Answer: D. Residuals (actual − predicted) are 2, −2, 1, −1 and 6, so MAE = 12/5 = 2.4 s and MSE = (4 + 4 + 1 + 1 + 36)/5 = 46/5 = 9.2 s². Option B squares the MAE (2.4² = 5.76) instead of averaging the squared residuals, A keeps the signs so they cancel, and C forgets to divide by 5. The single 6-s miss gives 36 of the 46: MSE punishes large errors more.
Source: Slides p. 41–42
Q7 · MAE and MSE: four runners
A model predicts the 10-km race times of four runners (in minutes). What are the MAE and the MSE?
actual <- c(40, 45, 50, 55)
predicted <- c(39, 48, 49, 58)Show answer
Answer: C. Residuals (actual − predicted) are +1, −3, +1, −3, so MAE = (1 + 3 + 1 + 3)/4 = 2 min and MSE = (1 + 9 + 1 + 9)/4 = 5 min². Option A squares the MAE (2² = 4) instead of averaging the squared residuals; B keeps the signs, and D forgets to divide by 4.
Source: Slides p. 41–42
Q8 · Overfitting with a small sample
On the least-squares slides, a new line Size = 0.4 + 1.3 × Weight passes exactly through the two red training points (sum of squared residuals =
0) but has a large sum of squared residuals for the green testing points. What does this show?
Show answer
Answer: B. Fitting two training points perfectly and failing on new data is overfitting, the problem ridge regression is meant to fix; a line with zero training error is not too simple (A), and minimising training residuals does not guarantee good test performance (C). Exam note: slide p. 49 literally says the line 'is overfit to the testing data'; it is overfitted to the training data and therefore fits the testing data badly.
Source: Slides p. 48–50
Q9 · Ridge regression calculation
Two training points and λ = 1. Line 1 (least squares) has slope 1.5 and residuals 0 and 0. Line 2 has slope 0.5 and residuals 0.4 and 0.2. The ridge total is the sum of squared residuals + λ × slope². Which statement is correct?
Show answer
Answer: C. Line 1 costs 0 + 1 × 1.5² = 2.25 and Line 2 costs 0.4² + 0.2² + 1 × 0.5² = 0.16 + 0.04 + 0.25 = 0.45; the lower total wins, so ridge prefers the flatter line (the same logic as the slide's 1.69 versus 0.74). Option A ignores the penalty and B uses |slope|, which is the LASSO penalty.
Source: Slides p. 51–52
Q10 · Ridge: the role of λ
Which statement about λ in ridge regression is correct?
Show answer
Answer: A. With no penalty (λ = 0) ridge is least squares, and a larger λ flattens the slope, so predictions become less sensitive (C is reversed). Training error is always lowest at λ = 0, so it cannot be used to choose λ (B), and setting slopes exactly to zero is LASSO's property, not ridge's (D). Exam note: the slide's 'lowest residuals for test data' means the validation folds inside cross-validation, not the final test set.
Source: Slides p. 53–59
Q11 · Ridge versus LASSO
You have 40 rowers and 25 candidate predictors of their 2000-m rowing-machine time, and you expect only a handful of predictors to matter. Which choice best follows the lecture's summary?
Show answer
Answer: D. The summary slide says LASSO outperforms ridge when many unnecessary features are included, because their slopes can become exactly 0; ridge shrinks slopes but never removes predictors (A). Ridge and LASSO are both more robust to overfitting than least squares in small data sets (B), and ridge with λ = 0 is simply least squares (C).
Source: Slides p. 58, 60
Q12 · Holdout and tuning
You fit LASSO models with 20 values of λ, compute each model's error on the test set, pick the λ with the lowest test error, and report that error as your final performance. What is the problem?
Show answer
Answer: B. The holdout slide says to train and tune with cross-validation on the training set and not to touch the test set 'until the very end'; once you pick λ on the test data, those data are no longer unseen, so the reported error looks better than it really is. Choosing on training error (C) would always pick λ = 0.
Source: Slides p. 43–45, 57
Q13 · Choosing and evaluating a regression model 4 pointscatch-up notes
You predict a numerical performance score for 45 athletes. You fit least squares and LASSO on the training athletes (λ chosen with cross-validation), then score both once on the locked test set (9 athletes) and compare them with a naive baseline that always predicts the training mean. State and motivate briefly, using the output below:
1) Which model would you report? (1 point)
2) What do the least-squares results and the baseline tell you? (2 points)
3) Why should you still be careful with the test result? (1 point)
Model MAE training MAE test
Least squares 1.2 4.6
LASSO 2.3 2.8
Naive baseline 3.9 4.0Show model answer
1) LASSO: it has the lowest error on the test set (2.8 versus 4.6), and predicting new athletes is what matters; the lower training error of least squares does not count. 2) Least squares fits the training athletes much better than the test athletes (1.2 versus 4.6): it is overfitted. On the test set it is even worse than the naive baseline (4.6 versus 4.0), so it adds nothing; LASSO clearly beats the baseline (2.8 versus 4.0), so its predictors add real predictive value. 3) The test set holds only 9 athletes, so the result depends heavily on which athletes ended up in it; a different split, or data from another lab (external validation), could give a different MAE.
How the points are earned
- 1 pt Report LASSO, because it has the lowest test error
- 1 pt Least squares: large gap between training and test error means overfitting
- 1 pt Least squares does no better than the baseline on the test set while LASSO beats it, so only LASSO adds predictive value
- 1 pt Small test set (9 athletes): the result depends on the split or needs external validation
Source: Slides p. 27, 41–45, 57–60; Recording 83:50–92:30
How did your answer compare?
Q14 · Ridge versus LASSO 4 points
A sports scientist wants to predict 2000-m rowing-machine time in 35 rowers from 20 candidate predictors, and considers ridge or LASSO regression instead of least squares. State and motivate briefly:
1) What do ridge and LASSO add to the least-squares rule, and what does λ control? (2 points)
2) How do ridge and LASSO differ in what they do to the slopes, and when would you prefer each? (2 points)
Show model answer
1) Both keep the sum of squared residuals but add a penalty for steep slopes: ridge adds λ × slope², LASSO adds λ × |slope|. They accept a slightly worse fit on the training data in return for flatter, less sensitive slopes, which makes them more robust to overfitting in small data sets. λ is the penalty strength: λ = 0 gives plain least squares and a larger λ gives flatter slopes; choose λ with cross-validation on the training data. 2) Ridge shrinks all slopes towards zero but keeps every predictor. LASSO can set slopes exactly to zero, which removes those predictors from the model. Prefer LASSO when many predictors are probably useless (likely here, with 20 candidates for 35 rowers), and ridge when most predictors are meaningful, for example a few measures chosen because physiology says each one matters.
How the points are earned
- 1 pt Both add a penalty on the slopes to the sum of squared residuals: ridge λ × slope², LASSO λ × |slope|
- 1 pt λ is the penalty strength: λ = 0 gives least squares, a larger λ gives flatter slopes; chosen with cross-validation
- 1 pt Ridge shrinks slopes but keeps all predictors; LASSO can set slopes exactly to zero and so removes predictors
- 1 pt LASSO when many predictors are unnecessary, ridge when most are meaningful
Source: Slides p. 50–60
How did your answer compare?
Q15 · Data leakage: splitting repeated measuresfrom the recording
You have 30 training sessions from each of 12 cyclists (360 rows) and want to predict session RPE (how hard the athlete rated the session) from workload variables. How should you create the test set, following the lecture?
Show answer
Answer: C. A random row split puts the same cyclist in both sets, so information leaks into training ('data leakage') and performance looks better than it would be for new people; the same rule applies to cross-validation folds. A random row split (A) is fine only when each person has a single row. Exam note: splitting by athlete and the term 'data leakage' come from the recording, not the slides.
Source: Recording 90:30–92:30; Slides p. 44
Q16 · Unsupervised learning generates hypothesesfrom the recording
A researcher gives a clustering algorithm only the 100-m times of a group of sprinters and long jumpers and asks for two clusters. Both clusters contain a mix of sprinters and long jumpers. According to the lecture, what is the best interpretation?
Show answer
Answer: A. Clustering gets no labels and finds groups itself; mixed clusters suggest discipline does not define sprint speed, which raises a new question, the lecturer's point that unsupervised learning generates hypotheses and supervised learning tests them. The algorithm never saw the labels, so this was not classification (C), and an exploratory clustering proves nothing (D).
Source: Recording 31:30–33:30; Slides p. 13
Q17 · Encoding categorical variablesfrom the recording
In a football data set with players from 100 countries, a student codes country as one number column (Italy = 1, …, United States = 100) and uses it to predict salary. What is the main problem the lecturer pointed out?
Show answer
Answer: D. A single number column implies an order and spacing that countries do not have, so a 'bigger' country number could appear to raise salary; 0/1 indicator columns remove that implied order. Option B describes one-hot encoding with many categories, a separate issue the lecturer solved by binning countries into continents. Exam note: on the slides one-hot = one column per category and dummy = one column fewer; the lecturer said it the other way round in the recording.
Source: Recording 74:30–77:30; Slides p. 32–35
Q18 · Shallow ML versus deep learningfrom the recording
Why did the lecturer advise students to use shallow (classical) machine learning rather than deep learning for their course data sets?
Show answer
Answer: B. The performance-versus-data curve shows both approaches performing about the same with little data, and only deep learning keeps improving as data grow; deep learning also extracts features itself, which costs interpretability and context, so the advice was to 'keep it simple'. Option C reverses the curve and D reverses the feature-extraction contrast.
Source: Recording 46:00–55:45; Slides p. 20–23
Lecture 6 · Classification and clustering
21 questions · lecture page
Q1 · Precisioncatch-up notes
Injured is the positive class. A classifier has TP = 12, FP = 3, FN = 8 and TN = 77. What is its precision?
Show answer
Answer: B. Precision = TP/(TP + FP) = 12/15 = 80%: of the players the model flagged as injured, 80% really were injured. 60% is recall, TP/(TP + FN) = 12/20, which starts from the players who really got injured; 89% is accuracy, (12 + 77)/100.
Source: Slides p. 26–31
Q2 · Bias–variance diagnosiscatch-up notes
A model that classifies training sessions as interval or endurance has a training error of 3% and a validation error of 15% (the validation error is measured on data kept apart from learning). What is the main problem?
Show answer
Answer: A. The lecture reads bias from the training error (3%, low) and variance from the gap between validation and training error (15 − 3 = 12 points, large): low bias with high variance is overfitting. Underfitting (high bias) would also show a high training error, as in the lecture's 15% vs 16%.
Source: Slides p. 9, 15–16
Q3 · Decision trees: choosing the first splitcatch-up notes
You want a decision tree that predicts whether a runner gets a stress fracture (yes/no) from previous injury (yes/no), sex and age in years. How does the tree choose its first question (the root)?
Show answer
Answer: C. Lower Gini means purer child groups (0 = only one class), and for a number such as age the candidate cut-offs are the means of neighbouring sorted values; the split with the lowest size-weighted Gini wins. Option A is the common trap: the highest Gini means the most mixed groups, which is the worst split.
Source: Supplementary slides p. 18–21
Q4 · Reading R output: the positive classcatch-up notes
A model predicts whether a youth player gets injured this season (Yes/No). Below is R's confusionMatrix() output for the test set; check which class R treats as positive. Which value is the share of the players who really got injured (Yes) that the model also predicted as injured?
Confusion Matrix and Statistics
Reference
Prediction No Yes
No 54 8
Yes 6 32
Accuracy : 0.86
95% CI : (0.7763, 0.9213)
No Information Rate : 0.6
P-Value [Acc > NIR] : 1.293e-08
Kappa : 0.7059
Mcnemar's Test P-Value : 0.7893
Sensitivity : 0.9000
Specificity : 0.8000
Pos Pred Value : 0.8710
Neg Pred Value : 0.8421
Prevalence : 0.6000
Detection Rate : 0.5400
Detection Prevalence : 0.6200
Balanced Accuracy : 0.8500
'Positive' Class : NoShow answer
Answer: B. R took 'No' as the positive class, so Sensitivity (0.90 = 54/60) is the share of real 'No' players predicted 'No'. Here the injured players are the negative class, so their recognition rate is Specificity = 32/(32 + 8) = 0.80. Accuracy (0.86) mixes both classes, and balanced accuracy (0.85) is the average of 0.90 and 0.80.
Source: Slides p. 26–33; Recording 20:35
Q5 · Reading R output: class imbalance
A model predicts overtraining (Yes/No) in 200 cyclists; overtraining is rare. Below is R's confusionMatrix() output for the test set. Which conclusion is most appropriate?
Confusion Matrix and Statistics
Reference
Prediction Yes No
Yes 6 4
No 14 176
Accuracy : 0.91
95% CI : (0.8615, 0.9458)
No Information Rate : 0.9
P-Value [Acc > NIR] : 0.3724
Kappa : 0.3571
Mcnemar's Test P-Value : 0.03389
Sensitivity : 0.3000
Specificity : 0.9778
Pos Pred Value : 0.6000
Neg Pred Value : 0.9263
Prevalence : 0.1000
Detection Rate : 0.0300
Detection Prevalence : 0.0500
Balanced Accuracy : 0.6389
'Positive' Class : YesShow answer
Answer: D. 90% of the cyclists are 'No', so always predicting 'No' already scores 0.90 accuracy (the no-information rate). With 'Yes' as positive, only 6 of 20 overtrained cyclists are caught (sensitivity 0.30), and balanced accuracy (0.30 + 0.98)/2 = 0.64 shows the weakness. Specificity is about non-overtrained cyclists correctly cleared, not about missed overtraining.
Source: Slides p. 26–33; Recording 20:35
Q6 · Choosing a metric for a rare outcome
In a running club, about 4% of the young runners get a stress fracture in a season. You build a model that predicts stress fracture (yes/no). Which metric should you use to evaluate it?
Show answer
Answer: C. This is classification with imbalanced classes: accuracy is dominated by the 96% easy 'no' cases, so it looks excellent even if no fracture is found. F1 (the harmonic mean of precision and recall) and balanced accuracy (the average of recall and specificity) expose that. MAE is a regression metric for a numeric target.
Source: Slides p. 27–33; Recording 20:35
Q7 · Fixing high bias
Your model has a training error of 14% and a validation error of 15%. Following the lecture's bias–variance flowchart, which action is most likely to help?
Show answer
Answer: B. A high training error with a small gap (1 point) means high bias (underfitting). The flowchart says: learn the training set better with a more complex model, more or better features, longer training or less regularization. More data, fewer features and more regularization are fixes for high variance. Exam note: the slide's 'Total error = bias + variance' is a simplification (strictly bias² + variance + noise); it does not change this diagnosis.
Source: Slides p. 9, 12–16
Q8 · Validation design in caret
This lecture code classifies training sessions as interval or endurance. Which statement is correct?
train_id <- createDataPartition(model_data$training_type, p = 0.7, list = F, times = 1)
train_data <- model_data[ train_id, ]
test_data <- model_data[-train_id, ]
fitControl <- trainControl(method = 'cv', number = 10, classProbs = T)
model <- train(training_type ~ ., data = train_data,
method = 'rf',
trControl = fitControl,
verbose = F,
metric = 'ROC')Show answer
Answer: C. p = 0.7 puts 70% of the sessions in train_data, and the 10-fold cross-validation runs only inside train_data; test_data is kept aside for the final check (the lecture's 'good practice'). In k-fold cross-validation every observation is validated once and used for training k − 1 times; leave-one-out would need k = the number of observations.
Source: Slides p. 20–22, 24–25
Q9 · Decision trees: splitting a numerical feature
To split on the numerical feature age, you follow the lecture's procedure: sort by age, use the mean of each pair of consecutive ages as a candidate threshold, and compute the weighted Gini impurity for each. Which threshold is chosen, and what is its weighted Gini?
Age: 15 17 20 22 26
Injured: No No Yes Yes NoShow answer
Answer: A. The candidate cut-offs are the midpoints 16, 18.5, 21 and 24. At 18.5 the left group {No, No} has Gini 0 and the right group {Yes, Yes, No} has 1 − (2/3)² − (1/3)² = 0.444, so the weighted Gini is 3/5 × 0.444 = 0.27, the lowest (16 and 24 give 0.40, 21 gives 0.47). Option D uses 20, an observed age, not a midpoint. Exam note: the supplementary slide's 'Gini impurity Age < 31.5 = 0.1905' is only half of the weighted sum (correct: 0.405); learn the procedure, not that number.
Source: Supplementary slides p. 18–21
Q10 · Elbow and silhouettecatch-up notes
You run k-means on standardised (z-scored) fitness data of 60 athletes for k = 1–8. Below are the values of the elbow plot (total within-cluster sum of squares) and of the silhouette plot (average silhouette score). What does each method suggest?
k 1 2 3 4 5 6 7 8
Within-cluster SS 900 560 330 150 132 118 107 99
Average silhouette – 0.38 0.44 0.51 0.53 0.60 0.49 0.42Show answer
Answer: D. The sum of squares always falls as k grows, so look for the bend after which extra clusters help little: the drops are 340, 230 and 180, then 18 or less, so the elbow is at k = 4. For the silhouette higher is better, and it peaks at k = 6. Option A picks the lowest sum of squares, which would always be the largest k.
Source: Slides p. 62–64, 82–83
Q11 · PCA: explained variancecatch-up notes
A PCA on four standardised fitness tests gives eigenvalues (the amount of variance along each component) of 12, 5, 2 and 1. Which statement is correct?
Show answer
Answer: A. Share of variance = eigenvalue / sum of all eigenvalues: 12/20 = 60% and 5/20 = 25%, together 85%. Option C divides by 12 + 5 only and forgets PC3 and PC4; explained variance is not a classification accuracy. Exam note: supplementary p. 60 writes '18/18+4 = 0.81 → 81%' and '17%'; the correct values are 81.8% and 18.2%.
Source: Supplementary slides p. 60–61; Slides p. 71
Q12 · PCA: purposecatch-up notes
A sports scientist has 40 fitness and GPS variables for each of 300 players. What is principal component analysis (PCA) mainly used for in this situation?
Show answer
Answer: D. PCA is unsupervised dimension reduction: PC1 follows the direction of largest spread, PC2 is perpendicular and catches most of what is left, and so on. It uses no target and assigns no clusters, although you can cluster on the PC scores afterwards; PC1 is fitted with perpendicular distances (like Deming regression), not vertical ones.
Source: Slides p. 70–74; Supplementary slides p. 54–61
Q13 · Case study: clustering cyclists
In van der Zwaard et al. (2019), elite track sprinters, team-pursuit cyclists and road cyclists were clustered with k-means on anthropometry (body size, composition and shape). Which result did the lecture report?
Show answer
Answer: B. The meso cluster held the 6 sprinters, and the tall and short meso-ecto clusters each held 4 pursuit and 5 road cyclists. This confirmed anthropometry-dependent specialization for sprint versus endurance but did not separate pursuit from road. Discipline was examined only after clustering, and the meso cluster showed higher sprint performance while the meso-ecto clusters showed higher endurance performance.
Source: Slides p. 77–79, 85–97
Q14 · Classification versus clustering 3 pointscatch-up notes
A sports scientist says: 'Classification and clustering both put athletes into groups, so they are the same thing.' State and motivate briefly:
1) the key difference between a classification model and a clustering model;
2) one sport or health example of a classification problem;
3) one example of a clustering problem.
Show model answer
1) Classification is supervised: it learns from examples whose correct category (label) is known and predicts that category for new athletes. Clustering is unsupervised: there are no labels; the algorithm (e.g. k-means) forms groups of athletes that are similar to each other and different from the other groups. 2) Predicting injured yes/no from preseason tests in youth football (Rommers et al.), or classifying a session as interval or endurance. 3) Grouping 24 elite cyclists by body measurements without telling the algorithm their discipline (van der Zwaard), or finding athlete profiles in fitness data.
How the points are earned
- 1 pt Difference: classification learns from known labels (supervised); clustering has no labels and finds the groups itself (unsupervised)
- 1 pt Valid classification example with a categorical target (e.g. injured yes/no)
- 1 pt Valid clustering example without labels (e.g. grouping cyclists by body build)
Source: Slides p. 4–7, 56–58
How did your answer compare?
Q15 · PCA: purpose and loadings 2 pointscatch-up notes
A researcher has 12 body measurements of 80 rowers and runs principal component analysis (PCA). PC1 explains 60% of the variance; its largest weights are 0.45 for body height and 0.41 for arm span, and the skinfolds have small weights (below 0.15). State and motivate briefly:
1) what PCA is used for here;
2) what these weights are called and what they tell you about PC1.
Show model answer
1) Dimension reduction: PCA replaces the 12 measurements by a few new perpendicular axes (principal components) that keep as much of the variance as possible, so the rowers can be plotted in 2-D or clustered on fewer variables. It is unsupervised: no target is used. 2) They are loadings: the weight of each original variable in the component. Height and arm span have the largest loadings, so PC1 mainly describes body size (length); the skinfolds contribute little to PC1.
How the points are earned
- 1 pt Dimension reduction: a few new axes that keep most of the variance (to plot or cluster); no target used
- 1 pt Loadings = weights; the larger the loading, the more the variable contributes, so PC1 ≈ body size/length
Source: Slides p. 70–74; Supplementary slides p. 56–61
How did your answer compare?
Q16 · Case study: injury prediction (Rommers et al., 2020) 3 points
Rommers et al. (2020) used XGBoost (a model built from many decision trees) on 29 preseason test results of 734 youth football players to predict who would get injured (yes/no). Precision, recall and F1 were 84/83/83% on the training set and 85/85/85% on the test set. State and motivate briefly:
1) what type of machine-learning problem this is;
2) whether the model over- or underfits;
3) whether the top feature in the SHAP plot, age at peak height velocity, can be called a cause of injury.
Show model answer
1) Supervised classification: the injury label (yes/no) is known for every player, and the target is a category. 2) Neither: training and test scores are almost equal (84/83/83 vs 85/85/85) and fairly high. Overfitting would show much better training than test scores; underfitting would show low scores on both. 3) No. SHAP shows which features the model relies on most for its predictions, not what causes injury; a strong predictor can be linked to injury through other factors (e.g. growth or exposure).
How the points are earned
- 1 pt Supervised classification, because the target is a known category (injured yes/no)
- 1 pt No over- or underfitting: training ≈ test scores, and both are high
- 1 pt SHAP shows predictive importance, not causation
Source: Slides p. 37–54; Recording 28:45–34:48
How did your answer compare?
Q17 · k-means in practice 3 points
You want to group 24 cyclists by height (cm), body mass (kg) and sum of skinfolds (mm) using k-means. State and motivate briefly:
1) why you first convert the variables to z-scores;
2) the steps of the k-means algorithm.
Show model answer
1) k-means groups by Euclidean distance, so a variable with large numbers or a large spread (height in cm) would dominate the distances. Z-scores (value minus mean, divided by the SD) give every variable mean 0 and SD 1, so all variables count equally. 2) Choose k and place k starting centroids at random (e.g. k random cyclists). Assign each cyclist to the nearest centroid. Move each centroid to the mean of the cyclists assigned to it. Repeat assigning and updating until the assignments no longer change. Because the start is random, run it from several starts and keep the solution with the lowest within-cluster sum of squares.
How the points are earned
- 1 pt Z-scores: k-means uses distances, so variables with bigger units or spread would dominate; z-scores make all count equally
- 0.5 pt Choose k and place k (random) starting centroids
- 0.5 pt Assign each cyclist to the nearest centroid
- 0.5 pt Move each centroid to the mean of its cluster
- 0.5 pt Repeat until the assignments stop changing
Source: Slides p. 58–67; Supplementary slides p. 41–43; Recording 54:23
How did your answer compare?
Q18 · Underfittingfrom the recording
Which model did the lecturer describe as the most extreme ('ultimate') example of underfitting?
Show answer
Answer: B. The naive baseline has no complexity at all, so it has maximal bias and performs poorly on both training and test data; your model has to beat it. A model through every point is the overfitting extreme, and 1% vs 11% error indicates overfitting (high variance).
Source: Recording 02:01–04:03
Q19 · Splitting on the targetfrom the recording
You want to classify sessions as extensive or intensive interval training. A random 70/30 holdout split happens to put almost all extensive sessions in the test set. What was the lecturer's advice?
Show answer
Answer: D. The model cannot learn to predict a class that is (almost) absent from the training data, so you arrange or partition on the target so that all classes, or all RPE values, are represented in both sets. Leave-one-out still splits the data (N times) and is meant for very small datasets.
Source: Recording 12:17; Slides p. 25 (code partitions on training_type)
Q20 · Standardising inputs for k-meansfrom the recording
Before k-means clustering of cyclists on height (cm), body mass (kg) and sum of skinfolds (mm), why did the lecturer prefer z-scores over dividing each variable by its maximum?
Show answer
Answer: A. k-means uses Euclidean distances on the raw numbers, so a variable with larger values or a larger spread dominates. Dividing by the maximum fixes the units but leaves different spreads, whereas z-scoring equalises the SDs. Both are linear rescalings that keep the athletes' order, and neither reduces the number of variables.
Source: Recording 54:23–56:26; Slides p. 66–67
Q21 · Case study: Rommers et al. (2020)from the recording
In Rommers et al. (2020), 50% of players got injured. Test precision, recall and F1 were all 85%, and training values were 84%, 83% and 83%. Which interpretation matches the lecture?
Show answer
Answer: C. With a 50/50 split, accuracy is not inflated by a majority class, and all metrics told the same story. Overfitting would show much better training than test performance, but here the difference is negligible, so 'they made a pretty good model'. Exam note: balance alone does not force precision = recall (that requires FP ≈ FN), but it removes the main reason to distrust accuracy.
Source: Recording 28:45–30:45; Slides p. 47
Lecture 7 · Feature operations
17 questions · lecture page
Q1 · Centering vs standardizationcatch-up notes
Before modelling, you subtract the mean from each predictor (centering) but you do not divide by its standard deviation (SD). What does this achieve?
Show answer
Answer: A. Subtracting the mean only shifts all values so the mean becomes 0; a feature in years still varies far more than one in metres and can still dominate a distance-based model such as KNN. Option B describes full z-scoring, which also divides by the SD.
Source: Slides p. 4–5
Q2 · Missing data: interpolationcatch-up notes
A heart-rate sensor records one value per second. One reading is missing between two valid readings. Which approach estimates the missing value from the readings just before and after it?
Show answer
Answer: C. Interpolation estimates a gap from its neighbours in an ordered series such as time, which suits a smooth signal sampled every second; like all imputation, the result is a best guess, not a measurement. Dummy encoding, binning and z-scoring change values that exist but do not fill a gap. Exam note: interpolation is not named on slide p. 15, but the official 2024 practice exam lists it as a way to handle missing values.
Source: Slides p. 15; 2024 practice exam Q27
Q3 · Scaling and KNN
In the lecture example, KNN (k =
2) classified students as male or female from height (training range 1.50–1.96 m) and age (training range 21–38 years). Without any scaling the test accuracy was only 0.46. What is the most likely explanation?
Show answer
Answer: D. KNN only compares distances, so the feature with the biggest numbers dominates: an age gap of 17 years outweighs a height gap of 0.45 m, and height, the feature that best separates the sexes, is almost ignored. Option A is wrong because KNN has no normality requirement. Exam note: the summary table gives min–max 0.85 but the min–max code printed 0.769; with 13 test students one student moves accuracy by about 0.08, so learn the lesson (scaling matters a lot), not an exact ranking.
Source: Slides p. 4, 6–12
Q4 · Z-score and min–max calculation
In a training set, height has mean 170 cm, SD 10 cm, minimum 150 cm and maximum 200 cm. A new player is 165 cm tall. Using these training values, what are the player's z-score and min–max value?
Show answer
Answer: B. z = (165 − 170) ÷ 10 = −0.5 (half an SD below the mean) and min–max = (165 − 150) ÷ (200 − 150) = 15 ÷ 50 = 0.3. Option A drops the minus sign, but a value below the mean always has a negative z-score; option C forgets to divide.
Source: Slides p. 5; numbers invented
Q5 · Standardization vs normalization
A sprint-speed feature contains a few extreme values caused by GPS errors. According to the lecture's cheat sheet (its summary slide), why is min–max normalization more affected by these outliers than z-score standardization?
Show answer
Answer: C. One extreme value sets the maximum, so with min–max it becomes 1 and everyone else is pushed close to 0; z-scores have no fixed range, so an outlier just gets a large z. Option A is wrong because neither method removes outliers. Exam note: the slides say standardization aims to 'acquire a normal distribution'; strictly, z-scoring keeps the shape of the distribution, so a skewed feature stays skewed.
Source: Slides p. 5, 13
Q6 · Log transformation
Training-load values 1, 2, 4 and 8 (each double the previous one) are log₂-transformed. Which statement describes the result?
Show answer
Answer: D. log₂ of 1, 2, 4 and 8 is 0, 1, 2 and 3, so every doubling (2-fold change) becomes the same step, which is why log suits right-skewed data. Option A is what min–max scaling would give, not log. Exam note: slide p. 5 files log under normalization ('fixed range e.g. 0–1'), but a log has no fixed range, needs positive values and changes the shape of the distribution.
Source: Slides p. 5, 10–11
Q7 · Imputation with mice
The lecture filled in the missing values of the nhanes data (columns age, bmi, hyp, chl) with the code below. Which statement is correct?
my_imp = mice(input_data, m = 5,
method = c("", "pmm", "logreg", "pmm"),
maxit = 20)
final_clean = complete(my_imp, 5)Show answer
Answer: B. m sets the number of completed data sets and complete(my_imp, 5) returns number 5; the 5 is an ID, not a quality score. Option A confuses maxit (the number of fill-in rounds) with m; also, "logreg" is used for the yes/no variable hyp and "" means age is not imputed at all. Exam note: the slide says 'choose the best matching imputation version'; strictly, multiple imputation analyses all m versions and combines the results.
Source: Slides p. 16
Q8 · Neural networks vs simple models
A physiotherapist has data from 40 patients and wants a model whose link between the predictors and recovery time she can explain to patients. Based on the lecture's comparison, why is a simple model (e.g. linear regression) a better choice here than a neural network?
Show answer
Answer: A. The slide: a simple model gives an interpretable relationship and simple control of bias and variance, while a network fits complex functions at the cost of interpretability, more data and more careful bias–variance control. Option B reverses the slide: networks learn relationships between inputs automatically, whereas with simple models you must handle them yourself.
Source: Slides p. 18
Q9 · Odds ratio calculation
In one season, 20 of 100 footballers with a high training load and 10 of 100 footballers with a low training load got injured. What is the odds ratio (OR) of injury for high versus low load?
Show answer
Answer: D. Odds = injured ÷ not injured: 20 ÷ 80 = 0.25 (high) and 10 ÷ 90 (low), so OR = (20 × 90) ÷ (80 × 10) = 1800 ÷ 800 = 2.25. 2.0 is the relative risk (20% ÷ 10%), which divides risks, not odds; 0.44 is the OR turned upside down and 0.10 the risk difference.
Source: Slides p. 21–22, 24; numbers invented
Q10 · Probability, odds and odds ratio
Which statement about probability, odds and the odds ratio (OR) is correct?
Show answer
Answer: B. Odds = p ÷ (1 − p) = 0.75 ÷ 0.25 = 3, and a ratio of 1 means no difference between the groups. Option C is tempting, but an OR of 0.5 halves the odds, not the probability; odds range from 0 to infinity (A) and likelihood is how well a model explains the observed data (D). Exam note: slide p. 24 says 'Odds = 0: no difference between groups'; the lecturer corrected this aloud to 1. The same slide's 'likelihood, e.g. R²' is loose: R² is not strictly a likelihood.
Source: Slides p. 21, 24
Q11 · Absolute vs relative risk
A headline says a supplement 'doubles the risk' of a rare side effect. In the study, 2 of 10,000 users and 1 of 10,000 non-users had the side effect. Which statement reports the result the way the lecture recommends?
Show answer
Answer: A. Absolute risk = 2 ÷ 10,000 = 0.02% versus 1 ÷ 10,000 = 0.01%, so RR = 2 but the absolute increase is only 0.01 percentage points: 1 extra case per 10,000 users. The slide's rule is never to report a relative risk without the absolute risk; option B confuses a 100% relative increase with percentage points.
Source: Slides p. 22–23; numbers invented
Q12 · Missing data strategies 3 pointscatch-up notes
In a data set of 200 youth footballers, the column 'sprints per match' is empty for 30 players. For 12 of them you find out that the GPS export writes NaN whenever a player made no sprints. State and motivate briefly:
1) What should you do with these 12 NaN values? (1 point)
2) Name two ways to deal with the remaining 18 missing values and give one drawback of each. (2 points)
Show model answer
1) Replace them with 0. The true value is known to be 0, so this is correcting a wrong value (the lecture's option 3), not guessing. 2) Any two, each with a drawback: delete all rows with a missing value (loss of information, and a biased sample if the missing players differ from the rest); delete only rows missing an essential variable (still loses data, and it is hard to decide what is essential); impute a best guess, e.g. the mean or with mice (the filled-in value is a guess, not a measurement; mean replacement also shrinks the spread).
How the points are earned
- 1 pt Replace NaN with 0 because the true value is known to be 0 (correcting a wrong value)
- 1 pt First valid method with a correct drawback (e.g. delete rows: loss of information or biased sample)
- 1 pt Second valid method with a correct drawback (e.g. impute with mean or mice: a guess, not a measurement; less spread)
Source: Slides p. 15–16; Recording 20:55
How did your answer compare?
Q13 · Odds ratio vs relative risk 3 points
A study selected 45 patients with heart disease and 45 people without heart disease and asked whether they smoke: 31 of the patients and 6 of the people without heart disease smoked (OR = 14.4). A journalist writes: 'Smokers are 14 times more likely to get heart disease.' State and motivate briefly:
1) What is wrong with the journalist's sentence? (1 point)
2) Why can this study report an odds ratio but not a relative risk? (1 point)
3) Write one correct sentence about the result. (1 point)
Show model answer
1) 14.4 is an odds ratio: it compares odds (people with the outcome ÷ people without it), not probabilities. 'Times more likely' is a statement about risk, and when the outcome is common the OR is much larger than the risk ratio. 2) The researchers chose people by disease status (45 with, 45 without: a case–control study). The share with heart disease is fixed by that choice, so the risk for smokers, and therefore an RR, cannot be calculated. The OR does not change with how many people with and without the disease were chosen. 3) 'The odds of heart disease were about 14 times higher in smokers than in non-smokers (OR = 14.4).' This shows an association, not proof that smoking causes heart disease.
How the points are earned
- 1 pt The OR compares odds, not probabilities/risk, so 'times more likely' overstates it
- 1 pt Case–control design: people were selected by outcome, so risk and RR cannot be calculated
- 1 pt Correct sentence: the odds of heart disease are about 14 times higher in smokers
Source: Slides p. 21–24
How did your answer compare?
Q14 · Scaling in the lifecyclefrom the recording
According to the lecturer, in which phase of the data science lifecycle does a log transformation of a right-skewed feature belong?
Show answer
Answer: B. The lecturer's rule: if only the scale changes (z-score, min–max, metres to centimetres) it is data preparation; once the shape of the distribution changes too, as with a log transformation, it is feature engineering. Option A is the rule for pure rescaling, which does not cover a log.
Source: Recording 00:00
Q15 · Neural network nodefrom the recording
A single node of a neural network decides whether an athlete trains today. Inputs: slept well x1 = 1, no muscle soreness x2 = 1, rain x3 = 0. Weights: w1 = 2, w2 = 3, w3 = 4. Threshold = 6. As in the lecture's surfing example, the node outputs 1 if the weighted sum minus the threshold is greater than 0, otherwise 0. What does the node compute?
Show answer
Answer: D. Multiply each input by its weight (1×2 + 1×3 + 0×4 = 5) and subtract the threshold: 5 − 6 = −1, which is not above 0, so the output is 0. Option A wrongly counts w3 although x3 = 0; an input of 0 contributes nothing.
Source: Recording 35:11–37:23
Q16 · Missing data optionsfrom the recording
In a match-statistics data set, the column 'goals scored' shows NaN for every player who did not score in a match. Which approach matches the lecturer's advice?
Show answer
Answer: A. The slide's option 3, 'change values that are wrong (e.g. 0 instead of NaN): great, do that!', was explained with exactly this case: a NaN that in truth is 0 is set to 0. Imputation (C, D) is a best guess for values you cannot recover, and here it would invent goals that were never scored.
Source: Recording 20:55; Slides p. 15
Q17 · Multiple imputationfrom the recording
You filled in missing values with mice using m = 5. According to the lecturer, what could you learn by running your model on each of the five completed data sets?
Show answer
Answer: C. Each version has slightly different filled-in values, so comparing model results across them shows how much the guesses matter; the lecturer called this out of scope and allowed picking one version for the project. Option B confuses m (number of completed data sets) with maxit (number of fill-in rounds), and accuracy cannot reveal which guesses are 'true' (A).
Source: Recording 25:01–29:04; Slides p. 16
Lecture 8 · Epidemiological data
16 questions · lecture page
Q1 · Exposure data choicecatch-up notes
You study the link between long-term (10-year) NO₂ exposure and death from any cause in adults across the Netherlands. NO₂ comes mostly from traffic and changes strongly over short distances. Which exposure data fit this question best?
Show answer
Answer: C. NO₂ changes a lot within metres of a road, so each person needs their own long-term value: the lecture used the 10-year average at the home address, predicted with LUR where nothing was measured. A national mean or one central station (A, B) gives everyone nearly the same value, so there is no difference between people to compare.
Source: Slides p. 17, 21–22, 33, 40
Q2 · Hazard, exposure and risk
A cyclist commutes for 2 hours a day along a busy road with high NO₂ levels. Using the lecture's definitions, which combination is correct?
Show answer
Answer: B. Slide p. 8: a hazard has the potential to harm, exposure is how much and how long you are subjected to it, and risk is the chance it actually causes harm. Option A swaps hazard and exposure: the 2 hours are part of the exposure, while NO₂ is the hazard.
Source: Slides p. 8
Q3 · Risk assessment paradigm
In the risk assessment paradigm, which step examines the probability and severity of a health outcome as a function of the dose (exposure)?
Show answer
Answer: D. Hazard characterization describes how harm depends on the dose. Hazard identification only asks whether the agent can cause harm at all, exposure assessment estimates who is exposed to which levels and for how long, and risk characterization asks what the consequences are at current levels and what reducing exposure would gain.
Source: Slides p. 9
Q4 · Exposure assessment: residential proxy
Why does the lecture's study use the average NO₂ level at each participant's home address instead of true personal exposure?
Show answer
Answer: A. Slide p. 22: true exposure across micro-environments is not feasible to measure for large populations, so the residential average is a proxy (stand-in). Option B is tempting but wrong: the proxy assumes people spend most, not all, of their time at home, so it misses exposure at work or while commuting.
Source: Slides p. 22
Q5 · Exposure measurement sources
A researcher needs measurements of a pollutant that the national network does not measure, at specific sites near schools. Which approach fits, and what is its main drawback according to the lecture?
Show answer
Answer: C. A designed campaign lets you choose the pollutants and locations, but it is costly and cannot cover large populations (slide p. 24). A routine network (A, D) is the opposite: low cost, long-running and continuous, but it may not measure your pollutant or have stations where you need them (p. 26).
Source: Slides p. 23–26
Q6 · Land-use regression (LUR)
In the case study's land-use regression (LUR) model, what is the outcome variable and what are the predictors?
Show answer
Answer: B. LUR is a statistical model: it learns how NO₂ measured at the stations depends on the roads, green space and buildings around them, then applies that rule to unmeasured places such as participants' homes. Option A describes the later Cox health model, and option D describes a deterministic (dispersion) model.
Source: Slides p. 31–40
Q7 · Health data sources
Which health data source does the lecture describe with these pros and cons: 'earlier in the chain; possibly less bias than clinical outcomes' versus 'can be invasive; expensive/time-consuming; instrument/observer effects'?
Show answer
Answer: A. These are the pros and cons of physiological measurements such as biomarkers (slide p. 48). Routine databases are large, cheap and consistent but depend on data quality, completeness, administrative factors and privacy rules (p. 46); remote sensing is an exposure source, not a health data source.
Source: Slides p. 45–48
Q8 · Correlation and confounding
Why does the lecture call the correlation coefficient between NO₂ and mortality 'not an interesting measure of association'?
Show answer
Answer: D. The class answer was 'confounding', shown with ice cream and drownings: both rise in summer, so they correlate without one causing the other. The Cox model therefore adjusts for measured confounders such as smoking, BMI, education and neighbourhood income.
Source: Slides p. 59–61, 68
Q9 · Hazard ratio interpretation
The adjusted Cox model reports HR = 1.04 (95% CI 1.03–1.04), p < 0.001, for a fixed step (increment) in NO₂ exposure. Which interpretation is correct?
Show answer
Answer: A. HR 1.04 multiplies the hazard by 1.04 (4% higher), and the CI excludes 1, so the association is significant. Option B is the classic trap: an HR is a relative measure and says nothing about percentage points of absolute risk; an observational study also cannot prove cause (C). Exam note: the conclusion slide says 'for every 1-unit increase … the mortality risk increases by 4%', but the printed coefficient (0.00392) fits 4% per 10 µg/m³, so always state which step an HR refers to.
Source: Slides p. 58, 63, 66, 69–72
Q10 · Cox model in R
The lecture's Cox model formula is shown below (shortened). Which statement is correct?
Surv(age_b, age_end, mort) ~ NO2_exposure + strata(sex) +
as.numeric(smoke_dur) + as.factor(smoking) + as.factor(bmi_cat) +
as.factor(edulev) + as.numeric(meaninc_neighbor_2011) + ...Show answer
Answer: C. Surv(age_b, age_end, mort) defines the follow-up (age at start and end) and the event (mort = 1 means died); NO2_exposure is the exposure and the other terms are confounders. Option D is wrong: strata(sex) gives men and women their own baseline hazard instead of excluding anyone.
Source: Slides p. 52, 64, 67–68
Q11 · Exposure model vs health model 3 pointscatch-up notes
The NO₂ study used two models: a land-use regression (LUR) model and a Cox model. State and motivate briefly:
1) What does each model predict or estimate, and from which data? (2 points)
2) Give one reason why the final result must be interpreted with care. (1 point)
Show model answer
1) LUR (the exposure model) is fitted on NO₂ measured at 69 monitoring stations, with traffic and land-use features around each station as predictors. It predicts the long-term NO₂ level at unmeasured places, such as each participant's home address. The Cox model (the health model) uses these predicted NO₂ levels, linked to each person's health record, to estimate the association between NO₂ and death during follow-up, adjusted for confounders such as smoking and income. The result is a hazard ratio. 2) One of: home-address NO₂ is only a proxy for what people really breathed; LUR predictions are estimates, so their errors carry into the Cox model; the study is observational, so unmeasured confounders can remain (association, not causation).
How the points are earned
- 1 pt LUR: trained on station measurements with land-use/traffic predictors; predicts NO₂ at unmeasured locations (homes)
- 1 pt Cox: estimates the association (hazard ratio) between NO₂ and mortality, adjusted for confounders
- 1 pt A valid caveat: proxy exposure, prediction error, or observational design / remaining confounding
Source: Slides p. 22, 33–41, 63–69
How did your answer compare?
Q12 · Statistical significance vs relevance 3 points
The study's main result is shown below. The lecturer asked: 'Is there a significant effect of NO₂ on mortality?' and 'Is it a substantial effect?' State and motivate briefly:
1) Is the effect statistically significant? (1 point)
2) Is it a substantial effect? (2 points)
Adjusted Cox model, n = 33,475 adults, 10-year follow-up
NO2 (per fixed increment): HR = 1.04 (95% CI 1.03–1.04), p < 0.001Show model answer
1) Yes: the 95% CI (1.03–1.04) does not include 1 and p < 0.001, so the association is very unlikely to be due to chance alone. With 33,475 people the interval is narrow, so even a small effect becomes significant. 2) Per person it is modest: about 4% higher hazard per step, an HR close to 1. But an HR is not a measure of health burden: whether it is substantial depends on how many people are exposed and to which levels (and on the baseline death rate). A small increase spread over millions of exposed people can still mean many extra deaths.
How the points are earned
- 1 pt Significant: the CI does not include 1 and/or p < 0.001
- 1 pt Per person the effect is small/modest (4% higher hazard, HR close to 1)
- 1 pt Real impact depends on how many people are exposed and to which levels (population health burden)
Source: Slides p. 65–66, 69–72
How did your answer compare?
Q13 · LUR model validationfrom the recording
After fitting a land-use regression (LUR) model for NO₂ on the 69 Dutch monitoring stations, the guest lecturer described the usual way to check that the model predicts well. Which procedure did she describe?
Show answer
Answer: C. She described a hold-out check: build the model on some stations, predict at the stations left out, and see whether the predictions match the measurements. Option B judges the model only on the data it learned from, which rewards overfitting; option A is circular, because the health result cannot check the exposure model.
Source: Recording 26:37–28:40; Slides p. 41
Q14 · Deterministic vs statistical exposure modelsfrom the recording
Which statement correctly distinguishes the two approaches the lecture described for predicting NO₂ at addresses without a monitoring station?
Show answer
Answer: A. Deterministic (dispersion or chemical transport) models follow pollution from known sources with physics and chemistry; the statistical LUR model learns from measured concentrations how roads, green space and houses explain differences between stations. Options B and D swap the two approaches.
Source: Recording 18:25–20:30; Slides p. 31–33
Q15 · Routine health databasesfrom the recording
In the Dutch mortality registry, more deaths are recorded on Mondays than on Sundays. According to the guest lecture, what does this illustrate?
Show answer
Answer: D. The lecturer gave this as an example of the registry drawback 'depends on administrative factors': deaths over the weekend can be registered on Monday. Option B is tempting because weekday NO₂ is higher, but she said explicitly that the Monday excess is not real.
Source: Recording 32:46; Slides p. 46
Q16 · Health data categoriesfrom the recording
A researcher cannot access diagnoses because medical records are too sensitive, but can obtain how much each person spent on specific therapies per year. In the lecture's grouping of health data (Galetsi et al., 2020), what type of data is this, and how can it be used?
Show answer
Answer: B. Costs and reimbursements of medical expenses are administrative data, and the lecturer noted that spending on certain therapies can stand in for clinical information that is too sensitive to access. Option D is tempting, but clinical data are the endpoints themselves (diagnoses, deaths), which this researcher cannot see.
Source: Recording 30:43; Slides p. 44
Lecture 9 · User-generated data
18 questions · lecture page
Q1 · Challenges of user-generated datacatch-up notes
A wearable company has millions of heart-rate measurements and all the analysis software it needs. But the heart-rate data are noisy, and there is no verified record of which users actually got ill. What is the main obstacle to building a useful illness model?
Show answer
Answer: A. The lecture names quality control (noisy and missing data) and the lack of reference data as the key challenges: 'you can have unlimited data and still have no use for it'. Option C is tempting, but a large sample is the strength of user-generated data; more users cannot replace accurate readings or verified outcomes.
Source: Slides p. 76–78, 81, 97, 134
Q2 · APIcatch-up notes
What is an API (application programming interface) in the context of sport and health data?
Show answer
Answer: C. An API lets two programs exchange data without a person typing anything; the running study used the Strava and TrainingPeaks APIs to get users' workouts. Option A is wrong: most wearables report no quality score at all (the lecture's 'no signal quality metric').
Source: Slides p. 26, 81, 98, 101, 113
Q3 · From data to knowledge: three steps
A company wants to use heart-rate variability (HRV) measured with the phone camera to study stress in thousands of users. In which order should the work proceed, according to the lecture?
Show answer
Answer: D. First show that the sensor measures what it claims (garbage in, garbage out), then reproduce a known lab result in the field (HRV drops after hard training) to show the whole pipeline works, and only then trust new findings. Option C swaps steps 2 and 3: confirming a known result must come before searching for new ones.
Source: Slides p. 57–68
Q4 · HRV and context
A user's morning HRV (rMSSD) is well below their normal range, and their resting heart rate is higher than usual. Based on the lecture's HRV4Training data, what is the best interpretation?
Show answer
Answer: B. HRV reacts to every stressor but points to none: hard training, sickness and heavy alcohol all lowered HRV and raised resting heart rate, so tags like 'sick' or 'alcohol' supply the reason. Option A is tempting, but sickness is only one possible cause; and after easy training HRV actually rose slightly (option D).
Source: Slides p. 17, 25, 65–70
Q5 · Limitations of typical lab studies
A lab study measured how HRV responds to training intensity in N = 10 male students. Which limitation does the lecture emphasize?
Show answer
Answer: A. Sport science studies often have N = 2–10, so results are valid only for the people measured, and extending them means running a new study (time, money). Option B is wrong: lab data are usually high quality; the problem is generalizability (whether results hold for other people), not quality.
Source: Slides p. 35–50
Q6 · Noisy data: estimating max heart rate
Without lab tests, an app estimates each user's maximum heart rate from the highest heart rates in their everyday workouts, so it can express training intensity as a percentage of that maximum. Which assumption does this need?
Show answer
Answer: C. The lecture's assumption is that there will be some hard sessions in the monitored period. Option A is not needed: glitches such as 234 or above 300 bpm can be removed by cleaning (e.g. the ±3 SD rule), but no cleaning can create a maximal effort that never happened ('but did they ever go hard?').
Source: Slides p. 83–91
Q7 · Missing data and selection bias
To simplify the analysis, you remove all users with missing workouts from a large app data set. What is the main risk the lecture points out?
Show answer
Answer: B. Users with gaps are not a random subset: keeping only consistent loggers keeps the more disciplined, probably fitter users, so results describe them rather than typical users (selection bias). Option A is wrong because a bigger sample shrinks random error, not a systematic difference in who is included; the lecture's advice is 'no universal answer, think critically'.
Source: Slides p. 91–93
Q8 · Reference data
Which question is NOT on the lecture's list of things to ask about reference data (the outcomes a model learns from) before collecting data in an app?
Show answer
Answer: D. The lecture's list asks what the outcomes are, whether we can track them, whether we ask too much of the user and whether collecting them is ethical. More sensor data cannot replace a missing outcome: 'you can have unlimited data and still have no use for it' (the MyFitnessPal story).
Source: Slides p. 97–98, 118–120, 134
Q9 · Running performance example
In the lecture's running study (about 2,100 runners), groups of features were added step by step to a regression model that predicts each runner's best 10 km time. Which statement matches the results?
Show answer
Answer: A. R² rose from 0.33 (resting physiology) via 0.71 (volume and speed) and 0.76 (intensity distribution) to 0.87 with previous performance: 'the best predictor of performance is performance'. Option C confuses R² (the share of differences between runners that the model explains) with accuracy; the target is a time, so this is regression. Exam note: the slide says 'N = 2100, RMSE = 2 minutes (4%)'; the paper reports 2,113 runners and an RMSE of 2.6 minutes, with 4% being the mean percentage error.
Source: Slides p. 101–113
Q10 · Pandemic example
During the first COVID-19 lockdown (March–May 2020), most people in the lecturer's poll (67.6%) expected resting heart rate to rise. What did data from about 5,500 app users show?
Show answer
Answer: C. In 2020 resting heart rate fell about 1.4 bpm below its January level, against about 0.4 bpm in 2019; less travel and more sleep are plausible explanations, not proven causes. Option D gets the sleep direction wrong: sleep time increased.
Source: Slides p. 124–130
Q11 · Why user-generated data matters
HRV4Training (managing stressors), Bloomlife (start of labour) and Oura (spotting infections) were the lecture's examples. What do these uses have in common?
Show answer
Answer: D. None of these uses was the original goal of the product; user-generated data made them possible thanks to contextual data recorded over time and reference points such as user-reported events. Option C contradicts the lecture's slogan 'it's not just the hardware anymore': the value lies in what the data can reveal.
Source: Slides p. 16–26
Q12 · Limitations of app-user data 2 pointscatch-up notes
A running app has morning heart data from 30,000 users. A researcher wants to use these data to give training advice to all athletes. State and motivate briefly:
1) one limitation that concerns who the users are;
2) one limitation that concerns the data themselves.
Show model answer
1) Selection bias: app users choose to use the app (health-interested, mostly men, e.g. 1,891 of 2,113 runners in the running study), so results may not hold for all athletes; a bigger N does not fix this. Also acceptable: the results are group averages that may not describe an individual athlete. 2) Data quality: consumer sensors are noisy, wrong readings (artifacts) are not flagged by the device, and workouts or outcomes are often missing, so conclusions rest on uncertain data.
How the points are earned
- 1 pt Names a limitation about who the users are (self-selected / not representative / group average) and explains it
- 1 pt Names a limitation about the data (noisy, artifacts, missing data, no reference outcome) and explains it
Source: Slides p. 81, 92, 129–130
How did your answer compare?
Q13 · API 2 pointscatch-up notes
A sleep app wants to know each user's training without asking users to type in their workouts. State and motivate briefly:
1) What is an API, and how could the app use one here?
2) Name one problem with user data that an API does not solve.
Show model answer
1) An API (application programming interface) is a defined connection through which one program automatically asks another program for data, without a person typing anything. The app could fetch each user's workouts (date, distance, time, heart rate) from Strava, TrainingPeaks or Garmin. 2) An API only moves data that already exist: it does not make sensor readings accurate, fill in missing workouts, make the users representative or check that an outcome is correct.
How the points are earned
- 1 pt Defines an API as a connection that lets one program request data from another automatically
- 0.5 pt Gives a concrete use: fetching workouts from another app (Strava, Garmin, TrainingPeaks)
- 0.5 pt Names one problem an API does not solve (accuracy, missing data, representativeness, outcome validity)
Source: Slides p. 26, 81, 92, 98, 101
How did your answer compare?
Q14 · Opportunities and limits of user-generated data 4 points
An app compared its users' morning resting heart rate during the first COVID-19 lockdown (spring 2020) with the same months in 2019. Most people in a poll expected resting heart rate to rise. State and motivate briefly:
1) Interpret the result below.
2) Why could these user data show this when a typical lab study could not?
3) Give two limitations of the conclusion.
App data: about 5,500 users, 3 months each, about 500,000 measurements
Change in resting heart rate from the January level, April–May:
2019: about −0.4 bpm
2020: about −1.4 bpm
Same months in 2020: travel strongly down, sleep time upShow model answer
1) Resting heart rate fell during the lockdown, clearly more than in the same months of 2019, the opposite of what most people expected. Less travel and more sleep are plausible explanations, not proven causes. 2) The app was already measuring thousands of people every day in real life, including 2019 for comparison, so an unplanned event could be studied at scale; a lab study would have to be planned in advance, has a small N, and you cannot schedule a pandemic. 3) The users are a self-selected group (people interested in tracking their health), so the result may not hold for the general population. The numbers are group averages: individuals may have changed in the opposite direction. (Also acceptable: noisy or missing measurements.)
How the points are earned
- 1 pt Resting heart rate fell (more than in 2019), opposite to expectation; travel/sleep are possible explanations, not proof
- 1 pt User data: already collected daily, at large scale, in real life, with 2019 for comparison; a lab study cannot be planned for an unforeseen event
- 1 pt Limitation 1: self-selected users, not the general population
- 1 pt Limitation 2: group averages hide individuals (or noisy/missing data)
Source: Slides p. 14, 54, 122–130
How did your answer compare?
Q15 · Sickness: detection vs predictionfrom the recording
A wearable company says its ring 'predicts' infections, because users' resting heart rate rose before they reported their first symptoms. How did the guest lecturer judge such a claim?
Show answer
Answer: C. He said that if the data show it, 'your body is already sick… you are detecting at best'; and the symptom day is only an uncertain reference point, since infection may have happened days earlier. Option A confuses 'earlier than a late reference point' with forecasting an illness before it starts.
Source: Recording 62:26 (also 24:46); Slides p. 69, 116–117
Q16 · False positives at scalefrom the recording
An app's sickness alert has a false-positive rate of 5% (5 in every 100 healthy users get a wrong alarm). Why is this a bigger problem for an app with 1,000,000 healthy users than for a study with 20 healthy participants?
Show answer
Answer: A. The false-positive rate is a property of the method; only the count grows with the number of users: 0.05 × 20 = 1 versus 0.05 × 1,000,000 = 50,000. Option B is tempting, but adding users does not change the rate itself.
Source: Recording 72:51
Q17 · Lifestyle stressors vs trainingfrom the recording
In the HRV4Training data, how did the morning changes in HRV and resting heart rate after sickness or heavy alcohol intake compare with those after hard training, and what did the lecturer conclude?
Show answer
Answer: D. HRV fell about 10–12% after sickness or alcohol versus about 3% after hard training, and resting heart rate rose about 6% versus under 1%; he called the lifestyle effects 'two or three times' as large. Options A–C contradict these plotted sizes and directions.
Source: Recording 22:43; Slides p. 66, 69–70
Q18 · Max heart rate from free-living datafrom the recording
You estimate users' maximum heart rate from months of everyday workouts. According to the guest lecturer, how can the data reveal users who probably never trained hard, and when is this a serious problem?
Show answer
Answer: B. A narrow spread of session heart rates flags people who always train at one intensity, so their maximum cannot be found; how much it matters depends on how common it is (3 of 7 vs 3 of a million). Option A describes a sensor glitch (artifact), not a missing hard effort; option C fails because age-based maximum heart rate cannot be trusted, given the large differences between people.
Source: Recording 48:03, 50:08; Slides p. 87–90
Mixed — across lectures
9 questions
Q1 · Scaling + validation for KNN
You build a KNN (k-nearest neighbours) classifier that predicts injury (yes/no) from weekly running distance (0–120 km) and sleep duration (5–10 h). Which workflow is most appropriate?
Show answer
Answer: B. KNN relies on distances, so unscaled, the km feature would dominate; and performance must be judged on data the model has not seen (a test set, or k-fold cross-validation). Option C lets the test set influence both the scaling and the choice of k, so the reported accuracy is too optimistic. Exam note: the lecture's KNN code scales the test set with its own statistics (scale(test), normalize(test)); strictly, the test set should be transformed with the numbers learned from the training set.
Source: L7 Slides p. 4, 8–12; L6 Slides p. 18–20
Q2 · Imbalanced outcome from wearable data
An app predicts whether a user will tag a day as 'sick'; sickness is tagged on about 3% of days. A model that always predicts 'not sick' reaches 97% accuracy. What should you do?
Show answer
Answer: C. Always predicting 'not sick' catches no sick day (recall 0, F1 0, balanced accuracy 50%), so accuracy misleads when one class is rare; and user tags are reference data whose start day is uncertain. Option A is the trap: 97% only reflects that sick days are rare.
Source: L6 Slides p. 27, 33; L9 Slides p. 97–98, 115–117
Q3 · Choosing the association measure
Match each design to the most suitable measure of association. (1) 100 injured and 100 uninjured runners are asked which shoe type they use. (2) 5,000 adults are followed for up to 10 years, with different follow-up times, and dates of death are recorded.
Show answer
Answer: A. In (1) the runners were chosen by outcome (a case–control design, like the lecture's 45 vs 45 heart-disease example), so the share injured is set by the researchers and risks cannot be estimated; only the odds ratio works. In (2) people are followed until death for different lengths of time, which the Cox model's hazard ratio handles. Option B is tempting, but a relative risk needs real risks, which a case–control design cannot give.
Source: L7 Slides p. 21–24; L8 Slides p. 58–60, 63
Q4 · Preprocessing free-living data
In a free-living running data set, 25% of users have missing resting-heart-rate values, and the heart-rate data contain impossible readings above 300 bpm. Which preprocessing plan is most defensible?
Show answer
Answer: D. A wrong value should be corrected (here: set to missing), and imputation fills gaps when deleting would lose too much; deleting users with gaps can cause selection bias. Option C is wrong: z-scoring only changes the units, so an impossible reading simply gets a large z-score.
Source: L7 Slides p. 5, 13, 15–16; L9 Slides p. 87, 92–93
Q5 · Model choice and validation 4 points
A running platform wants to predict each user's 10 km race time (in minutes) from training data imported from other apps: weekly distance, average speed, heart rate during training and previous best time. Data are available for 2,000 users. State and motivate briefly:
1) Which type of model would you choose and why?
2) How would you split the data and test the model, and why?
Show model answer
1) Regression, because the target (10 km time) is a number. For example multiple linear regression, which is easy to interpret, or LASSO/ridge regression, whose penalty protects against overfitting (LASSO can set unhelpful predictors to zero). 2) Hold out a test set (e.g. 80/20) that is used only once at the end, so the model is judged on users it has never seen. Inside the training set, use k-fold cross-validation to compare models or choose λ. Keep each user entirely on one side of the split, and report the error in minutes (e.g. MAE).
How the points are earned
- 1 pt Regression, because the target is a number
- 1 pt Names a suitable model with a reason (linear regression: interpretable; LASSO/ridge: against overfitting; random forest)
- 1 pt Holdout test set (e.g. 80/20 or 70/30), used once, to test on unseen users
- 1 pt k-fold cross-validation within the training data to tune/compare models
Source: L5 Slides p. 41–44, 57–60; L6 Slides p. 18–20; L9 Slides p. 101–113
How did your answer compare?
Q6 · Preparing free-living data 3 points
In the same running data set, some resting-heart-rate values are missing and some heart-rate readings are above 300 bpm. State and motivate briefly:
1) Give two preprocessing steps you would take.
2) Name one limitation of these data that no preprocessing can fix.
Show model answer
1) Treat impossible readings (above 300 bpm) as artifacts: set them to missing or remove them, e.g. with a ±3 SD rule. Handle the missing resting heart rate: impute it (e.g. mice with pmm) or delete those users only after checking that they do not differ from the rest (selection bias). (Also acceptable: scale the predictors with numbers from the training set.) 2) The users are self-selected app users (mostly men, health-interested), so results may not hold for all runners; or: a runner who never ran hard has no true best time in the data; or: the best workout is not a standardized race.
How the points are earned
- 1 pt Step 1: impossible values treated as artifacts (set to missing / removed, e.g. ±3 SD)
- 1 pt Step 2: missing values imputed (e.g. mice) or deleted with a check for selection bias
- 1 pt One limitation preprocessing cannot fix (self-selected users, never ran hard, non-standard outcome)
Source: L7 Slides p. 4–5, 15–16, 18; L9 Slides p. 87, 92, 101–113
How did your answer compare?
Q7 · Relative vs absolute risk
In app data from 4,000 runners, 6% of months in which morning HRV was more than 10% below the runner's normal range ended in an injury, against 2% of other months. What are the relative risk (low-HRV months compared with other months) and the absolute difference in risk?
Show answer
Answer: C. Relative risk = 6% ÷ 2% = 3, and the absolute difference = 6% − 2% = 4 percentage points; the lecture's rule is to always report both, because 'three times the risk' alone can sound bigger than it is. Option A swaps the two calculations (it divides where it should subtract and vice versa).
Source: L7 Slides p. 22–23
Q8 · Communicating a wearable-based risk finding 3 points
A coach sees that injuries were three times as common in months with low morning HRV (6% vs 2%) in data from 4,000 app users, and asks whether athletes should stop training whenever HRV drops. State and motivate briefly:
1) Give two reasons why this result does not prove that low HRV causes injuries.
2) Give one reason why the result may not apply to the coach's own athletes.
Show model answer
1) Confounding: hard training, sickness or alcohol lower HRV and may also raise injury risk, so a third factor can explain the link (HRV is sensitive but not specific). The data are observational: no experiment assigned low HRV, and an injury or illness that is already starting may itself lower HRV. (Also acceptable: noisy or missing consumer data.) 2) The app users are self-selected (health-interested, mostly men), and a group-level association need not hold for an individual athlete. So the finding supports looking at context, not an automatic 'stop training' rule.
How the points are earned
- 1 pt Reason 1: confounding by a third factor (hard training, sickness, alcohol)
- 1 pt Reason 2: observational association / reverse timing / noisy data
- 1 pt Generalizability: self-selected users or group average vs individual
Source: L8 Slides p. 59–61; L9 Slides p. 65–70, 81, 92, 129–130
How did your answer compare?
Q9 · Pooled vs subgroup results 4 pointscatch-up notes
A coach pooled data from sprinters and endurance athletes and found a strong positive correlation between weekly gym hours and jump height. The output below shows the results per group. State and motivate briefly:
1) Interpret the results.
2) What should the analysis have done instead?
3) Would you advise all athletes to do more gym hours?
Group n mean gym h/week mean jump (cm) r (gym vs jump)
All athletes 60 4.5 45 0.70
Sprinters 30 7.0 52 0.05
Endurance 30 2.0 38 0.45Show model answer
1) The strong pooled correlation comes mainly from the difference between the groups: sprinters do more gym hours and also jump higher. Within the groups the picture differs: almost no association in sprinters (r = 0.05), a moderate positive one in endurance athletes (r = 0.45). 2) Analyse each group separately (stratify), or include group in the model, because group (type of athlete) acts as a confounder of the pooled association. 3) No. A correlation does not show that more gym hours cause higher jumps, and for sprinters there is no association at all; at most, more gym work might be tested for endurance athletes. Group results also need not hold for each individual.
How the points are earned
- 1 pt Pooled correlation is driven by the group difference (sprinters: more gym hours and higher jumps)
- 1 pt Within groups: no association in sprinters, moderate positive in endurance athletes
- 1 pt Analyse groups separately (stratify) or include group in the model
- 1 pt No general advice: correlation is not causation, and no association in sprinters
Source: 2024 practice exam Q35; L8 Slides p. 59–61; L9 Slides p. 130
How did your answer compare?