Skip to content

Instantly share code, notes, and snippets.

Show Gist options
  • Select an option

  • Save patricksavalle/5b55a469d44802b269a9a1acdd643c57 to your computer and use it in GitHub Desktop.

Select an option

Save patricksavalle/5b55a469d44802b269a9a1acdd643c57 to your computer and use it in GitHub Desktop.
Apple Watch VO2max accuracy vs lab CPET, Cooper test, 2000m row, and FTP — a source-traced investigation, ranked (2026-05-29)

How accurate is the Apple Watch VO₂max — versus a lab test, the Cooper test, a 2000m row, and FTP?

A source-traced investigation. Access date for all fetched sources: 2026-05-29.

Every load-bearing factual claim below carries a warrant label: (traced) = primary fetched and read this session (URL in source table); (traced — search summary) = obtained via a web-search summary that quotes the primary, primary not independently opened this session; (deferred to consensus) = relying on a named body of knowledge; (memory — unverified) = recalled, not verified. The point is to make verification — and its absence — visible.

Pre-registered hypotheses

Registered before searching, to keep the search from selecting the question:

  • H1: Laboratory CPET (breath-by-breath gas analysis) is the gold standard; the maximal field proxies (Cooper run, 2000m row, FTP) are more accurate than the Apple Watch because they rest on actual maximal performance, while the Watch infers from a submaximal heart-rate response. Provenance: analyst framing.
  • H2: The Apple Watch is competitive with the field proxies for tracking/ranking purposes even if biased in absolute terms. Provenance: steelman of the device.
  • H3: The ranking is sport- and individual-dependent — each proxy is accurate only within its own modality (and only for someone proficient in it), and the Watch is systematically biased for non-runners. Provenance: analyst framing.

Bottom line

VO₂max — the maximum rate at which the body can take up and use oxygen — is the reference measure of aerobic fitness. It can only be measured in a lab, via gas analysis during a maximal exertion test (CPET). Every other number, including the one on your wrist, is an estimate produced by a population-calibrated formula. (deferred to consensus)

Ranked by agreement with lab CPET:

  1. Lab CPET (gas analysis) — direct measurement; the reference by definition.
  2. Cooper test (12-minute run) — the best estimate, if you are a trained runner (r ≈ 0.90–0.93 with measured VO₂max). (traced)
  3. 2000m row (Concept2/Hagerman) — strong for trained rowers, but error climbs sharply once technique and pacing fall short. (traced — search summary)
  4. FTP (cycling) — an excellent performance metric but a mediocre VO₂max estimator (correlation ranges r ≈ 0.45 to 0.80). (traced)
  5. Apple Watch — the least accurate: a systematic underestimation of ≈6 mL/kg/min and a typical error of 13–16%, but the only option that works passively, without maximal effort. (traced)

The single most important correction to popular framing: the binding limitation is not that the Watch's sensor is bad, but that (a) it estimates from a submaximal signal during outdoor walking/running only, (b) it underestimates structurally — even in fit users — by about 6 points, and (c) every field proxy that beats it is itself a modality-locked estimate, not a measurement. The Watch trades accuracy for convenience; treat it as a trend line, not an absolute number.


Measured vs. estimated — why the methods differ

A lab measures VO₂max directly: you breathe through a mask to exhaustion while a metabolic cart measures oxygen consumption litre by litre. (deferred to consensus — exercise physiology) Every other method derives VO₂max from a performance or a heart-rate response, via an equation fitted to a sample. The methods therefore differ on three axes: whether maximal effort is required, whether the modality matches the person, and how wide the calibrating equation's error band is.

The Apple Watch is unique in the list: it requires no maximal effort. It estimates VO₂max from the heart-rate response during outdoor walking, running, or hiking (terrain under ~5% grade), requiring roughly a 30% increase in heart rate above resting. No gas analysis, no maximal test — and, critically, it does not work from cycling, rowing, indoor exercise, or strength work. (traced)


The ranking: accuracy relative to lab CPET

Rank Method What it is What the research says about deviation vs. lab CPET Verdict
1 Lab CPET (gas analysis) Direct measurement of O₂ uptake during a maximal test The criterion itself; residual variation is mainly biological day-to-day variation Gold standard
2 Cooper test (12-min run) Maximal field test; distance → equation r ≈ 0.90–0.93 with measured VO₂max; tends to overestimate (~3 mL/kg/min in one validation); unreliable for non-runners Best proxy (within running)
3 2000m row (Concept2/Hagerman) Maximal field test; time + sex + weight → equation Strong association in trained rowers; "accuracy decreases substantially" for non-rowers (technique, pacing) Good within its own sport
4 FTP (functional threshold power) 20-min threshold power on a bike Moderate-to-fair link to VO₂max (r ≈ 0.45 in trained cyclists; up to r ≈ 0.80 with age-adjusted equations). Better as a performance than a VO₂max metric Strong performance metric, weak VO₂max estimator
5 Apple Watch Passive estimate from submaximal HR during walking/running Systematic underestimation 6.07 mL/kg/min; MAE 6.92; MAPE 13.3% (Series 9/Ultra 2) to 15.8% (Series 7); limits of agreement −6.11 to +18.26 mL/kg/min Trend only

Caveat on the rank order. Places 2–4 are modality-dependent. For a trained rower the 2000m test is more accurate than the Cooper test; for a cyclist, FTP is the most relevant number (though still a weak VO₂max estimator). The ranking assumes a person is tested in the sport they are proficient in. And every proxy (2–5) converts a performance or heart rate into VO₂max via a population equation — none is lab-grade.


Method-by-method

Lab CPET — the reference

Direct indirect-calorimetry measurement during a maximal protocol (e.g. modified Åstrand on a treadmill with a COSMED metabolic cart in the validation below). (traced) It defines accuracy here; its own limitation is cost, access, and the fact that it captures one day's physiology. (deferred to consensus)

Cooper 12-minute run — best proxy, modality-locked

A validation in 88 male university students found r = 0.93 between Cooper distance and measured VO₂max, with the original equation overestimating by ~3 mL/kg/min (predicted 42.8 vs measured 39.8), prompting a population-specific re-fit. (traced) It is maximal and matched to running — but it measures running ability and pacing as much as aerobic ceiling for anyone who isn't a runner. (traced — search summary)

2000m row — accurate as your technique

Concept2's estimate uses Fritz Hagerman's formulas (calibrated on thousands of rowers via gas analysis), selecting different equations for trained vs untrained rowers and using 2k time, sex and weight. (traced — search summary) The estimate stands or falls on rowing economy: an equally fit non-rower posts a slower 2k through inefficient technique and is handed a falsely low VO₂max. The literature's general rule: run an estimated-VO₂max test in the modality you know best. (traced — search summary)

FTP — a performance number masquerading as fitness

In moderately trained cyclists, relative FTP (W/kg) predicted race performance (r = −0.74) far better than VO₂max did (r = −0.37) — but FTP's correlation with VO₂max in that same study was only r = 0.45, and the authors warn against treating FTP as a VO₂max surrogate. (traced) Other equations (age + FTP) reach r ≈ 0.80 / R² = 0.776, so the link is real but inconsistent and equation-dependent. (traced — search summary) FTP tells you how hard you can ride for an hour; it is a noisy proxy for maximal oxygen uptake.

Apple Watch — largest error, but zero-effort

The strongest independent validation (PLOS One 2025, Lambe et al., 28 analysed, vs a COSMED cart) found a mean underestimation of 6.07 mL/kg/min, MAE 6.92, MAPE 13.31%, with limits of agreement of −6.11 to +18.26 — i.e. an individual reading can be off by nearly 20 points. (traced) An Apple Watch Series 7 study reported MAPE 15.79% / RMSE 8.85, with better accuracy in fitter participants. (traced — search summary) Tellingly, 21 of 30 PLOS participants were "excellent/superior" fitness — so the error was measured in the group where the Watch should do best. (traced)


Three nuances the ranking can't hold

  1. The Watch underestimates structurally, and in the fit. The −6 mL/kg/min bias is not a tail effect; it is the mean, measured in a fit cohort. (traced)
  2. FTP is a performance metric, not a fitness meter. Its popularity as a "fitness number" rests on convenience, not physiological equivalence to VO₂max. (traced)
  3. All field proxies are estimates of an estimate. Each converts a maximal performance into VO₂max via a sample-fitted equation; none measures gas exchange. The "better than the Watch" ranking is relative, not a claim of lab accuracy. (traced / traced — search summary)

Verdict on the hypotheses

  • H1 (lab gold standard; maximal proxies beat the Watch): Supported. The Watch carries the largest error (MAPE 13–16%; bias −6 mL/kg/min; LoA −6 to +18), while the Cooper test reaches r ≈ 0.93 and the row is strong in trained rowers. (traced)
  • H2 (Watch competitive for tracking/ranking): Partially supported. For intra-person trend tracking the Watch is usable; but its wide limits of agreement undermine cross-person ranking, and its absolute value is the worst in the list. The "competitive" claim holds only for the narrow use of watching one's own line move over time. (traced)
  • H3 (ranking is sport-/person-dependent): Supported. The proxy order reshuffles by modality, FTP is a weak VO₂max estimator even for cyclists, and every proxy's accuracy collapses for someone unskilled in that sport. (traced)

Synthesis: Lab CPET is the only path to a reliable absolute number. For an at-home estimate, run a maximal test in your sport and read the result as a band, not a point. The Apple Watch is valuable for answering "is my line trending up?" — not "how fit am I?" — and sits at the bottom of the accuracy ranking with the lab unambiguously at the top.


What would change this assessment

  • A direct head-to-head in which the same people undergo lab CPET, Cooper, 2000m row, FTP, and an Apple Watch estimate within a short window — not found in this search round.
  • Apple Watch validation in average/below-average fitness cohorts (current studies skew fit and young; n≈28).
  • A traceable, primary standard-error figure for the Cooper test in mL/kg/min (the one fetched paper reported an implausible 0.193, almost certainly a units artifact; the reliable takeaways are r ≈ 0.93 and ~3 mL/kg/min bias).

Bias self-audit (Rule 6)

Would the verdict be the same if the socially expected answer were reversed? The expected answer ("lab best, wearable worst") is confirmed, which carries confirmation-bias risk. Two counter-intuitive findings indicate the analysis tracked sources rather than the cliché: FTP — treated by many cyclists as a fitness number — is a weak VO₂max estimator, and even the "better" field tests are population estimates, not measurements. The numbers support the order independently of expectation (Watch MAPE 13–16% exceeds the Cooper spread). No source carried a manufacturer interest that colored the verdict; the Apple validations are independent (PLOS One, JMIR), not run by Apple — which, if anything, cuts against the expected pro-device tilt.


Source table (warrants)

# Source Used for Warrant
1 Lambe et al., Investigating the accuracy of Apple Watch VO₂ max measurements: A validation study, PLOS One 2025 — PMC12080799 Watch bias −6.07; MAE 6.92; MAPE 13.3%; LoA −6.1 to +18.3; method (submaximal, outdoor walk/run); fit-skewed sample (traced) — fetched 2026-05-29
2 Assessing the Accuracy of Smartwatch-Based VO₂max Estimation (Apple Watch Series 7), JMIR 2024 — PMC11325102 MAPE 15.79%; RMSE 8.85; better in fitter participants (traced — search summary)
3 Validity of Cooper's 12-minute run test for estimation of VO₂max, 2015 — PMC4314605 Cooper r = 0.93; ~3 mL/kg/min overestimation; population-specific equation (traced) — fetched 2026-05-29
4 The Validity of FTP and VO₂max for Cycling Performance in Moderately Trained CyclistsPMC6835290 FTP↔VO₂max r = 0.45; FTP predicts performance (r = −0.74) better than VO₂max (r = −0.37) (traced) — fetched 2026-05-29
5 Linear relationships VO₂max–Pmax–FTP (search result) Alternative equations: age + FTP → VO₂max r ≈ 0.80; R² = 0.776 (traced — search summary)
6 Estimating VO₂max with a 2000m Row (9to5strength, explaining Hagerman/Concept2) — link 2k equations (time+sex+weight, trained/untrained); accuracy drops sharply for non-rowers; technique/pacing dependence (traced — search summary)
7 Concept2 VO₂max calculator Hagerman methodology (gas analysis on thousands of rowers) (traced — search summary) — primary page returned 403

Search volume: 4 web searches + 6 documents fetched this session (4 primary studies fully read). The npj Digital Medicine 2025 meta-analysis fetched in a prior session corroborates the Watch figure (−6.07 mL/kg/min from this same PLOS study). Sufficient for a focused method comparison; the missing piece is a single-cohort head-to-head, flagged above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment