Skip to content

Instantly share code, notes, and snippets.

Show Gist options
  • Select an option

  • Save patricksavalle/854d21c7a353a3b9e852767db3507a4f to your computer and use it in GitHub Desktop.

Select an option

Save patricksavalle/854d21c7a353a3b9e852767db3507a4f to your computer and use it in GitHub Desktop.
Apple Watch: how reliable are its measurements and derived biomarkers? A source-traced investigation (2026-05-29)

How reliable are the Apple Watch's measurements and derived biomarkers?

A source-traced investigation. Access date for all fetched sources: 2026-05-29.

Every load-bearing factual claim below carries a warrant label: (traced) = primary fetched and read this session (URL in source table); (traced — search summary) = obtained via a web-search summary that quotes the primary, primary not independently opened this session; (deferred to consensus) = relying on a named body; (memory — unverified) = recalled, not verified. The point is to make verification — and its absence — visible.


Bottom line

The popular framing — "one good heart-rate monitor, one usable ECG, an acceptable activity and saturation meter, and everything else is a statistical model that isn't clinical-grade" — is broadly correct in its conclusion but imprecise in its premises, and it understates two things in opposite directions.

  1. The "accurate core" is more accurate on average but noisier per-reading than the framing implies. Heart-rate mean error is near zero (≈0.2–0.3 bpm), not "1–3 bpm" — but the limits of agreement are roughly ±7–8 bpm for any single reading, which the framing omits. (traced)
  2. The "derived models" are not uniformly soft. Some derived features (AFib screening, the new hypertension screen) have regulatory clearance and peer-reviewed validation, while others (sleep stages, stress/HRV, energy expenditure, VO₂max) have weak or essentially no independent validation. Lumping them together as one "statistical model" loses this gradient. (traced)

The most important correction is directional: the binding limitation of the Apple Watch is not raw sensor accuracy. It is that (a) every "biomarker" beyond heart rate, ECG rhythm, and SpO₂ is an inference whose error is rarely shown to the user, and (b) for the screening features, a positive result in a low-risk wearer is wrong most of the time (low positive predictive value), and a negative result does not rule the condition out (low sensitivity). The device is a wellness-trend instrument with a few genuine screening functions bolted on — not a diagnostic one.


What the device physically measures

Per the claim under examination and corroborated by the validation literature, the Apple Watch's sensors are: an optical heart sensor (PPG), an electrical ECG sensor, an SpO₂ sensor, skin-temperature sensors, accelerometer + gyroscope, GPS, and a microphone. (traced — search summary) Everything labelled a "biomarker" in the Health app is computed from these few raw channels. This premise is accurate and is the correct starting point: the question is not whether the signals are real, but whether the inferences drawn from them are reliable.


Tier 1 — Direct or near-direct measurements (best evidence)

The strongest single source is a living systematic review and meta-analysis in npj Digital Medicine (Lambe et al., University College Dublin / Insight Research Ireland; 82 studies, 430,052 participants, pooled mean age 41.3; funded by Science Foundation Ireland; authors declare no competing interests). (traced) Note its own stated limitations: 56% of included studies were rated high risk of bias, the evidence is biased toward physically active males, older adults and people with comorbidities are under-studied, and many tested now-discontinued models. (traced)

Metric Pooled result Reading
Heart rate (all conditions) mean bias −0.27 bpm (95% CI −0.72–0.17); LoA −7.19 to +6.64 Excellent on average; a single reading can be off by ~7 bpm either way. (traced)
Heart rate (resting) bias 0.21 bpm; LoA −8.14 to +8.56 Spread is wider at rest than during exercise. (traced)
Heart rate (exercise) bias −0.63 bpm; LoA −6.86 to +5.60 (traced)
SpO₂ bias −0.04%; LoA −4.01 to +3.94 Good in normoxia; agreement weaker in low-oxygen ranges. A 4-point band matters clinically near thresholds. (traced)
AFib detection sensitivity 0.79 (0.61–0.90), specificity 0.91 (0.81–0.96) Good rule-out of normal rhythm; misses ~1 in 5 AFib episodes. See caveats below. (traced)

Correction to the common framing. The claim that resting heart rate has "an average error of 1 to 3 bpm" is not what the best meta-analysis reports: the mean bias is essentially zero, but the limits of agreement (±8 bpm at rest) are the number a user should actually care about and are usually omitted. The framing is too pessimistic on the average and silent on the spread. (traced)


Tier 2 — Derived biomarkers, ranked by how well-validated they actually are

Step count, distance, energy expenditure, VO₂max — accuracy degrades sharply down this list

  • Step count: "moderate accuracy… moderate correlation and wide limits of agreement." (traced)
  • Energy expenditure (calories): error "often large and varied considerably." This is one of the least reliable common numbers on the watch. (traced)
  • VO₂max: a clinically significant mean difference of −6.07 mL/kg/min with wide limits of agreement — from a single 30-participant study. (traced)

Sleep — sleep/wake good, sleep stages unreliable

Two findings converge. The npj review: "good differentiation between sleep and wake states, but moderate-to-poor differentiation between physiologically similar sleep stages." (traced) A six-device head-to-head against polysomnography (the clinical gold standard) found the Apple Watch Series 8 was the best of the six tested, yet that still meant: sleep-detection sensitivity 96%, but wake-detection specificity only 52%, Cohen's κ = 0.53 (moderate), and per-stage epoch accuracy of ~83% (light), 69% (REM), 51% (deep), 52% (wake). The authors: useful for tracking "prolonged and significant changes in sleep architecture" but "do not serve as replacements for PSG in clinical diagnoses." (traced) Takeaway: trust "how long you slept" as a trend; do not trust a specific night's deep/REM breakdown.

Stress / HRV — valid only under narrow conditions

HRV from the Apple Watch agrees well with a chest-strap reference (>0.9) when manually triggered, at rest, and can detect HRV drops under mental stress. But: the RR-interval series has frequent gaps (~5 gaps per recording, ~6.5 s each) that distort frequency-domain HRV, and the automatically collected background data has not been validated. (traced — search summary) The "stress" and overnight-HRV trends shown passively are therefore the least anchored to a validated measurement.

Respiratory rate, wrist temperature, sedentary behaviour — unvalidated

The npj review states plainly these "remain unvalidated in included literature." Features built on wrist temperature (e.g. cycle/ovulation estimates) and respiratory rate are inferences without a published accuracy base in this corpus. (traced)


The skin-tone question — more nuanced than usually stated

A common claim is "PPG weakens with dark skin tones." A 2024 systematic review/meta-analysis in JMIR (note conflict: four authors are employed by or hold equity in Verily Life Sciences, a Google health subsidiary — a hostility flag) splits this: (traced)

  • SpO₂ (pulse oximetry): accuracy-RMS exceeded the FDA's ≤3% threshold in all pigmentation groups (3.96 / 4.71 / 4.15% for light / medium / dark); dark skin showed the largest overestimation bias (+1.27%). So oximetry is imperfect across the board, with a real dark-skin overestimation tendency. (traced)
  • Wearable pulse rate (heart rate): "no statistically or clinically significant bias based on skin pigmentation" in the mean — but the error spread for dark skin was far wider (LoA −33.7 to +32.5 bpm vs −16.0 to +13.5 for light). (traced)

Net: the skin-tone problem is real but lives mainly in SpO₂ accuracy and in per-reading noise, not in average heart-rate bias. The simplified "PPG fails on dark skin" claim is half right.


AFib: the gap between "screening sensitivity" and real-world value

Meta-analytic sensitivity/specificity (0.79 / 0.91) describes performance in studies enriched with AFib patients. In the general, low-prevalence population that actually wears the watch, the number that matters is the positive predictive value, and it collapses:

  • Apple Heart Study: only ~0.5% of participants ever got an irregular-pulse notification; among those who completed follow-up, only ~34% were confirmed with AFib. (traced — search summary)
  • Real-world clinical yield (Mayo Clinic, 264 patients evaluated after an Apple Watch alert): only 11.4% received any clinically actionable cardiovascular diagnosis; just 4.9% were diagnosed with AFib; "7 patients needed to be evaluated to establish 1 diagnosis," rising to 15 for asymptomatic alert-recipients. (Funded by FDA/NIH; authors disclose a Mayo–AliveCor partnership.) (traced)
  • A growing cardiology literature documents overdiagnosis harms: false alerts are associated with measurable declines in self-perceived physical wellbeing, and young low-risk users — the heaviest smartwatch users — bear a disproportionate false-positive burden. (traced — search summary)

So the AFib feature is genuinely good at confirming a normal rhythm and at catching paroxysmal AFib it happens to observe, but a notification in a healthy 30-year-old is more likely a false alarm than real AFib. This is a property of low disease prevalence, not a sensor defect — but it is the single most consequential reliability fact for the typical wearer.


Hypertension notifications (2025) — the newest, and most easily misunderstood, derived metric

This feature does not measure blood pressure. It analyses PPG/heart-rate signals over a 30-day window and infers signs of chronic hypertension, FDA-cleared as a screening tool only. (traced — search summary) Apple's own validation (a vendor source — hostility flag): sensitivity 41.2% (53% for stage II), specificity 92.3%. (traced — search summary) An independent JAMA analysis (Univ. of Utah, Feb 2026) found accuracy is strongly age-dependent and that, critically, absence of an alert still leaves ~34% probability of undiagnosed hypertension. (traced — search summary)

Reading: it misses roughly half of hypertension cases, so a non-alert is not reassurance; but because specificity is high, an alert is worth acting on (especially in older wearers). It is a net-useful nudge toward a cuff measurement — not a blood-pressure monitor.


Verdict on the three hypotheses

  • H1 (the framing: core sensors fine, rest is soft estimation): Largely supported, with two corrections. Correct that the impressive dashboard rests on a few raw signals and that most derived numbers are inferences. Imprecise on heart-rate error figures, and wrong to treat all derived metrics as equally soft — AFib and hypertension screening are validated and cleared; sleep-stages/stress/calories/VO₂max are not. (traced)
  • H2 (optimist: "derived ≠ unreliable; it's all validated"): Rejected as a general claim. Validation exists for specific features (sleep/wake, AFib screening, hypertension screening) but is weak or absent for sleep stages, stress/HRV background, energy expenditure, respiratory rate, and wrist temperature. (traced)
  • H3 (skeptic: even the core has understated failure modes): Supported. Per-reading limits of agreement (±7–8 bpm HR; ±4 pts SpO₂), SpO₂ dark-skin bias, and — above all — the low real-world PPV of AFib alerts are real, under-communicated failure modes. (traced)

Synthesis: treat the Apple Watch as (1) a reliable heart-rate trend monitor with noisy single readings, (2) a usable single-lead ECG and a decent rule-out for normal rhythm, (3) a reasonable activity/SpO₂ estimator, and (4) a set of screening prompts — AFib, hypertension — whose alerts mean "go get a real test," never "you have/don't have this." Everything passive and continuous (sleep stages, stress, HRV trends, calories, VO₂max) is wellness-grade trend data, not a biomarker.


What would change this assessment

  • Free-living, demographically representative validation (current evidence is biased toward active young males; 56% of studies are high risk of bias). (traced)
  • Outcome trials showing wearable AFib/hypertension screening reduces hard endpoints (stroke, CV death) enough to justify the false-positive burden — currently unresolved. (traced — search summary)
  • Per-reading uncertainty surfaced to users (Apple shows point estimates, not the ±8 bpm / ±4-pt bands the literature reports).

Bias self-audit (Rule 6)

Would the verdict be the same if the socially expected answer were reversed? The culturally expected answers point in two opposite directions — tech-enthusiast ("it's basically a medical device") and tech-skeptic ("it's a toy"). The evidence rejects both poles, which is some evidence the analysis tracked the sources rather than a prior. The clearest residual risk is deference to the strongest single source: the npj meta-analysis is well-conducted but itself flags 56% high-risk-of-bias inputs and an active-male skew, so even the "Tier 1" numbers are optimistic relative to a sedentary, older, or darker-skinned wearer. Two key sources carry vendor conflicts (Apple's hypertension paper; Verily authors on the skin-tone meta-analysis) and are labelled as such; the AFib real-world yield comes from a Mayo group with a disclosed competitor (AliveCor) partnership. None of these were treated as independent of their funders.


Source table (warrants)

# Source Used for Warrant
1 Lambe et al., The accuracy of Apple Watch measurements: a living systematic review and meta-analysis, npj Digital Medicine 2025 — PMC12823594 HR, SpO₂, AFib, sleep, calories, VO₂max, unvalidated metrics, risk-of-bias (traced) — fetched 2026-05-29
2 Impact of Skin Pigmentation on Pulse Oximetry & Wearable PR Accuracy, JMIR 2024 — jmir.org/2024/1/e62769 Skin-tone SpO₂ bias; no HR bias but wider dark-skin spread (traced) — fetched 2026-05-29; conflict: Verily/Google-affiliated authors
3 Six-device wrist-wearable sleep-staging vs PSG, SLEEP Advances 2025 — zpaf021 Sleep-stage κ, per-stage accuracy, "not a PSG replacement" (traced) — fetched 2026-05-29
4 Clinical evaluation following abnormal pulse on Apple Watch, Mayo Clinic — PMC7526465 Real-world post-alert yield (11.4% / 4.9% AFib / NND 7) (traced) — fetched 2026-05-29; discloses Mayo–AliveCor partnership
5 Empirical Health, Apple Watch hypertension alerts explainedlink HTN feature mechanism (PPG/30-day inference), 41.2%/92.3%, JAMA age-dependence (traced) — fetched 2026-05-29 (secondary; quotes Apple + JAMA Utah)
6 Apple, Hypertension Notifications Validation (Sept 2025) + FDA clearance coverage HTN sensitivity/specificity, screening-only framing (traced — search summary); vendor source — hostility flag
7 Apple Heart Study (NEJM 2019 / Stanford) AFib notification rate 0.5%; ~34% follow-up confirmation (traced — search summary) — primary not independently opened this session
8 HRV validation (Sensors 2018, PMC6111985) + serial-HRV validity studies HRV valid manual/at-rest; RR gaps; auto-data unvalidated (traced — search summary)
9 Overdiagnosis literature (Scand. J. Prim. Health Care 2026; JACC Advances 2025; PMC10358285 on false-alert wellbeing) AFib overdiagnosis harms; low-risk-user false-positive burden (traced — search summary)
10 Source gist under examination — gist 8833dded… The framing being tested (traced) — fetched 2026-05-29

Search volume: 7 web searches + 6 primary/secondary documents fetched and read this session. Sufficient for a moderate-to-complex consumer-health-tech audit; a dedicated deep-research pass would add free-living and demographic-subgroup primaries (currently the weakest part of the evidence base).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment