The short answer
In public Apple Watch data from 31 participants, a research model using heart-rate-derived information and temporal context achieved agreement of κ=0.221 with PSG across five stages. A separate clinical-ECG experiment in 35 participants achieved κ=0.244, but these were not the same people, so the two signals cannot be considered equivalent. The 67-record nightly-summary analysis used clinical ECG, not Apple Watch. These results cannot tell us how accurate Apple's native feature is, or whether REM percentages from wearable apps in general are accurate.
Key points
- This evaluates a research model using heart-rate-derived information and temporal context, not Apple's native Sleep Stages feature; its detailed design is unpublished
- Agreement was κ=0.221 in 31 Apple Watch participants and κ=0.244 in a separate 35-participant clinical-ECG cohort; the separate cohorts do not establish equivalence
- Removing the surrounding time information reduced κ to 0.092 and 0.119, suggesting that temporal context helped this model
- The 67-record nightly-summary analysis used clinical ECG and did not test REM percentages from Apple Watch
- Held-out AUROC across three cohorts was 0.621–0.649, a range of 0.028; this does not prove that the model is free of overfitting
- The results are an unreviewed internal secondary analysis with no public external preregistration
First, separate this from Apple's native feature
Apple's native Sleep Stages feature uses Apple Watch accelerometer signals as its input and estimates four states: Awake, REM, Core and Deep. Apple has published a technical report on that feature, and an independent study has compared Apple Watch Series 8 with PSG.
Our internal analysis is different. We used heart-rate-derived information and surrounding time information from a public Apple Watch dataset in a research model that estimated five stages: Wake, N1, N2, N3 and REM. The specific feature design, preprocessing and model architecture are unpublished. Its input and stage definitions differ from Apple's native feature, so its results cannot be substituted for Apple's own validation results.
What was the reference, and which data were used?
We used sleep stages from polysomnography (PSG) as the reference. PSG records several signals, including brain activity, eye movements and muscle activity, and trained scorers assign a stage to each 30-second epoch.
The Apple Watch experiment used 31 participants and 25,813 epochs from the public dataset published by Walch and colleagues, which recorded Watch and PSG data during the same nights. The separate clinical-ECG experiment used 35 HMC participants. Both used participant-grouped five-fold evaluation, but they were separate experiments with different participants and signals.
Result 1: the point estimates were close in two small experiments
The two point estimates were close, but this was not a comparison in which the same participants wore an Apple Watch while clinical ECG was recorded. We did not estimate a confidence interval for the difference or run an equivalence or non-inferiority test. We therefore cannot conclude that Apple Watch heart rate is equivalent to clinical ECG.
Cohen's κ measures agreement after accounting for agreement expected by chance. We report the values directly rather than assigning a fixed label such as "good" or "adequate". Fitness for a particular use would also require uncertainty estimates and stage-by-stage performance.
- Clinical PSG-derived HRV: macro AUROC 0.711, agreement κ 0.244
- Public Apple Watch dataset: macro AUROC 0.708, agreement κ 0.221
Result 2: surrounding time information helped the model
We compared the research model with and without surrounding time information. Agreement was higher when that context was included; the exact context range and how it entered the model are unpublished.
- Clinical PSG-derived HRV, with context → without context: κ 0.244 → 0.119
- Public Apple Watch dataset, with temporal context → without temporal context: κ 0.221 → 0.092
Result 3: the 67 nightly summaries came from a separate clinical-ECG analysis
We next summarized predicted stages across a night in 67 clinical-ECG recordings: 35 from HMC, 15 from CAP and 17 from the MIT-BIH Polysomnographic Database. No Apple Watch recordings were included in this analysis. We compared Pearson correlations with summaries derived from PSG against a prespecified internal target; the target details are unpublished.
- Wake after sleep onset (WASO): r = 0.42
- Share of deep sleep (N3%): r = 0.32
- Sleep efficiency: r = 0.30
- Sleep latency: r = −0.007
- Share of REM sleep: r = 0.003
Result 4: report the range across three held-out cohorts
For the three clinical-ECG cohorts, we also trained on two cohorts and tested on the remaining one, an evaluation called leave-one-dataset-out, or LODO. Held-out AUROC was 0.621 for HMC, 0.647 for CAP and 0.649 for SLPDB. The corresponding κ values were 0.117, 0.172 and 0.107.
The difference of 0.028 is the range between the largest and smallest held-out AUROC. It is not a measured "performance drop" from an in-domain result to an external dataset. Similar point estimates across these three cohorts do not prove that the model is free of overfitting or will perform the same way in new users.
What can and cannot be concluded
Within these small public datasets, heart-rate-derived features plus nearby time information allowed our custom model to distinguish sleep stages to a limited degree. Agreement without EEG remained limited, and nightly summaries in the separate clinical-ECG analysis did not reach the internal target.
The study does not establish the accuracy of Apple's native feature, equivalence between Apple Watch and clinical ECG, the accuracy of REM percentages from wearables in general, or the ability to diagnose a sleep disorder. Each of those questions needs a study designed for that specific product, signal and population.
Transparency and limitations
- This is an unreviewed internal secondary analysis by Feelmo and has not yet been reproduced by independent researchers
- Evaluation criteria were specified in internal materials, but the detailed targets are unpublished and were not registered publicly with a third-party service before the results were examined
- The Apple Watch experiment had 31 participants, the clinical-ECG experiment had 35, and the nightly-summary analysis had 67 recordings; the analysis does not provide a full account of uncertainty with confidence intervals
- The Apple Watch and clinical-ECG results used different participants, signals and features, so they are not a direct comparison of the two modalities
- The LODO AUROC range of 0.028 does not prove an absence of overfitting
- All evaluations used public datasets and no Feelmo user data
- The analysis was not designed to diagnose sleep disorders, guide treatment, or evaluate Apple's native Sleep Stages feature
References
- Estimating Sleep Stages from Apple Watch (technical report on the native feature) — Apple
- Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults — Sensors
- Sleep stage prediction with raw acceleration and PPG heart rate data — Sleep
- Motion and heart rate from a wrist-worn wearable and labeled sleep from PSG — PhysioNet
- Haaglanden Medisch Centrum sleep staging database — PhysioNet
- CAP Sleep Database — PhysioNet
- MIT-BIH Polysomnographic Database — PhysioNet
