How accurate is Apple Watch sleep tracking?
In short, Apple Watch can be a useful reference for reviewing whether you were asleep or awake and the broad trend in total sleep time. It is not, however, a device that establishes “exactly how many minutes of deep sleep” you had with the same accuracy as polysomnography (PSG).
The most important point is not to mix the following three kinds of numbers.
- Results from Apple’s validation of its own Sleep Stages feature
- Results from independent researchers comparing a commercially available Apple Watch with PSG
- Results from Feelmo Paper 17 evaluating a research model that uses heart-rate-derived information and temporal context (the detailed feature design, preprocessing, and model architecture are unpublished)
Apple’s official development dataset contains 1,171 nights. By contrast, the public dataset in Paper 17 with paired Apple Watch heart rate and PSG contains 31 people. Even if the numbers look similar at a glance, the manufacturer’s development data and Feelmo’s internal analysis are different studies.
What is PSG?
PSG records several signals at the same time, including brain activity, eye movements, muscle activity, ECG, and breathing. A trained specialist assigns a sleep stage every 30 seconds. It is the reference standard for sleep-stage comparisons.
Apple Watch estimates sleep from small sensors on the wrist. That makes nightly measurement convenient, but it does not observe the same signals as PSG.
What did Apple’s official validation show?
According to Apple’s 2025 technical document, the native Sleep Stages classifier uses accelerometer data as its input. It does not classify sleep from HRV alone. It classifies Wake, REM, Core, and Deep every 30 seconds from acceleration signals that include subtle wrist movements related to breathing.
| Population in Apple’s document | People | Records | Main result |
|---|---|---|---|
| Development data | 858 | 1,171 nights | Used to develop the model; not an independent evaluation |
| Validation data excluded from development | 166 | 299 nights | Mean four-stage κ 0.63 (SD 0.13) for the earlier version |
| Updated watchOS 26 version on the same validation data | 166 | 299 nights | Median four-stage κ 0.68 (SD 0.11) |
| Separate clinical-PSG cohort, updated version | 236 | 390 watch sessions | Median four-stage κ 0.66 (SD 0.16) |
Kappa (κ) measures agreement after accounting for agreement expected by chance. A value closer to 1 indicates greater agreement, while 0 is around chance. Results reported as means and medians, from different populations and operating-system versions, cannot simply be placed side by side as if they were directly comparable.
This is a large and important report, but it remains a manufacturer validation report. The raw training data and model are not public, and Apple itself does not present Sleep Stages for clinical use.
What did research independent of Apple find?
In a 2024 study, 35 adults without sleep disorders completed one night of PSG, and analysable Apple Watch Series 8 data were available for 29 people. Agreement between Apple Watch and PSG was 93% with κ=0.60 for sleep versus wake, and 75.0% with κ=0.60 across four stages.
However, compared with PSG, Apple Watch estimated an average of 43 fewer minutes of deep sleep and 45 more minutes of light sleep. The sample was small and included only one night in healthy adults, so the findings cannot be applied unchanged to people of different ages or with sleep disorders. The study was independent of Apple, but it received funding from competitor Oura Ring Inc., and the first author was a member of Oura’s Medical Advisory Board. Both the results and the funding and conflicts of interest need to be considered.
A systematic review published in 2026 likewise concluded that Apple Watch distinguishes sleep from wake relatively well, while discrimination between similar sleep stages is moderate to low. In other words, total sleep time and the minutes assigned to each stage do not have the same degree of certainty.
What did Paper 17 evaluate?
Paper 17 did not reproduce Apple’s native Sleep Stages feature. It was an unreviewed internal feasibility study using a research model with heart-rate-derived information and temporal context from public cohorts to estimate five PSG stages. The specific feature design, temporal reference range, preprocessing, and model architecture are unpublished.
| Feelmo internal model evaluation | Sample | Macro AUROC | κ | Result |
|---|---|---|---|---|
| Public Walch Apple Watch / PSG data | 31 people | 0.708 | 0.221 | Did not meet the internal target |
| HMC clinical PSG data (participant-level five-fold evaluation) | 35 people | 0.711 | 0.244 | Did not meet the internal target even within the same cohort |
| HMC + CAP + SLPDB clinical PSG data (LODO) | 67 records total | Mean 0.639 (SD 0.013) | Mean 0.132 (SD 0.028) | Performance fell further on an unseen cohort |
AUROC describes how well each stage is ranked apart from the other stages; κ describes final stage agreement. Within HMC, AUROC was 0.711 while κ for five-stage agreement was only 0.244. In leave-one-dataset-out (LODO) evaluation, which tested a whole cohort excluded from training, mean AUROC fell to 0.639 and mean κ to 0.132. Numbers from different evaluation designs should not be treated as the same performance measure.
For five sleep summaries calculated from predicted stages across 67 records in three cohorts, Pearson correlations with PSG ranged from −0.007 to 0.417, and none met the prespecified internal target. A separate exploratory analysis reached a maximum correlation of 0.545 for the proportion of deep sleep, but still failed to meet the revised internal target across all five summaries. In another analysis, increasing heterogeneous data from 53 to 101 records left κ essentially unchanged, from 0.336 to 0.330. Simply adding more mixed data did not solve the problem. The model’s feature design, temporal reference range, preprocessing, and architecture remain proprietary and unpublished.
What Paper 17 supports
With these public data and this research model, five-stage estimates based on heart-rate-derived information and temporal context did not replace PSG. This is neither an evaluation of Apple’s native feature nor proof that every future model must fail.
How should you read the numbers each morning?
- Look at trends over several weeks using the same device and settings, rather than drawing a conclusion from one night
- Do not assume total sleep time and the minutes labelled REM, Core, or Deep are equally certain
- Do not conclude that one night with “little deep sleep” means illness or inadequate recovery
- An operating-system update can change the algorithm, so watch for a step change before and after an update
- If severe sleepiness, snoring, breathing pauses, or insomnia persist, do not rely on watch numbers alone; consult a healthcare professional
References
- Apple. Estimating Sleep Stages from Apple Watch. October 2025. Official technical document
- Robbins R, et al. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors. 2024;24:6532. doi:10.3390/s24206532
- Lambe R, et al. The accuracy of Apple Watch measurements: a living systematic review and meta-analysis. npj Digital Medicine. 2026. doi:10.1038/s41746-025-02238-1
- Walch O, et al. Sleep stage prediction with raw acceleration and photoplethysmography heart rate data derived from a consumer wearable device. Sleep. 2019;42(12):zsz180. doi:10.1093/sleep/zsz180
- Walch O. Motion and heart rate from a wrist-worn wearable and labeled sleep from polysomnography, version 1.0.0. PhysioNet. doi:10.13026/hmhs-py35
Medical caution
Apple Watch, Feelmo, and the Paper 17 model do not replace diagnosis by PSG. This page explains research findings; it is not diagnostic or treatment advice.
※ This research explanation does not establish that the Feelmo app is effective.