How accurate is Apple Watch sleep tracking?

In short, Apple Watch can be a useful reference for reviewing whether you were asleep or awake and the broad trend in total sleep time. It is not, however, a device that establishes “exactly how many minutes of deep sleep” you had with the same accuracy as polysomnography (PSG).

The most important point is not to mix the following three kinds of numbers.

  1. Results from Apple’s validation of its own Sleep Stages feature
  2. Results from independent researchers comparing a commercially available Apple Watch with PSG
  3. Results from Feelmo Paper 17 evaluating a research model that uses heart-rate-derived information and temporal context (the detailed feature design, preprocessing, and model architecture are unpublished)

Apple’s official development dataset contains 1,171 nights. By contrast, the public dataset in Paper 17 with paired Apple Watch heart rate and PSG contains 31 people. Even if the numbers look similar at a glance, the manufacturer’s development data and Feelmo’s internal analysis are different studies.

What is PSG?

PSG records several signals at the same time, including brain activity, eye movements, muscle activity, ECG, and breathing. A trained specialist assigns a sleep stage every 30 seconds. It is the reference standard for sleep-stage comparisons.

Apple Watch estimates sleep from small sensors on the wrist. That makes nightly measurement convenient, but it does not observe the same signals as PSG.

What did Apple’s official validation show?

According to Apple’s 2025 technical document, the native Sleep Stages classifier uses accelerometer data as its input. It does not classify sleep from HRV alone. It classifies Wake, REM, Core, and Deep every 30 seconds from acceleration signals that include subtle wrist movements related to breathing.

Population in Apple’s documentPeopleRecordsMain result
Development data8581,171 nightsUsed to develop the model; not an independent evaluation
Validation data excluded from development166299 nightsMean four-stage κ 0.63 (SD 0.13) for the earlier version
Updated watchOS 26 version on the same validation data166299 nightsMedian four-stage κ 0.68 (SD 0.11)
Separate clinical-PSG cohort, updated version236390 watch sessionsMedian four-stage κ 0.66 (SD 0.16)

Kappa (κ) measures agreement after accounting for agreement expected by chance. A value closer to 1 indicates greater agreement, while 0 is around chance. Results reported as means and medians, from different populations and operating-system versions, cannot simply be placed side by side as if they were directly comparable.

This is a large and important report, but it remains a manufacturer validation report. The raw training data and model are not public, and Apple itself does not present Sleep Stages for clinical use.

What did research independent of Apple find?

In a 2024 study, 35 adults without sleep disorders completed one night of PSG, and analysable Apple Watch Series 8 data were available for 29 people. Agreement between Apple Watch and PSG was 93% with κ=0.60 for sleep versus wake, and 75.0% with κ=0.60 across four stages.

However, compared with PSG, Apple Watch estimated an average of 43 fewer minutes of deep sleep and 45 more minutes of light sleep. The sample was small and included only one night in healthy adults, so the findings cannot be applied unchanged to people of different ages or with sleep disorders. The study was independent of Apple, but it received funding from competitor Oura Ring Inc., and the first author was a member of Oura’s Medical Advisory Board. Both the results and the funding and conflicts of interest need to be considered.

A systematic review published in 2026 likewise concluded that Apple Watch distinguishes sleep from wake relatively well, while discrimination between similar sleep stages is moderate to low. In other words, total sleep time and the minutes assigned to each stage do not have the same degree of certainty.

What did Paper 17 evaluate?

Paper 17 did not reproduce Apple’s native Sleep Stages feature. It was an unreviewed internal feasibility study using a research model with heart-rate-derived information and temporal context from public cohorts to estimate five PSG stages. The specific feature design, temporal reference range, preprocessing, and model architecture are unpublished.

Feelmo internal model evaluationSampleMacro AUROCκResult
Public Walch Apple Watch / PSG data31 people0.7080.221Did not meet the internal target
HMC clinical PSG data (participant-level five-fold evaluation)35 people0.7110.244Did not meet the internal target even within the same cohort
HMC + CAP + SLPDB clinical PSG data (LODO)67 records totalMean 0.639 (SD 0.013)Mean 0.132 (SD 0.028)Performance fell further on an unseen cohort

AUROC describes how well each stage is ranked apart from the other stages; κ describes final stage agreement. Within HMC, AUROC was 0.711 while κ for five-stage agreement was only 0.244. In leave-one-dataset-out (LODO) evaluation, which tested a whole cohort excluded from training, mean AUROC fell to 0.639 and mean κ to 0.132. Numbers from different evaluation designs should not be treated as the same performance measure.

For five sleep summaries calculated from predicted stages across 67 records in three cohorts, Pearson correlations with PSG ranged from −0.007 to 0.417, and none met the prespecified internal target. A separate exploratory analysis reached a maximum correlation of 0.545 for the proportion of deep sleep, but still failed to meet the revised internal target across all five summaries. In another analysis, increasing heterogeneous data from 53 to 101 records left κ essentially unchanged, from 0.336 to 0.330. Simply adding more mixed data did not solve the problem. The model’s feature design, temporal reference range, preprocessing, and architecture remain proprietary and unpublished.

What Paper 17 supports

With these public data and this research model, five-stage estimates based on heart-rate-derived information and temporal context did not replace PSG. This is neither an evaluation of Apple’s native feature nor proof that every future model must fail.

How should you read the numbers each morning?

  • Look at trends over several weeks using the same device and settings, rather than drawing a conclusion from one night
  • Do not assume total sleep time and the minutes labelled REM, Core, or Deep are equally certain
  • Do not conclude that one night with “little deep sleep” means illness or inadequate recovery
  • An operating-system update can change the algorithm, so watch for a step change before and after an update
  • If severe sleepiness, snoring, breathing pauses, or insomnia persist, do not rely on watch numbers alone; consult a healthcare professional

References

  1. Apple. Estimating Sleep Stages from Apple Watch. October 2025. Official technical document
  2. Robbins R, et al. Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults. Sensors. 2024;24:6532. doi:10.3390/s24206532
  3. Lambe R, et al. The accuracy of Apple Watch measurements: a living systematic review and meta-analysis. npj Digital Medicine. 2026. doi:10.1038/s41746-025-02238-1
  4. Walch O, et al. Sleep stage prediction with raw acceleration and photoplethysmography heart rate data derived from a consumer wearable device. Sleep. 2019;42(12):zsz180. doi:10.1093/sleep/zsz180
  5. Walch O. Motion and heart rate from a wrist-worn wearable and labeled sleep from polysomnography, version 1.0.0. PhysioNet. doi:10.13026/hmhs-py35

Medical caution

Apple Watch, Feelmo, and the Paper 17 model do not replace diagnosis by PSG. This page explains research findings; it is not diagnostic or treatment advice.

※ This research explanation does not establish that the Feelmo app is effective.

How accurate is Apple Watch sleep tracking? | Feelmo Research