Is HRV really “anonymous data”? Heart-rate variability and re-identification risk
If names and email addresses are removed, does heart-rate variability (HRV) become anonymous health data? In an internal reaudit using public data, HRV sections separated from the same recording could be linked within a candidate set of 202 people. Details of how the sections and comparison were constructed are unpublished.
This does not mean that “anyone can verify identity from Apple Watch” or that HRV can replace a fingerprint. The result came from a closed candidate set within the same measurement session. It did not establish that a person can be tracked across days or devices. A separate internal model also failed its evaluation across sleep and wake.
Research status
This page explains an unreviewed internal secondary analysis of existing public data. It did not directly test Apple Watch data and does not demonstrate the clinical performance of a medical device, authentication product, or Feelmo itself.
What question did the study ask?
HRV summarises changes in the interval from one heartbeat to the next. Wellness products sometimes present it alongside sleep or recovery, but HRV does not directly measure either. Age, body size, autonomic regulation of the heart, breathing, posture, lifestyle, and measurement conditions all contribute, so recordings may also contain patterns that recur within a person.
The study asked the following question:
Using a research comparison based on HRV-derived information, can two sections separated from the same resting ECG recording be linked among a set of candidates?
This was not an authentication test for unlocking a passcode. It tested same-record linkability—whether anonymised record fragments could be connected. That is one step before re-identification with a real name.
Data and methods
Paper 3 as a whole combined records from 1,227 people across six public cohorts: Autonomic Aging, Fitbit PMData, MMASH, Fantasia, MIT-BIH Normal Sinus Rhythm, and CinC 2018. The data included resting ECG, RR intervals during walking, sleep-related recordings, and longer-term wearable records.
The research used HRV-derived information and kept records from the same person from being mixed between training and evaluation partitions.
The public-facing reaudit used the Autonomic Aging cohort. Of 211 people held out from training, 202 met prespecified eligibility conditions. Separate sections from the same recording were used, but the exact segmentation, selection, and comparison rules are unpublished.
The research comparison used HRV-derived information. “Top five” means that the correct person appeared somewhere among the five candidates returned. The feature design, section construction, quality and selection conditions, preprocessing, and comparison equation are unpublished. Independent researchers therefore cannot fully reproduce the same analysis, which is an important limitation.
Main results
| Evaluation | Observed | Chance under random ranking | 95% Wilson interval |
|---|---|---|---|
| Top one | 62 / 202 (30.69%) | 0.50% | 24.74–37.36% |
| Top five | 112 / 202 (55.45%) | 2.48% | 48.55–62.13% |
| Top ten | 133 / 202 (65.84%) | 4.95% | 59.06–72.03% |
In a sensitivity analysis reversing the comparison direction, top-five accuracy was 55.94%. A test that permuted only the correct pairings was run 20,000 times. With that number of repetitions, the smallest reporting unit for a one-sided p value is 0.00005. The public aggregate file does not include the number of repetitions that equalled or exceeded the observed result, so the observed p value cannot be recalculated from that file alone.
A bootstrap that kept the 202 candidates fixed and resampled query people gave a conditional 95% interval of 48.51–61.88% for top-five accuracy, close to the Wilson interval. Neither interval guarantees performance for future users or real-world operation.
An important failure: the method did not work the same way across sleep and wake
A separate unreviewed model in Paper 3 failed the prespecified evaluation condition that person-level patterns would remain stable across sleep and wake. That model cannot be compared directly with the procedure above. Among 15 MMASH participants who had both states, the results were:
- Enrol during sleep and match during wake: top five 56.13%
- Enrol during wake and match during sleep: top five 42.99%
The results did not reach the prespecified internal criterion. Its exact value and decision rule are unpublished. Breathing, posture, activity, and neural regulation of heart rate all change between sleep and wake. These results do not establish state-invariant linkability and do not support the idea of a permanent “heartbeat fingerprint.”
Top-five accuracy is not authentication accuracy
This is the share of cases in which the correct person appeared among five candidates in a known, closed set. The study did not evaluate rejection of unknown people, resistance to impersonation, or false-accept rates, so it cannot be compared with the authentication performance of a security product.
What did this show—and not show—about Apple Watch?
This study did not directly evaluate Apple Watch. The six cohorts included ECG, RR intervals, sleep-related recordings, and long-term Fitbit data collected under different conditions. They did not include a prospectively collected Apple Watch dataset following one common protocol.
The following claims therefore cannot be made:
- “Apple Watch HRV can identify a person with high accuracy”
- “Apple Watch can provide biometric authentication”
- “Feelmo identifies users from their heartbeat”
On Apple Watch, the optical sensor, sampling timing, missingness caused by movement, and operating-system processing can all affect the result. Direct performance would require a prospective Apple Watch-specific validation with device, posture, and time of day controlled.
Why does this matter for privacy?
Linkability can matter even when it is far below 100%. A record fragment could narrow a candidate list, and other information—such as sleep timing, age, activity, or location—might narrow it further. This study’s 55.45% result alone does not establish that such narrowing works across days or devices.
At minimum, the following precautionary principles follow:
- Do not assume that detailed RR intervals or HRV features are “anonymous” merely because names were removed
- Treat feature vectors and AI embeddings, not only raw beat intervals, as potentially identifying health information
- Store only the duration and granularity needed, and delete data when it is no longer needed
- Process on the device when possible and minimise server transmission and third-party sharing
- Do not release person-level embeddings, nearest-neighbour tables, or matchable time series unchanged in research datasets
- Include contractual prohibitions on matching and re-identification when data are shared
How this affects Feelmo’s design
Feelmo treats body data such as HRV and sleep as sensitive information that could be linked to a person, not merely as numbers for display. By default, it processes and stores this information on the device and does not automatically send it to Provider servers. See Data and privacy for details.
This research does not support a feature that “identifies someone from HRV.” It supports the opposite design principle: the more information data can reveal, the more reason there is to keep it on the person’s device.
Limitations and next tests
- This is an unreviewed internal secondary analysis and needs independent replication
- Sensors, measurement environments, and preprocessing differ between public cohorts
- The main result comes from separate sections of the same resting ECG recording and does not establish persistence across days
- Nine of the 211 held-out people did not meet prespecified eligibility conditions; the detailed selection rules are unpublished
- Differences in acquisition, including ECG lead configuration, may have affected results across cohorts
- A separate internal evaluation also missed its target and did not establish invariance across sleep and wake
- This was closed-set top-five linkage, not a test of real-world authentication
- Apple Watch was not directly evaluated
- Future work requires prospective Apple Watch-specific data, external cohorts, state-specific evaluation, and open-set tests that include unknown people
Research resources
- Public reaudit: Aggregate JSON without personal information. It contains no person IDs, person-level ranks, features, preprocessing code, or comparison procedure
- Internal analysis: Paper 3 (unreviewed and not yet generally public). Feature design, quality conditions, preprocessing, comparison equations, person-level arrays, and intermediate artefacts remain unpublished
- Scale of analysis: six cohorts, 1,227 people
- Main public result: 202 eligible people among 211 held out from Autonomic Aging; same-record top-five 55.45% (112/202)
- Main data source: Schumann A, Bär KJ. Autonomic Aging. Paper doi:10.1038/s41597-022-01202-y · Data doi:10.13026/2hsy-t491
- State-specific data: Rossi A, et al. A Public Dataset of 24-h Multi-Levels Psycho-Physiological Responses in Young Healthy Adults. Data. 2020;5(4):91. doi:10.3390/data5040091 · MMASH v1.0.0 doi:10.13026/cerq-fc86
- Research status: unreviewed internal secondary analysis of public data
※ This internal research does not establish that the Feelmo app is effective.