Can HRV alone read emotion? What we tested across three cohorts
Heart-rate variability (HRV) contains clues about autonomic activity. But if an AI receives only those numbers, can it identify a person’s current emotion—such as “pleasant” or “distressing”?
The conclusion from this internal research is limited:
The HRV-derived information and research model tested here did not predict emotion in unseen people with practically useful accuracy across three public cohorts.
This does not prove that HRV contains no information related to emotion. The sections below explain what was tested and what remains unknown.
Research status
This is an unreviewed internal reanalysis of public data. It is not a clinical trial validating a diagnostic method or the effectiveness of Feelmo itself.
HRV and emotion are not the same thing
HRV describes how much beat-to-beat intervals vary. It is affected by breathing, posture, exercise, sleep, time of day, medication, illness, and many other factors.
Emotion can be described along at least two dimensions:
- Valence: pleasant versus unpleasant
- Arousal: calm versus activated
Even if heart rate rises and HRV changes, the reason could be excitement, anxiety, exercise, or caffeine. Because the same bodily response can have several meanings, assigning an emotion label from HRV alone is a difficult problem.
How was it tested?
We used three public datasets containing emotion labels and cardiac signals.
| Cohort | Main signal | How emotion was elicited |
|---|---|---|
| CASE | ECG | Video |
| AMIGOS | ECG | Short and long videos |
| DEAP | Finger PPG | Music videos |
The research used HRV-derived information and separated people so that data from one person did not appear in both training and evaluation. It then predicted valence and arousal for people unseen during training. This cross-person evaluation matters: a model that merely memorises one person’s patterns is not necessarily useful for a new user.
We tested several tasks, including continuous-value regression, binary classification, and multiclass classification of videos, and statistically combined cohort-level results. The exact feature design, preprocessing, model architecture, and training conditions are proprietary and unpublished. Independent researchers therefore cannot fully reproduce the same model, which is an important limitation.
Reading the result as a beginner
The pooled result for valence was R² = −0.40 (internal reference 95% interval −0.61 to −0.19). The nominal source samples were 30 people in CASE, 40 in AMIGOS, and 32 in DEAP—102 in total—but CASE contributed 29 analysed folds because the internal artefact did not record why one person was absent. The pooled analysis therefore included 101 participant folds. It also combined only three cohorts (k=3). The interval is a normal approximation based on the variance of the leave-one-subject-out R² folds and is fragile with few cohorts and a skewed R² distribution. It should be treated as a reference interval, not as certainty about a broad population.
R² indicates how much variation in emotion the model explains. A value near 1 is better; 0 means performance similar to simply returning the group average. A negative R² does not mean a negative emotion. It means the prediction was worse than a method that simply answered with the average.
The direction was consistent across the three cohorts: the HRV-derived information and research model tested here did not generalise to people unseen during training. However, strict two-sided equivalence testing was not established in every analysis. We therefore cannot conclude that the true effect is exactly zero or that every possible AI model must fail.
What the study does and does not support
What it supports
- Under the tested conditions, the HRV-derived information and research model did not generalise to unseen people across CASE, AMIGOS, and DEAP
- Including both ECG and PPG still did not produce practically useful valence prediction under the same analysis conditions
- Products should be cautious about assigning an emotion label from HRV alone
What it does not support
- It does not prove that HRV contains no information related to emotion
- It does not prove that future models, different features, foundation models, or larger datasets must fail
- It does not rule out within-person estimation that combines a long-term personal baseline with breathing, movement, voice, context, and self-reports
- It does not show that HRV is useless as one clue when considering sleep, recovery, stress, or other physiological contexts
A model that “did not work” still provides important evidence
AI research should report not only high accuracy, but also the conditions under which a model failed to generalise. Clearly defining that boundary reduces overclaiming to users and makes the next data needs more concrete.
What should be tested next?
Improving accuracy requires more than simply making the model larger. The following should be evaluated in order:
- Follow the same person over time and learn changes from that person’s usual pattern
- Record contextual factors that alter HRV, such as sleep, breathing, movement, time of day, and caffeine
- Collect self-reported “pleasant / unpleasant” labels at the same time as the measurement
- Re-evaluate performance on people, devices, and environments excluded from training
- Report calibration, failure cases, and uncertainty—not only accuracy
How this affects Feelmo’s design
Feelmo summarises available HRV-derived information and the person’s own past trends through proprietary on-device processing for review. Details are unpublished. It does not directly infer a specific emotion from HRV.
The AI display should never outrank what the person actually feels. Feelmo avoids declaring an emotion label and instead uses its product display as a prompt for the person to check their own experience and context.
Research resources
- Internal analysis: Paper 1 (unreviewed and not yet generally public). Analysis code and person-level intermediate artefacts will remain unpublished until reproducibility, licensing, and privacy audits are complete
- CASE: Sharma K, et al. A dataset of continuous affect annotations and physiological signals for emotion analysis. Scientific Data. 2019. doi:10.1038/s41597-019-0209-0
- AMIGOS: Miranda-Correa JA, et al. AMIGOS: A Dataset for Affect, Personality and Mood Research on Individuals and Groups. IEEE Transactions on Affective Computing. 2021. doi:10.1109/TAFFC.2018.2884461
- DEAP: Koelstra S, et al. DEAP: A Database for Emotion Analysis Using Physiological Signals. IEEE Transactions on Affective Computing. 2012. doi:10.1109/T-AFFC.2011.15
- Review of HRV and emotion: Kreibig SD. Autonomic nervous system activity in emotion: A review. Biological Psychology. 2010. doi:10.1016/j.biopsycho.2010.03.010
※ This internal research does not establish that the Feelmo app is effective.