The Science of Sleep Tracking: What Wearables Measure and What They Miss

Maya Chen

Maya Chen

July 7, 2026

The Science of Sleep Tracking: What Wearables Measure and What They Miss

Sleep tracking is one of the most popular features in consumer wearables—Apple Watch, Fitbit, Oura Ring, Garmin, Whoop, and Samsung Galaxy Watch all offer sleep stage analysis, sleep score calculations, and detailed nightly summaries. Millions of people wake up and check how they slept, using the data to inform decisions about lifestyle, training load, and recovery. The data is compelling, the visualizations are satisfying, and the promise—quantified sleep insight without a clinic visit—is genuinely appealing.

What the apps often don’t tell you is how inaccurate these measurements can be compared to clinical gold standards, which metrics are reasonably reliable and which are essentially algorithmic guesses, and whether this information actually helps people sleep better or just adds another source of anxiety. The science here is more interesting and more complicated than the marketing suggests.

The Clinical Gold Standard: Polysomnography

The definitive measurement of sleep is polysomnography (PSG)—the sleep study conducted in a clinical lab. A PSG typically measures:

  • EEG (electroencephalogram): Brain electrical activity via electrodes on the scalp, the primary measure of sleep stage
  • EOG (electrooculogram): Eye movement to detect REM sleep
  • EMG (electromyogram): Muscle activity, particularly jaw and leg muscles
  • ECG/heart rate: Cardiac activity
  • Pulse oximetry: Blood oxygen saturation
  • Respiratory effort and flow: To detect breathing irregularities

From this data, sleep technologists can classify each 30-second “epoch” of sleep into stages: Wake, N1 (light non-REM sleep), N2 (moderate non-REM), N3 (deep slow-wave sleep), and REM (rapid eye movement sleep). The EEG is central to this classification—specific brain wave patterns (sleep spindles and K-complexes in N2, delta waves in N3, desynchronized activity in REM) are what actually define the sleep stages.

Consumer wearables measure none of this EEG data. Electrodes on the wrist cannot detect brain electrical activity with any useful fidelity. Wearables instead use proxy measurements—primarily heart rate, heart rate variability, movement (via accelerometer), and in some devices optical blood oxygen—to infer sleep stages through machine learning models trained on populations where wearable and PSG data were collected simultaneously.

What Wearables Actually Measure

Movement (actigraphy): An accelerometer detects physical motion. When you move a lot, you’re likely awake or in light sleep. When you’re still, you’re likely asleep. This is the most reliable measure wearables make, and simple actigraphy is fairly accurate at detecting sleep onset, wake time, and total sleep time—typically within 10–20 minutes of PSG for most sleepers under most conditions.

The limitation: actigraphy can’t distinguish between different sleep stages—N1, N2, N3, and REM all involve similar levels of stillness. It also confuses sleepers who naturally move a lot (or sleep partners who move, if you share a bed) with those who don’t. A sleeper lying still but awake (common in insomnia) will register as asleep on actigraphy alone.

Sleep study polysomnography lab with EEG electrodes measuring brain wave activity

Heart rate and heart rate variability (HRV): Heart rate is measured optically in most wearables—a green LED illuminates skin, and a photodetector measures how much light is absorbed by blood in capillaries, which varies with each heartbeat (photoplethysmography, or PPG). This provides beat-to-beat heart rate data from which HRV can be derived.

Heart rate patterns do correlate with sleep stages: heart rate is typically lower in deep sleep, variable in REM, and higher in light sleep and wake. HRV tends to be higher in certain sleep stages. The correlations are real but messy—the relationship between heart rate patterns and sleep stages varies substantially between individuals, with sleep disorders, and with factors like alcohol consumption, medications, and cardiovascular conditions.

Respiratory rate: Some wearables estimate respiratory rate from the photoplethysmography signal. Breathing causes subtle variations in blood volume detected by the optical sensor. Respiratory rate slows in deeper sleep and changes during REM. This adds another input to sleep stage models but is a derived measurement with its own error characteristics.

Blood oxygen (SpO2): Some wearables use red and infrared LEDs to estimate blood oxygen saturation during sleep, which can indicate sleep apnea-related oxygen desaturations. The measurement is less accurate than a clinical pulse oximeter and typically presented as a range rather than a continuous reading on most consumer devices. When significantly abnormal patterns are detected, it’s a reason to seek clinical evaluation, not a diagnosis.

Skin temperature: Several devices (Oura Ring, Garmin) measure skin temperature, which provides information about autonomic nervous system state, illness, menstrual cycle phase, and potentially sleep quality. Elevated skin temperature during sleep correlates with poorer sleep quality. This is a useful signal, but its relationship to specific sleep stages is indirect.

Sleep Stage Accuracy: What the Research Shows

Multiple independent studies have compared consumer wearable sleep staging to simultaneous polysomnography. The results are consistently in the same direction: wearables do better than chance but worse than clinical measurement, with specific patterns of error.

A 2022 meta-analysis published in JAMA Network Open that reviewed 22 validation studies found that consumer sleep trackers generally:

  • Accurately detect total sleep time within about 10–30 minutes on average
  • Overestimate deep (N3) sleep in many subjects
  • Show highly variable accuracy for REM sleep detection (sensitivity ranges from 65–85% across devices and studies)
  • Are poor at detecting N1 light sleep specifically
  • Perform worse for sleepers with sleep disorders, compared to the healthy young adults typically used in validation studies

The overall epoch-by-epoch accuracy (correctly classifying each 30-second period) typically ranges from 60–80%, compared to inter-rater agreement between sleep technologists reviewing the same PSG data (typically 85–90%). This means consumer sleep stage data is in a gray zone: better than nothing, but not reliable enough to diagnose sleep disorders or make individual clinical decisions.

Heart rate variability data visualization from a wearable health sensor

Sleep Scores: Useful or Misleading?

Most wearables reduce nightly sleep data to a single score (Garmin’s Body Battery, Apple Watch’s sleep data, Oura’s Sleep Score). These scores aggregate multiple metrics—duration, stage distribution, consistency, HRV—into a number designed to be legible and actionable.

The problem is that the scoring algorithms are proprietary, differ between devices, and their clinical validity hasn’t been independently established for most products. Two devices worn simultaneously by the same person on the same night can produce substantially different scores. Whether a score of 78 versus 82 means anything at all is not clear from published research.

What scores can do is track trends for an individual over time—if your score consistently declines during a period of stress, alcohol use, or illness, that trend is likely meaningful even if the absolute number isn’t clinically validated. Intra-individual consistency (does your score go down when you know you slept badly?) may be more useful than absolute accuracy for population benchmarks.

Orthosomnia: When Sleep Tracking Creates Problems

One unintended consequence of sleep tracking has been documented in clinical literature: orthosomnia—a term coined in a 2017 paper in the Journal of Clinical Sleep Medicine for the phenomenon where patients become so focused on achieving perfect sleep data that their anxiety about sleep tracking itself impairs their sleep.

People who become preoccupied with their sleep scores, who modify their behavior to optimize tracked metrics rather than actual rest, or who experience increased anxiety when their sleep data is “bad” may be worse off with tracking than without. This isn’t a universal outcome—many people find tracking helpful—but it’s a real phenomenon that clinicians have observed and that the wearable companies rarely acknowledge prominently.

The irony is sharp: a device designed to improve sleep quality can, in certain individuals and circumstances, make sleep worse.

Where Sleep Tracking Adds Genuine Value

Despite its limitations, consumer sleep tracking has real use cases where the data, understood properly, is genuinely useful:

Tracking trends over time: Longitudinal data on sleep duration, consistency, and heart rate patterns during sleep can reveal how lifestyle changes, travel, stress, or alcohol affect sleep quality—even if individual nightly data is imprecise.

Identifying potential sleep apnea: Tracking devices with SpO2 monitoring can flag patterns suggestive of sleep-disordered breathing—frequent oxygen desaturations, elevated resting heart rate, unrefreshing sleep despite adequate duration. This doesn’t diagnose apnea, but it’s a reason to seek a proper sleep study. Given how commonly sleep apnea goes undiagnosed, this is a genuine public health benefit.

Measuring HRV as a recovery metric: HRV measured during sleep is one of the more reliable outputs from consumer wearables and is used by athletes as an indicator of training readiness and recovery status. The HRV literature supports its utility for this purpose, even though absolute accuracy of sleep staging from the same device is limited.

Consistency and regularity: Tracking when you go to sleep and wake up—much easier to do accurately than staging—can help establish and reinforce consistent sleep schedules, which is one of the most well-supported sleep hygiene behaviors.

The Honest Summary

Consumer sleep trackers are good at measuring total sleep duration and detecting wake periods. They’re passable at rough approximations of sleep stage distributions for healthy sleepers. They’re significantly less accurate for individuals with sleep disorders, who may be the people most interested in tracking their sleep. Their clinical utility for diagnosing specific sleep conditions is limited—a wearable cannot replace a sleep study for anyone with suspected sleep apnea, insomnia disorder, or other sleep pathology.

Used as a way to monitor long-term trends and notice when something is wrong—rather than as a precise clinical measurement of each night’s architecture—they offer something genuinely useful. The key is calibrated expectations: understanding what the sensors actually measure, what the algorithms infer from those measurements, and where the uncertainty lies. With that context, the data can be informative without becoming a new source of anxiety about something most people already find stressful enough.

More articles for you