The short version: every consumer sleep tracker is good at telling you when you were asleep and bad at telling you when you were awake. Across every independent study we could find, sleep-detection sensitivity runs 0.91 to 0.99 while wake-detection specificity runs 0.18 to 0.54. A device that calls quiet wakefulness "sleep" will overstate your total sleep, understate your night wakings, and inflate your sleep efficiency — in one case by about 45 minutes in both directions.
The four-stage breakdown on your app's graph — light, deep, REM — agrees with laboratory scoring at a kappa of roughly 0.21 to 0.65 depending on device. For reference, trained human scorers agree with each other at kappa 0.76, and only 0.24 on light (N1) sleep.
And a correction up front, because it is repeated everywhere: the most-cited multi-device validation study, Chinoy 2021, did not test Oura, Whoop, Apple Watch or the Fitbit Sense. We say below what it actually tested.
What the validation studies found
Chinoy 2021 — the most-cited study, and what it actually measured
**Chinoy ND, Cuellar JA, Huwa KE, et al. SLEEP 2021;44(5):zsaa291. n=34 healthy young adults (mean age 28.1 ± 3.9), three nights of full polysomnography. Funded through US Navy research support (Naval Health Research Center); no manufacturer funding identified.**
Devices tested: Fatigue Science Readiband, Fitbit Alta HR, Garmin Fenix 5S, Garmin Vivosmart 3, EarlySense Live, ResMed S+, SleepScore Max — plus an Actiwatch 2 research actigraph as a reference comparator. No Oura, no Whoop, no Apple Watch.
| Device | Sleep sens. | Wake spec. | Accuracy | Total sleep bias | Night-waking bias | Sleep efficiency bias |
|---|---|---|---|---|---|---|
| Actiwatch 2 (research device) | 0.97 | 0.54 | 0.90 | +23.9 min | −16.6 min | +5.0% |
| Fitbit Alta HR | 0.98 | 0.44 | 0.91 | +2.6 (ns) | −2.1 (ns) | +0.9% (ns) |
| ResMed S+ | 0.97 | 0.51 | 0.90 | −0.3 (ns) | −3.4 (ns) | 0.0% (ns) |
| EarlySense Live | 0.96 | 0.48 | 0.87 | +13.6 | −15.3 | +2.9% |
| SleepScore Max | 0.95 | 0.38 | 0.85 | +7.5 (ns) | −12.1 | +1.6% (ns) |
| Fatigue Science Readiband | 0.93 | 0.35 | 0.85 | +13.3 (ns) | −12.5 (ns) | +2.8% (ns) |
| Garmin Fenix 5S | 0.99 | 0.18 | 0.88 | +43.7 min | −49.5 min | +10.6% |
| Garmin Vivosmart 3 | 0.99 | 0.19 | 0.87 | +46.8 min | −47.6 min | +10.1% |
(ns = not statistically significant)
The Garmin numbers are the headline. A wake specificity of 0.18 means the device correctly identified fewer than one in five epochs that the laboratory scored as wake. The consequence is an overstatement of total sleep by about three quarters of an hour and an understatement of time awake in bed by about the same.
On stages: all six staging devices differed significantly from polysomnography on light sleep, mostly overestimating it. Three overestimated deep sleep and three underestimated REM. The authors' own conclusion was that "device sleep stage assessments were inconsistent."
Chinoy 2022 — the Oura paper, and a weaker reference standard
**Chinoy ND, et al. Nature and Science of Sleep 2022;14.** n=21 healthy young adults, multiple nights at home, funded by the Office of Naval Research; authors reported no conflicts and no device-company involvement.
Important design caveat: the reference was a Dreem 2 EEG headband, not full polysomnography.
| Device | Sleep sens. | Wake spec. |
|---|---|---|
| Polar Vantage V Titan | 0.96 | 0.35 |
| Actiwatch 2 | 0.95 | 0.35 |
| Fatigue Science Readiband v5 | 0.94 | 0.40 |
| Oura Ring Gen 2 | 0.94 | 0.41 |
| Fitbit Inspire HR | 0.93 | 0.45 |
Stage performance was described as "mixed and highly variable," with better accuracy on nights of consolidated sleep.
The best independent data on Whoop and Apple Watch
***Sleep Advances 2025;6(2):zpaf021.* n=62 adults (52 men, mean age 46.0 ± 12.6) with full polysomnography. Funding not specified in the source we read.
| Device | Subset | Kappa | Sleep sens. | Wake spec. | Total sleep bias | Night-waking bias |
|---|---|---|---|---|---|---|
| Apple Watch 8 | 20 | 0.53 | 96.3% | 52.2% | +19.6 min | −21.2 min |
| Fitbit Sense | 37 | 0.42 | 93.3% | 48.8% | +6.3 min | −12.4 min |
| Fitbit Charge 5 | 39 | 0.41 | 91.7% | 47.5% | +11.1 min | −16.9 min |
| Whoop 4.0 | 40 | 0.37 | 93.6% | 40.1% | +24.5 min | −19.2 min |
| Withings ScanWatch | 41 | 0.22 | 94.3% | 31.1% | +39.9 min | −47.9 min |
| Garmin Vivosmart 4 | 25 | 0.21 | 95.9% | 29.4% | +38.4 min | −38.3 min |
No device exceeded about 52% wake specificity. Whoop 4.0 — marketed hardest on sleep and recovery — came fourth of six on kappa, behind both Fitbits.
The manufacturer-funded study, and how it differs
**Robbins R, et al. Sensors 2024;24(20):6532. n=35 healthy adults aged 20–50, one night in hospital. Funded by Oura Ring Inc.**, with Harvard Catalyst and NIH support.
- Sleep/wake sensitivity ≥95% for all three devices; epoch agreement Oura 92%, Apple 93%, Fitbit 91%
- Four-stage kappa 0.52–0.65
- Apple underestimated deep sleep by 43 minutes and overestimated light sleep by 45. Fitbit was −15 and +18
- Oura showed no significant nightly-total differences
- Apple and Fitbit failed to record data for 6 and 2 participants respectively
Oura's marketing line that it was "5% more accurate than Apple Watch and 10% more accurate than Fitbit" comes from this study — an Oura-funded, single-night study of 35 healthy people. It is not fabricated, and it is also not independent.
A discrepancy worth noting: the journal page we read gave wake-stage kappas of Oura 0.60, Apple 0.60, Fitbit 0.52, while Oura's own blog reports 0.65, 0.60 and 0.55. We could not resolve which figure is which.
What happens in people who actually have sleep problems
**Herberger S, et al. Scientific Reports 2025.** 45 participants from a university sleep-lab clinical cohort with a mix of sleep disorders. Funding not stated in the page we read.
Oura Gen 3:
- Total sleep bias +11.7 minutes (SD 37.1)
- Sleep/wake kappa 0.43
- Overall staging accuracy 53.2% — against Oura's own marketing claim of 79%
- About 85% sleep/wake accuracy
The authors noted that individual-level differences "often remained large." Two competing rings in the same study scored 50.5% and 35.1% staging accuracy.
This is the most important study in the article for most readers, because the people most likely to buy a sleep tracker are the people who think they sleep badly — and that is the population where the devices perform worst.
What has not been validated
We could find no independent polysomnography validation of the Oura Ring 4, Whoop 5.0 or MG, Apple Watch Series 9 or 10 staging, the Fitbit Charge 6, or any recent Garmin. The accuracy claims for those products rest on manufacturer pages. That is not an accusation; it is the state of the literature, and it means the device you can buy today has probably never been independently tested.
The ceiling problem: humans don't agree either
A device can only be measured against a reference, and the reference is noisy.
**Lee YJ, et al. J Clin Sleep Med 2022** — meta-analysis of inter-rater reliability in polysomnography scoring:
- Pooled Cohen's kappa 0.76 overall
- Per stage: Wake 0.70, N1 (light) 0.24, N2 0.57, N3 (deep) 0.57, REM 0.69
**Rosenberg RS & Van Hout S, J Clin Sleep Med 2013;9(1):81–87 — the American Academy of Sleep Medicine's inter-scorer reliability programme, over 3.2 million epoch decisions from more than 2,500 scorers: 82.6% average agreement with the majority score. Per stage: Wake 84.1%, N1 63.0%, N2 ~85%, N3 67.4%**. (This paper reports percentage agreement, not kappa — don't quote a kappa from it.)
So trained humans disagree on roughly one epoch in six, and score light sleep at "fair" agreement. When your app shows "1 h 12 m of deep sleep," bear in mind that two certified technicians reading the same EEG would agree on deep sleep at kappa 0.57.
But don't let this argument be abused. It is sometimes deployed to excuse devices generally, and it doesn't stretch that far: **wake is the stage humans score well** (kappa 0.70, 84% agreement). The devices' poor wake specificity is not explained away by reference noise. That failure is the devices'.
Wake detection, and why it matters most
Every independent study above shows the same pattern. High sleep sensitivity, low wake specificity, errors all in one direction:
- Chinoy 2021: wake specificity 0.18–0.54
- Chinoy 2022: 0.35–0.45 (Oura 0.41)
- Sleep Advances 2025: 0.29–0.52
The devices call quiet wakefulness sleep. If you lie still in the dark with your eyes open, a wrist or finger sensor sees a still body and a resting heart rate, and scores it as sleep. The results are overstated total sleep time, understated time awake after sleep onset, and inflated sleep efficiency.
This is precisely backwards for the person most likely to care. If you wake at 3 am and lie there for an hour, that hour is the single most clinically relevant feature of your night — and it is the thing the device is worst at seeing. Insomnia is characterised by exactly this, and the device will tell you that you slept well.
How far can we push that claim? Chinoy 2022 reports better accuracy on consolidated-sleep nights, and the Herberger clinical cohort shows large individual errors. But we did not find a dedicated polysomnography validation in diagnosed insomnia, so "accuracy degrades in insomnia" is supported indirectly rather than established. We are flagging that rather than asserting it.
Heart rate variability: the one thing that works
Staging is the headline feature and the worst one. Overnight heart-rate variability is the quiet success.
**Dial MF, et al. Physiological Reports 2025. n=13 healthy adults, 536 nights**, against ambulatory ECG, nocturnal RMSSD. (Funding not retrieved.)
| Device | Concordance | Mean absolute % error |
|---|---|---|
| Oura Gen 4 | 0.99 | 5.96% |
| Oura Gen 3 | 0.97 | 7.15% |
| Whoop 4.0 | 0.94 | 8.17% |
| Garmin Fenix 6 | 0.87 | 10.52% |
| Polar Grit X Pro | 0.82 | 16.32% |
A finger ring measuring HRV during sleep is genuinely close to the gold standard. Two caveats: the sample is thirteen people, and this is overnight HRV from a ring, which is not the same measurement as a watch's spot HRV at the wrist.
On that point, two studies via a secondary source (journals unconfirmed): O'Grady et al. 2024 (n=39, Apple Watch Series 9 and Ultra 2) found HRV underestimated by 8.31 ms with a 28.9% mean error, outside the equivalence margin; Bonneval et al. 2025 (n=78, Series 6, at rest) found a 31.3% error on beat intervals. We could not verify the journals for these two and are not treating them as settled.
Whoop's "99% accurate" HRV claim is a Whoop claim, not an independent finding. The independent figure above — a mean error of 8.17% for Whoop 4.0 — is good, and it is not 99%.
We did not retrieve device-versus-ECG resting heart rate figures. Resting heart rate is generally the easiest of these measurements and we would expect it to be the most accurate, but we are not printing a number we don't have.
Readiness, recovery and sleep scores
We found no peer-reviewed evidence that Oura Readiness, Whoop Recovery or Garmin Body Battery predicts injury, illness, performance, or any health outcome.
There is a structural reason for this, and it is worth understanding: there is no gold standard for "recovery." You cannot validate a score against a thing that has no independent measurement. These are proprietary composites of inputs — some of which (HRV, resting heart rate) are measured well, some of which (sleep stages) are measured badly — combined by an undisclosed formula into a number out of 100.
Whoop's accuracy claims are about the measurement of HRV, not about the score built on top of it. Those are different claims and the marketing does not always distinguish them.
We should be precise about our own limits here: we found no outcome study, which is not the same as proving none exists. But the burden of proof sits with the company selling the score.
The algorithm changes under you
Whoop changed its sleep staging algorithm on 24 February 2025. Whoop says the update raised four-stage accuracy by 7% and wake precision by 3%, citing two years of research at Central Queensland University and the University of Arizona. Those are Whoop's figures.
Whoop staff stated in May 2025 that "your past sleep data stays the same" — history was not rescored. Users reported shifts in awake time and restorative sleep after the update.
Oura announced Sleep Staging Algorithm 2.0 in November 2022, claiming 79% four-stage agreement based on 1,200+ nights. We could not confirm whether historical data was recalculated. We found no documented silent change for Fitbit, which is not evidence that none occurred.
Why this matters: the entire proposition of a sleep tracker is longitudinal — watch your trends, see what helps. A silent algorithm change breaks that. A step down in your deep sleep in March 2025 might be a change in your sleep or a change in the software, and you have no way to tell. If you compare this year to last year on a Whoop, you are comparing two different measuring instruments.
Orthosomnia: real, but handle the numbers carefully
**Baron KG, Abbott S, Jao N, Manalo N, Mullen R. "Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?" J Clin Sleep Med 2017;13(2):351–354. This is the paper that named the phenomenon — a case series** from a sleep clinic describing patients whose preoccupation with tracker data was itself worsening their sleep. The authors recommend cognitive behavioural therapy for insomnia and integrating tracker data into treatment rather than banning it.
Because it is a case series, it provides no prevalence.
On the prevalence figure you may have seen: Goel R, Quan S, Ejikeme C, Weaver MD, Sablan J, Czeisler C, Robbins R. "Exploring the Prevalence of Sleep Tracking and Orthosomnia in a National Survey of Adults in the US." SLEEP 2026;49(Supplement_1):A246.
- It is a conference abstract, not a peer-reviewed paper
- n=1,280 trackers, of whom 30.9% screened positive — not 30.9% of the whole sample
- The instrument was an adapted two-item measure asking whether participants felt "nervous, anxious or on edge" about their tracker data. It is not a validated orthosomnia scale, and it measures tracker-related anxiety rather than a diagnosis
- Funded by gift funding from Oura Ring Ltd. The last author is also lead author of the Oura-funded 2024 validation study above
So: orthosomnia is a recognised clinical phenomenon with a named case series behind it, and "roughly a third of sleep trackers have orthosomnia" is not a conclusion this abstract supports. We flag it because the figure is already circulating.
A cross-sectional study, "Prevalence of Orthosomnia in a General Population Sample" (Brain Sciences, November 2024), appears to exist and would be a better source. We could not read it.
What is actually regulated
This is where the distinction between a medical device and a wellness feature matters.
Apple Watch sleep apnoea notification — TGA approval reported 29 May 2025, for Series 9, 10 and Ultra 2, users 18 and over. It uses the accelerometer and reviews data over 30-day periods.
Apple's own validation white paper (September 2024) is unusually transparent, and the numbers deserve to be read:
- n=1,499 enrolled, 1,448 completed. Mean age 46, 56.5% female, mean BMI 32
- Enriched cohort: 559 normal, 362 mild, 216 moderate, 201 severe
- Reference: at least two nights of home sleep apnoea testing scored to AASM criteria
- Overall sensitivity 66.3% (95% CI 62.2–70.3); specificity 98.5% (98.0–99.0)
- Severe OSA sensitivity 89.1% (83.7–93.2)
- Moderate OSA sensitivity 43.4% (36.5–50.5)
That last number is the one to carry away: the Apple Watch misses roughly 57% of moderate obstructive sleep apnoea. Specificity is excellent — if it notifies you, take it seriously — but a silent watch is not a negative test. There is no independent validation, and it is not a diagnostic device.
Samsung received TGA approval a little over a week earlier (around 20 May 2025, date not confirmed), for Galaxy Watch 4 and later, users 22 and over, using blood oxygen over two-night periods. We could not find Samsung's validation figures.
What we could not confirm: the ARTG entries and TGA classification for either product, and the regulatory status of each device's other sleep features. The general shape — that a screening or diagnostic claim brings a product into the TGA framework while sleep staging and readiness scores are marketed as wellness — is our understanding of how the regime works, but we are not asserting the per-feature split without checking the register. If you want certainty on a specific product, the ARTG is searchable at tga.gov.au.
What it costs over three years
Retrieved 1 October 2026. Prices move and promotional pricing is heavy in this category.
| Device | Upfront (AUD) | Subscription | 3-year total |
|---|---|---|---|
| Apple Watch Series 11 (42 mm GPS) | $679 RRP, seen at $427–447 | None | $679 (or ~$447 on sale) |
| Oura Ring 4 | $398 RRP, $317 at JB Hi-Fi | $9.99/mo or $109.99/yr | $728 on the annual plan ($647 with the sale ring) |
| Whoop 5.0 | $99–139 | ~$300/yr | $999–1,039 |
| Whoop MG | $249 | ~$450/yr | $1,599 |
| Garmin | Device price only | None for core features | Device price |
| Fitbit | Charge 6 price not retrieved | Premium ~$123/yr; Google Health Premium reported at $15.30/mo | Not computed |
Notes, because the fine print is where the money is:
- Oura. Membership is required for the Sleep, Readiness and Activity scores, the Advisor, and the detailed breakdowns. Without it, the ring still records and CSV export still works — so you keep your data, you lose the interpretation.
- Whoop. The Australian hardware/membership split comes from a single trade report (29 May 2026) describing it as a "limited-time trial in select markets," with existing members keeping old pricing for a time. The US page lists US$199/239/359 per year for its three tiers with a 30-day refund window. We could not confirm what happens to the strap if you stop paying, and we are not repeating the common claim that it bricks without a source.
- Apple. The sleep apnoea notification costs nothing extra. This is the only device here with no subscription at all.
- Garmin. Garmin Connect+ exists at US$6.99/month, and Garmin states existing features stay free. We did not retrieve an Australian price.
- Fitbit. A report of Google Health Premium at $15.30/month and a screenless "Fitbit Air" at around $150 is a pre-launch item we could not confirm has shipped.
The pattern is clear enough: a ring or a strap costs $650 to $1,600 over three years and most of that is subscription. The one-off watch costs a fraction of it and has the only TGA-approved sleep feature in the group.
So what is a sleep tracker good for?
Genuinely useful:
- When you went to bed and when you got up. Trivially measured, and the single most actionable sleep variable there is. Consistency of timing matters more for how you feel than any stage breakdown.
- Total sleep duration as a trend. Not any given night — the error bars are tens of minutes — but a month-long average is informative.
- Overnight heart rate and HRV, on a ring in particular, where the measurement is close to ECG.
- Noticing something has changed. A sustained shift in resting heart rate or HRV is a real signal, even if the "readiness score" built on it isn't validated.
- Behaviour change through attention. If wearing it makes you go to bed earlier, that is a real benefit regardless of measurement accuracy.
It cannot do:
- Tell you how much deep sleep or REM you got. Kappa 0.21 to 0.65 against a reference that humans themselves score at 0.57 for deep sleep. The pie chart is an estimate dressed as a measurement.
- Reliably detect being awake. Wake specificity 0.18–0.54 in every independent study. If you lay awake, your device probably didn't notice.
- Diagnose anything. The best-validated feature in the group — Apple's apnoea notification — misses about 57% of moderate OSA. A silent watch is not reassurance.
- Support a verdict on one night. Single-night totals carry errors of 20 to 45 minutes.
- Be compared across an algorithm update. Whoop's February 2025 change was not applied retrospectively.
And if you are sleeping badly, see a GP rather than your app. The devices are worst in precisely the population that needs an answer — in a clinical cohort, Oura's staging accuracy was 53% against the 79% on its marketing page. Snoring with daytime sleepiness, long night wakings, or unrefreshing sleep despite adequate time in bed are reasons for a referral, not reasons for a better ring.
Frequently asked
Which sleep tracker is the most accurate? For sleep-versus-wake and staging, the best independent data put Apple Watch slightly ahead (kappa 0.53) of Fitbit (0.41–0.42), Whoop 4.0 (0.37) and Garmin (0.21) in a 62-person polysomnography study. An Oura-funded study of 35 healthy people put Oura first. In a clinical population with actual sleep disorders, Oura's staging accuracy was 53%. The Oura Ring 4, Whoop 5.0, Apple Watch Series 9 and 10 staging and Fitbit Charge 6 have no independent polysomnography validation we could find.
Can a sleep tracker measure deep sleep and REM? Only approximately. Four-stage agreement with laboratory scoring runs from kappa 0.21 to 0.65 depending on the device. Trained human scorers agree with each other at 0.57 for deep sleep and 0.24 for light sleep, so the reference standard is itself noisy — but the devices sit well below even that ceiling. Treat the stage breakdown as an estimate, not a measurement.
Why does my tracker say I slept well when I know I was awake? Because that is the devices' central weakness. Wake-detection specificity runs 0.18 to 0.54 across every independent study: the device sees a still body and a low heart rate and scores quiet wakefulness as sleep. Two Garmin models overstated total sleep by about 45 minutes. Your own experience of the night is the better record.
Is the Apple Watch sleep apnoea feature any good? It is TGA-approved (May 2025) and its specificity is excellent at 98.5%, so a notification is worth acting on. But Apple's own validation data show overall sensitivity of 66.3% and only 43.4% for moderate sleep apnoea — it misses roughly 57% of moderate cases. It is a screening prompt, not a diagnostic test, and not getting a notification does not mean you don't have apnoea.
Are readiness and recovery scores validated? We found no peer-reviewed evidence that any of them predicts injury, illness or performance. There is a structural reason: "recovery" has no independent gold-standard measurement, so there is nothing to validate the score against. The underlying HRV measurement, at least on a ring, is accurate; the composite built from it is proprietary and untested.
Does orthosomnia actually exist? Yes, as a described clinical phenomenon — a 2017 sleep-clinic case series named it, describing patients whose preoccupation with tracker data worsened their sleep. The widely circulated "30.9% of trackers" figure comes from a 2026 conference abstract, not a peer-reviewed paper, using a two-item anxiety screen rather than a validated instrument, with gift funding from a ring manufacturer. Treat the phenomenon as real and the prevalence figure as unestablished.
Do I need the subscription? With Oura, the ring keeps recording and you can still export your data without a membership — you lose the scores and interpretation. Whoop is membership-based and we could not confirm what the hardware does if you stop paying. Garmin's core sleep features are free. The Apple Watch has no subscription at all and happens to carry the only TGA-approved sleep feature in the group.
Last reviewed 1 October 2026. Manufacturer funding is flagged on every validation study where the source discloses it. Prices are Australian retail at the date of review. Where we could not verify a figure or could not find a study, we have said so rather than filling the gap. This page is general information, not medical advice: persistent poor sleep is a reason to see a doctor, not a reason to buy a different device.