What '95% Accurate' Actually Means on a Sleep Tracker Box
Dovy Paukstys
Founder, Komori Care

The Number on the Box
Every sleep tracker wants you to see one number. "95% accurate." "93% agreement with clinical sleep studies." Sometimes there's an asterisk. Usually there isn't.
That number is close to meaningless. And the people printing it know why.
I'm going to show you the trick, then show you the three numbers that actually matter, what real devices score on them, and exactly what to ask a manufacturer if you want a straight answer.
This isn't a hit piece on wearables. Some of them are decent. It's a piece about a measurement convention that lets a good device and a useless device print the same headline.
Key Facts
- Wrist actigraphy scores 86.3% accuracy but only 32.9% specificity — it catches sleep well and misses wake badly (Sleep, 2013)
- Seven consumer trackers vs. a sleep lab: sensitivity ranged 0.93 to 0.99, specificity only 0.18 to 0.54 (Sleep, 2021)
- Sleep periods contain far more sleep than wake, which produces cases of "high agreement but low kappa" (Sleep, 2021)
- Cohen's kappa strips out the credit a device gets from lucky guessing; below 0.60 is generally called inadequate (Biochemia Medica, 2012)
- Human sleep scorers agree with each other only 82.6% of the time, and just 63.0% on the lightest stage of sleep (JCSM, 2013)
First, How Sleep Gets Scored
Sleep labs chop the night into 30-second slices. Each slice is called an epoch, which is just a fancy word for "one small chunk of time." A trained human looks at the brain waves in each epoch and stamps it: asleep or awake, and if asleep, which stage.
A wearable does the same chopping. It looks at wrist motion, or heart rhythm, or both, and stamps each 30-second slice with its own guess.
Accuracy is then simple. Line the two lists up, count how many stamps match, divide by the total. That's the number on the box.
Now watch what goes wrong.
The Trick: You're Asleep Almost the Whole Time
Say you're in bed for 8 hours. That's 960 epochs. Say you sleep 7 of those hours, which is a normal-ish night — measured sleep efficiency in one large validation study ran about 85.6% in people without insomnia and 83.1% in people with it [^1].
So roughly 840 epochs of sleep and 120 epochs of wake.
Now build the world's laziest sleep tracker. It has no sensors. It's a rock. It stamps every single epoch "asleep."
It gets 840 out of 960 right. That rock is 87.5% accurate.
Put it in a box, print "87% agreement with polysomnography," and technically you haven't lied. The rock told you nothing. It cannot find a single minute of the night you spent staring at the ceiling. But the headline number looks respectable.
This is not a hypothetical concern researchers invented. The authors of the biggest actigraphy validation study said it outright: "most of the study sleep period is occupied by sleep, thus the high accuracy of actigraphy is largely explained by high sensitivity" [^1].
Translation: the number is inflated by the easy part of the job.
Sensitivity and Specificity, In Plain Words
This is why researchers stopped reporting accuracy alone decades ago. They split it into two questions.
Sensitivity answers: when you were actually asleep, did the device say asleep? It's graded only on the sleep epochs. Our rock scores a perfect 1.00 here, because it says sleep every time.
Specificity answers: when you were actually awake, did the device say awake? It's graded only on the wake epochs. The rock scores 0.00. Zero. It never once says wake.
Notice that specificity is the harder job and the one people care about. If you're lying awake at 3 a.m., you already know you slept badly. What you want from a tracker is confirmation of how much you were awake and when — which is exactly the number these devices are worst at.
For more on what those metrics mean across different sensor types, see our comparison of sleep tracking methods.
Cohen's Kappa: Grading Against Luck
There's a third number, and it's the one to actually ask for.
Cohen's kappa takes the raw agreement and subtracts out the agreement you'd expect from pure chance given how lopsided the categories are. The formula is straightforward: observed agreement minus expected-by-chance agreement, divided by how much room was left to improve [^2].
Run it on the rock. Observed agreement: 0.875. Expected by chance: also 0.875, because a device that always says "sleep" agrees with an 87.5%-sleep night 87.5% of the time by definition.
Kappa = 0.00. No skill whatsoever. Exactly right.
Now run it on a plausible real device — sensitivity 0.99, specificity 0.18, which is the low end of the seven consumer trackers tested against a sleep lab [^3]. On that same 840/120 night, that device gets about 89% accuracy and a kappa of roughly 0.25.
On the standard interpretation scale, a kappa of 0.21 to 0.39 is called "minimal" agreement, and anything under 0.60 is generally considered inadequate for healthcare use [^2]. So the device with the 89% accuracy claim is scoring "minimal."
That gap — 89% versus minimal — is the entire scam in one line.
One caveat: kappa itself is sensitive to how lopsided the categories are. Sleep researchers have proposed a prevalence-adjusted version (PABAK) precisely for this reason, and the standardized testing framework published in Sleep recommends it for sleep-versus-wake agreement [^4]. If a company reports it, they've read the literature.
What Real Devices Actually Score
Here's the picture from the two validation studies I'd trust most.
Wrist actigraphy, the research-grade reference for decades, tested on 77 adults: accuracy 0.863, sensitivity 0.965, specificity 0.329 1. It finds 96 out of every 100 sleep epochs and 33 out of every 100 wake epochs.
Seven consumer devices tested against polysomnography in 34 healthy young adults: sensitivity 0.93 to 0.99, specificity 0.18 to 0.54 2. Most of them beat actigraphy's 0.39 specificity in that study, which the authors noted was previously actigraphy's greatest weakness. That's real progress. It's also still a coin flip on your worst nights.
A December 2025 preprint (not yet peer reviewed) trained a deep learning model on wrist accelerometry from 453 adults undergoing clinical sleep testing. The authors open by stating flatly that "previous works demonstrated poor wake detection." Their model reports sensitivity 0.87 and specificity 0.78 — a much healthier balance — and they got there by deliberately training on people with low sleep efficiency and lots of arousals 3. That's the honest fix: feed the model the hard nights.
Because it's a preprint, treat those numbers as promising and unconfirmed until they clear peer review.
The Same Problem, One Level Up
Sleep staging has the identical disease. A separate August 2026 preprint (also not yet peer reviewed) argues that photoplethysmography — the green-light blood-flow sensor in wrist trackers and smart rings — "trails EEG-based methods by a substantial margin" at classifying sleep stages, and that the standard 30-second epoch actually hides the signal, because pulse features shift within seconds at stage boundaries 4.
That paper's title tells you where the field is: Rethinking PPG-based Sleep Staging: Datasets, Metrics, and Benchmarks. Researchers are re-examining the yardstick, not just the devices.
Which brings up the humbling part.
Even the Humans Disagree
The American Academy of Sleep Medicine ran more than 2,500 certified scorers through the same sleep records. Overall stage agreement: 82.6%. For deep N3 sleep: 67.4%. For light N1 sleep: 63.0% 5.
Sit with that. A device advertising "95% accurate sleep staging" is claiming to beat the humans who wrote the answer key by twelve points. On a task where the humans themselves can't agree two-thirds of the time.
That claim isn't impressive. It's a red flag about what they compared against.
The Lookup Table
| The number | The question it answers | What the "always asleep" rock scores | What to ask for |
|---|---|---|---|
| Accuracy | Of all 30-second slices, how many matched? | ~88% | Nothing. Ignore it alone. |
| Sensitivity | When you were asleep, did it say asleep? | 1.00 (perfect) | Report it, but it's the easy half |
| Specificity | When you were awake, did it say awake? | 0.00 | The real test. Demand this one. |
| Cohen's kappa | How much better than lucky guessing? | 0.00 | Under 0.60 is generally inadequate 6 |
| Bias on wake time | Does it systematically undercount time awake? | Undercounts by 100% | Near zero, with tight error bounds |
| Who was tested | Healthy 25-year-olds, or people with real sleep problems? | — | Ask. It changes everything. |
Five Questions for a Manufacturer
Send these. The response tells you more than the spec sheet.
- "What's your specificity for wake detection?" If they only have accuracy, they either didn't measure it or didn't like the answer.
- "What's your Cohen's kappa, or PABAK?" The published standardized framework asks for it 7. A company that's done the work will have it.
- "Who was in the validation sample?" Thirty-four healthy young adults 2 is a different claim than 453 people at a clinical sleep lab 3. Ask about age range, and whether anyone in the study had insomnia or apnea.
- "What were you compared against?" In-lab polysomnography with EEG, or another consumer device, or a self-reported sleep diary? Those are wildly different bars.
- "Was it peer reviewed, or is it your own white paper?" Both can be fine. Only one has strangers trying to poke holes in it.
If a company answers all five clearly, that's a good sign about the whole product, not just the stats.
What This Means for Your Own Numbers
The practical takeaway isn't "throw out your tracker." It's narrower and more useful.
Trust the trends, distrust the wake minutes. Almost every consumer device undercounts how long you were awake, because that's the failure mode the math rewards. If your tracker says you were awake 12 minutes and you remember lying there for an hour, your memory is probably closer than the device.
Trust nothing from a single night. These are noisy instruments measuring a noisy system. One number, one night, is not a finding.
And if a device's headline metric is a single percentage with no error bars, that's the tell. This is part of why we've argued against compressing your whole night into one score — the compression is where the honesty leaks out.
Where Komori Fits
We're not going to put an accuracy percentage on the box. Not because we're modest — because per the argument above, the number would be marketing, not information.
Komori is being built as a contactless device with no camera and nothing worn on your body. When it ships, it's designed to log sleep position, movement and restlessness, bed-exit events, and room conditions like temperature, humidity, CO2, light, and noise. It's a general wellness device. It doesn't diagnose anything, and it isn't a sleep study.
What it's designed to give you is many nights instead of one, so that patterns rather than single-night numbers do the talking. A shift in your typical position mix over six weeks is a real observation. A number from last Tuesday is not.
The Bottom Line
A percentage on a box, with no specificity and no kappa behind it, is a decoration.
The three questions worth asking about any sleep tracker are: how well does it find wake, how much better is it than guessing, and who did you test it on? Any company that can't answer those in one email is telling you something.
Ask for the boring numbers. The boring numbers are the ones with information in them.
Footnotes
-
Marino M, et al. "Measuring Sleep: Accuracy, Sensitivity, and Specificity of Wrist Actigraphy Compared to Polysomnography." Sleep, 2013. ↩
-
Chinoy ED, et al. "Performance of seven consumer sleep-tracking devices compared with polysomnography." Sleep, 2021. ↩ ↩2
-
Montazeri N, et al. "A robust generalizable device-agnostic deep learning model for sleep-wake determination from triaxial wrist accelerometry." arXiv preprint — not peer reviewed, 2025. ↩ ↩2
-
Zheng S, et al. "Rethinking PPG-based Sleep Staging: Datasets, Metrics, and Benchmarks." arXiv preprint — not peer reviewed, 2026. ↩
-
Rosenberg RS, Van Hout S. "The American Academy of Sleep Medicine inter-scorer reliability program: sleep stage scoring." Journal of Clinical Sleep Medicine, 2013. ↩
-
McHugh ML. "Interrater reliability: the kappa statistic." Biochemia Medica, 2012. ↩
-
Menghini L, et al. "A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code." Sleep, 2021. ↩
You can choose a position at lights-out. Knowing what you held until morning is the hard part.
Komori is a contactless monitor that logs which position you slept in, through blankets, with no camera and nothing to wear. Pre-launch — join the list and we'll tell you when it ships.
Keep reading
Want to see your sleep position data?
Get the Insider Pass and be first to experience Komori when it ships.


