Every wearable now reports a recovery or readiness score. Almost none of them tell you how it was computed. That is a product decision, and we think it is the wrong one.
The inputs are noisier than the outputs suggest
Heart rate variability is genuinely informative and genuinely noisy. It moves with hydration, alcohol, room temperature, measurement window, body position and where you are in a respiratory cycle. Day-to-day variation in a healthy person routinely exceeds the difference between a 'good' and 'poor' score on most consumer scales.
Sleep staging is an estimate derived from movement and heart rate, not a measurement. Consumer devices agree with polysomnography reasonably well on total sleep time and poorly on stage boundaries. A product that reports forty-two minutes of deep sleep to the minute is communicating a precision it does not have.
What we did instead
- Baseline per person, not per population. A score is a deviation from your own rolling baseline, which takes about three weeks to establish — and Pulsr says so rather than showing a confident number on day two.
- Trends over points. The seven-day direction is displayed at least as prominently as today's value, because that is where the signal actually lives.
- Show the contribution. Tapping the score reveals which inputs moved it and by how much, so a user can tell 'you drank last night' from 'you trained hard on Tuesday'.
- Admit uncertainty. When inputs are missing or contradictory, the app widens the range instead of guessing.
A number without its reasoning is not insight. It is an instruction the user has no way to evaluate.
The evaluation harness mattered more than the model
The engineering result we did not anticipate: the most valuable artefact was not the model but the harness around it. Held-out cohorts, replay of historical data, drift monitoring on every input stream, and an alert when the distribution of scores shifts without a corresponding shift in inputs.
That harness is now used by the group's enterprise clients for entirely unrelated time-series inference. A consumer product paid for infrastructure that made the B2B work better — which is the argument for running both.
Working on this?
We run a paid two-week diagnostic that ends with an architecture record, a risk register and a costed plan — yours to keep either way.
Talk to an engineer