A single overnight sleep recording may soon carry far more clinical weight than a diagnosis of insomnia or sleep apnea. A foundation model called SleepFM, described in a peer-reviewed Nature Medicine study, was trained on roughly 585,000 hours of polysomnography from approximately 65,000 participants and, according to that paper, can predict future risk for 130 diseases from one night of data. The work draws on some of the largest federally maintained sleep cohorts in the United States and raises a pointed question: can an algorithm trained across decades of recordings and multiple clinical sites hold up when applied to new patients in new hospitals?
Why overnight sleep signals are becoming diagnostic currency
Polysomnography, or PSG, captures brain waves, heart rhythm, breathing effort, oxygen levels, and muscle activity while a person sleeps. Clinicians have long used it to diagnose obstructive sleep apnea and narcolepsy. What makes the SleepFM approach different is scale and ambition. Rather than flagging a single disorder, the model treats the full sensor output of a sleep study as a biological fingerprint and maps it against years of follow-up health records to estimate risk across a wide range of conditions.
The training data came from several distinct sources. The Sleep Heart Health Study, registered as NCT00005275, recruited participants from nine NHLBI cohorts and added in-home polysomnography with follow-up approximately four years after baseline. The Multi-Ethnic Study of Atherosclerosis, registered as NCT00005487, includes the Exam 5 sleep cohort, which contributed PSG data from a racially and ethnically diverse population originally assembled to study cardiovascular disease. A third stream came from Massachusetts General Hospital, where clinical polysomnograms were recorded between 2009 and 2016, as documented in a deep-learning analysis published in the journal SLEEP.
According to Stanford Medicine’s institutional announcement, the largest single training cohort comprised roughly 35,000 clinic patients whose PSGs were recorded between 1999 and 2024 and linked to electronic health records. That figure sits alongside the Nature Medicine paper’s aggregate count of approximately 65,000 participants. The discrepancy has not been publicly reconciled: it is unclear whether the 35,000-patient clinical set overlaps with SHHS or MESA participants or represents an entirely separate pool. Both numbers should be read with that gap in mind.
Calibration drift across sites and decades of sleep data
Training on recordings that span 25 years and at least three institutional sources introduces a specific technical risk: calibration drift. PSG equipment, electrode placement protocols, and scoring conventions have changed over time. The SHHS baseline recordings date to the mid-1990s. The MGH clinical set runs from 2009 to 2016. The broader clinic cohort, per Stanford Medicine, extends to 2024. Sensor hardware, digital sampling rates, and even the demographics of who gets referred for a sleep study shifted across those windows.
That variation matters because a model’s accuracy on one disease category can degrade differently than on another when the input signal subtly changes. Cardiovascular endpoints, for example, depend heavily on heart-rate variability and oxygen desaturation patterns, both of which are sensitive to sensor calibration. Neurological conditions such as epilepsy or neurodegenerative disease leave signatures primarily in EEG channels, which are recorded with different montages in research-grade PSG versus routine clinical EEG. A model pretrained on the combined SHHS, MESA, and MGH data could plausibly show stronger discrimination for neurological endpoints, where the EEG signal is the dominant input, and weaker discrimination for cardiovascular ones, where peripheral sensor drift has more influence. No public release of disease-specific AUC values or confidence intervals from the SleepFM paper has confirmed or refuted that pattern.
Parallel work reinforces the broader direction. A self-supervised model called SleepJEPA, described in an open-access paper, learns from at-home sleep signals and estimates disease risk with time horizons up to 15 years. That study uses simpler wearable-grade data rather than full PSG, which sidesteps some equipment-drift concerns but introduces its own signal-quality tradeoffs. Together, the two projects suggest that the field is converging on long-horizon disease prediction from sleep data, though neither has published the granular performance breakdowns clinicians would need to trust a risk score for any individual condition.
Missing performance data and what patients should watch for
Several gaps stand between a promising research result and a tool a doctor can act on. The Nature Medicine paper reports that 130 diseases achieved measurable predictive performance, but neither the paper’s public summary nor the Stanford announcement provides per-disease AUC values, subgroup analyses by race or ethnicity, or breakdowns by recording site. The MESA cohort was specifically designed to capture cardiovascular risk across a multi-ethnic population, yet there is no public accounting of whether SleepFM performs equally well for Black, Hispanic, Asian, and white participants when trained on that mix of data. Without those numbers, it is impossible to know whether the model amplifies, narrows, or simply shifts existing health disparities.
Patients are unlikely to see “SleepFM score” on their charts in the near term, but they may start hearing about AI-derived sleep risk indices in research settings or specialty clinics. When that happens, there are concrete questions they and their clinicians can ask. One is whether the model has been validated on data from the same type of lab and equipment used for their own study. Another is whether the reported risk estimates have been calibrated for their age group, sex, and racial or ethnic background, or whether a single threshold is being applied across all patients.
Even if future publications fill in the missing performance tables, a second challenge looms: interpretability. A cardiologist can explain why a high coronary calcium score signals elevated heart attack risk. A neurologist can point to specific EEG abnormalities. By contrast, a foundation model trained on millions of unlabeled signal segments may surface patterns that are statistically predictive but physiologically opaque. For some conditions, that may be acceptable if the risk score is used only to prompt further testing. For others, particularly where treatment carries side effects or cost, clinicians will need more than a black-box probability to justify action.
Regulatory and ethical questions on the horizon
As SleepFM-like systems move from research to practice, regulators will have to decide how to classify them. If a single model outputs risk estimates for 130 diseases, is it one medical device or 130? Does each disease-specific prediction require its own validation study and post-market surveillance plan, or can they be bundled under a single approval? Current frameworks for software as a medical device were not built with multi-task foundation models in mind.
Ethically, the prospect of forecasting long-term disease risk from one night of data raises consent and data-governance issues. Many historical PSG recordings were collected for narrow clinical reasons, such as confirming sleep apnea. Patients may not have anticipated that their data would later be used to predict dementia, cancer, or psychiatric disorders. As health systems consider deploying foundation models trained on legacy cohorts, they will need to revisit how they communicate secondary uses and whether new consent processes are warranted.
There is also a risk of over-medicalization. If every sleep study comes bundled with a dashboard of future disease probabilities, clinicians may feel pressure to act on marginal signals, ordering cascades of tests that generate anxiety and cost without clear benefit. Conversely, patients with reassuringly low scores might be falsely reassured and neglect standard preventive care. Careful guidelines will be needed to define when a sleep-derived risk estimate should change management and when it should simply be documented and monitored.
What comes next for sleep-based disease prediction
For now, SleepFM is best understood as a proof of concept that compresses a complex overnight recording into a reusable representation of physiological health. The most immediate impact may not be direct clinical deployment but rather acceleration of research: investigators can apply the model to existing PSG archives to explore links between sleep and disease without training new algorithms from scratch. Over time, that could clarify which endpoints are robustly predictable and which remain too noisy or confounded.
Whether that promise translates into routine care will depend on data that have not yet been shared. Clinicians will want to see per-condition metrics, subgroup fairness analyses, and external validation on hospitals and home devices that were never part of the training pipeline. Patients, in turn, will need clear explanations of what their sleep-derived risk scores mean and how they will be used. Until those pieces are in place, the idea that one night in a sleep lab can map decades of health will remain more of a research milestone than a clinical reality.
More from Morning Overview
*This article was researched with the help of AI, with human editors creating the final content.