Unlocking Sensitive Data with SPHERE in the Age of AI
Sensitive human data underpin discoveries across medicine, biology and the social sciences, yet privacy regulation often prevents sharing them with collaborators or artificial intelligence (AI) systems. We introduce SPHERE, a model-free method that makes sensitive datasets directly usable by AI and shareable for open science as a synthetic twin, while the original records never leave the local…
The SPHERE model, introduced in our report, enables the utilization of sensitive human data for AI analysis and scientific research while maintaining privacy. This breakthrough allows researchers to share synthetic twins of data with AI systems, ensuring that the original records remain within local environments untouched. Our trial across 33 datasets from medicine, biology, and social sciences demonstrates that SPHERE effectively safeguards individual privacy against adversarial reidentification attacks, all while preserving statistical properties of the data, such as means, variances, and correlations.
Moreover, SPHERE maintains the utility of nonlinear machine learning algorithms and enables researchers to generate synthetic twins in mere seconds on a standard laptop.
When Frontier AI agents run on the SPHERE twin, they achieve the same scientific conclusions as they would with the original data. These findings are not only applicable to genome and proteome-wide data at the scale of the UK Biobank but also extend to recover landmark study results across three independent cohorts and consortia. The algorithm's capabilities aren't limited to specific data types; it applies to deep-learning embeddings in language, vision, and time-series domains with minimal loss in utility.
To demonstrate the practical application of SPHERE, we have made the Stanford Alzheimer's Disease Research Center cohort openly available. This is the first time the dataset, which spans nine modalities, has been made accessible for analysis by any registered researcher without undergoing an approval process. To further ensure the privacy and fidelity of the data, we have provided certification for each SPHERE twin.
Additionally, we have developed an AI agent capable of autonomously executing research tasks on sensitive data without ever accessing it.
Our ultimate goal is to transform sensitive datasets, which are currently inaccessible for research, into routine inputs for open science and AI. This progression will pave the way for significant discoveries across various scientific domains.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.