Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics
Mass spectrometry-based proteomics increasingly relies on machine learning, yet existing models are trained for defined supervised tasks such as peptide identification, de novo sequencing or fragment intensity prediction, limiting transfer across datasets, instruments and acquisition methods. Here we present InstaNovo-FM, a self-supervised foundation model for bottom-up proteomics trained to…
Mass spectrometry-based proteomics heavily depends on machine learning techniques; however, current models are trained for specific supervised tasks like peptide identification, de novo sequencing, or fragment intensity prediction, restricting their application across datasets, instruments, and acquisition methods. To address this issue, researchers have introduced InstaNovo-FM, a self-supervised foundation model designed for bottom-up proteomics. This model is trained to reconstruct masked regions of tandem mass spectra.
The team behind InstaNovo-FM has compiled a diverse training corpus, comprising 1.47 billion MS/MS spectra and 184.6 million high-confidence annotations. These vast amounts of data enable the model to learn about various aspects of experimental and biological properties, including the type of fragmentation method utilized, sequence characteristics, and post-translational modifications.
Interestingly, the InstaNovo-FM embeddings manage to encode these fundamental properties without requiring any peptide labels. This self-supervised learning approach allows the model to capture essential information about peptide fragmentation spectra, establishing a unified representation space for the entire proteomics ecosystem.
As a result, InstaNovo-FM directly empowers various downstream applications, such as de novo peptide sequencing, database-free identification, and analytical run classification. By providing a shared representation for mass spectrometry data, the self-supervised foundation model significantly enhances the transferability and interoperability of proteomics research across different platforms and methodologies.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.