{
  "id": 6134073,
  "title": "Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics",
  "url": "https://urgent.news/2026/09/07/learning-from-tandem-mass-spectra-at-scale-with-a-self-supervised",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-07T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.03.747733v1?rss=1"
  },
  "original_language": "en",
  "account": "Mass spectrometry-based proteomics heavily depends on machine learning techniques; however, current models are trained for specific supervised tasks like peptide identification, de novo sequencing, or fragment intensity prediction, restricting their application across datasets, instruments, and acquisition methods. To address this issue, researchers have introduced InstaNovo-FM, a self-supervised foundation model designed for bottom-up proteomics. This model is trained to reconstruct masked regions of tandem mass spectra.\n\nThe team behind InstaNovo-FM has compiled a diverse training corpus, comprising 1.47 billion MS/MS spectra and 184.6 million high-confidence annotations. These vast amounts of data enable the model to learn about various aspects of experimental and biological properties, including the type of fragmentation method utilized, sequence characteristics, and post-translational modifications.\n\nInterestingly, the InstaNovo-FM embeddings manage to encode these fundamental properties without requiring any peptide labels. This self-supervised learning approach allows the model to capture essential information about peptide fragmentation spectra, establishing a unified representation space for the entire proteomics ecosystem. As a result, InstaNovo-FM directly empowers various downstream applications, such as de novo peptide sequencing, database-free identification, and analytical run classification. By providing a shared representation for mass spectrometry data, the self-supervised foundation model significantly enhances the transferability and interoperability of proteomics research across different platforms and methodologies.",
  "summary": "Mass spectrometry-based proteomics increasingly relies on machine learning, yet existing models are trained for defined supervised tasks such as peptide identification, de novo sequencing or fragment intensity prediction, limiting transfer across datasets, instruments and acquisition methods. Here we present InstaNovo-FM, a self-supervised foundation model for bottom-up proteomics trained to…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}