{
  "id": 4060767,
  "title": "OmicsFM brings proteomics into the foundation model era",
  "url": "https://urgent.news/2026/08/28/omicsfm-brings-proteomics-into-the-foundation-model-era",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-28T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.25.747021v1?rss=1"
  },
  "original_language": "en",
  "account": "Foundation models have demonstrated their ability to learn biological representations from large transcriptomic atlases. However, the potential of proteomics data in this regard remained unexplored until now. To address this gap, researchers have introduced OmicsFM – a modality-agnostic transformer pretrained through masked abundance reconstruction on an extensive proteomics data corpus of 48,837 quality-filtered profiles from 1,397 reprocessed PRIDE projects.\n\nRemarkably, despite training on 14-to-93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, OmicsFM matches their performance. On held-out projects, the attention networks of OmicsFM recover more molecular relationships compared to co-expression methods and existing single-cell foundation models across nine reference databases. These databases reveal pathway-level organization, showcasing OmicsFM's prowess.\n\nFurthermore, sample-level embeddings preserved biological structure across independent studies, indicating the robustness of OmicsFM's representations. The model's success extended beyond its primary task, transferring effectively to cell-type classification, gene-essentiality prediction, and perturbation-response prediction. Notably, OmicsFM consistently outperformed task-specific models across these diverse applications.\n\nCritical insights emerged from the findings: proteomics and transcriptomics representations capture complementary aspects of biology. This discovery underscores the importance of training highly performant proteomics-based foundation models, which hold great promise in modeling and uncovering fundamental biological processes.",
  "summary": "While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}