{
  "id": 5477707,
  "title": "scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning",
  "url": "https://urgent.news/2026/09/03/screp-a-latent-space-self-distilled-foundation-model-for-single-cell",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-03T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.31.747784v1?rss=1"
  },
  "original_language": "en",
  "account": "The article introduces scRep, a novel compact latent-space self-distillation framework designed for single-cell representation learning. Unlike existing methods that reconstruct masked gene expression values, scRep focuses on aligning different perturbed views of the same cell using a momentum-updated teacher-student architecture with self-distillation objectives at both the cell and gene levels. This approach aims to capture stable biological information that remains consistent across incomplete and perturbed transcriptomic observations.\n\nPretrained on roughly 2.8 million cells, scRep pretrained achieves the best performance among frozen-representation benchmarks, showcasing remarkable sample efficiency. A more extensive scRep model, trained on 30.72 million cells, further validates the framework's effectiveness when scaled up to a larger and more diverse dataset. Unlike other approaches that emphasize cell identity, scRep prioritizes well-established marker genes and recovers transcription factor-associated gene programs that exhibit cell-type-specific activity. Additionally, it maintains continuous developmental structure, facilitating graph-based pseudotime inference.\n\nThe pretraining performance of scRep is closely linked to biological diversity; by optimizing cell-type coverage while reducing redundant cells, the model can match or surpass the performance of larger but less balanced training corpora. In summary, latent-space self-distillation emerges as a promising alternative to expression reconstruction for single-cell foundation modeling, highlighting that efficient scaling relies not only on the quantity of cells but also on the learning objective and the biological diversity of the pretraining corpus.",
  "summary": "Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learning stable biological representations. This challenge is particularly relevant to…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}