Urgent.News

What's breaking now, across thousands of outlets.

Science

scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning

Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learning stable biological representations. This challenge is particularly relevant to…

The article introduces scRep, a novel compact latent-space self-distillation framework designed for single-cell representation learning. Unlike existing methods that reconstruct masked gene expression values, scRep focuses on aligning different perturbed views of the same cell using a momentum-updated teacher-student architecture with self-distillation objectives at both the cell and gene levels.

This approach aims to capture stable biological information that remains consistent across incomplete and perturbed transcriptomic observations.

Pretrained on roughly 2.8 million cells, scRep pretrained achieves the best performance among frozen-representation benchmarks, showcasing remarkable sample efficiency. A more extensive scRep model, trained on 30.72 million cells, further validates the framework's effectiveness when scaled up to a larger and more diverse dataset.

Unlike other approaches that emphasize cell identity, scRep prioritizes well-established marker genes and recovers transcription factor-associated gene programs that exhibit cell-type-specific activity. Additionally, it maintains continuous developmental structure, facilitating graph-based pseudotime inference.

The pretraining performance of scRep is closely linked to biological diversity; by optimizing cell-type coverage while reducing redundant cells, the model can match or surpass the performance of larger but less balanced training corpora. In summary, latent-space self-distillation emerges as a promising alternative to expression reconstruction for single-cell foundation modeling, highlighting that efficient scaling relies not only on the quantity of cells but also on the learning objective and the biological diversity of the pretraining corpus.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Thursday 3 September →