{
  "id": 2741645,
  "title": "On the robustness of scRNA-seq foundation models for plant perturbation response prediction under cross-experiment shift",
  "url": "https://urgent.news/2026/08/22/on-the-robustness-of-scrna-seq-foundation-models-for-plant",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-22T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.21.746324v1?rss=1"
  },
  "original_language": "en",
  "account": "Foundation models for analyzing single-cell transcriptomics data have shown potential in learning generalizable representations of cellular states. However, recent studies have indicated that these models often underperform compared to basic machine learning methods. Even more concerning is the limited understanding of how well these models can generalize to new experimental conditions, especially in plants where comprehensive evaluations beyond cell type annotation and batch integration are scarce. To tackle this issue, researchers have developed an Arabidopsis thaliana foundation model called scAraFM and tested it on various perturbation conditions using three progressively challenging protocols: random splits from a single experiment, replicate-based splits, and cross-experiment transfer learning. The findings reveal that using random splits tends to overestimate the model's performance by up to 30 points when compared to cross-experiment evaluations. Additionally, preserving gene identity appears to be more effective than using standard pooled embeddings across different representation strategies. Even more striking is the observation that simple baselines utilizing raw reads can perform competitively in single-experiment scenarios, casting doubt on the notion that foundation models inherently possess universal advantages. In contrast, when cross-experiment transfer learning is employed, pretrained representations demonstrate added value, particularly when dealing with limited labeled samples. This suggests that the real benefits of foundation models become apparent in the very situations where practical deployment is crucial. In summary, the study underscores the significant impact of evaluation design on conclusions drawn about foundation models. Furthermore, maintaining the per-gene structure proves beneficial for generalization in downstream tasks, providing evidence for robust predictions across uncharted experimental contexts.",
  "summary": "Foundation models for single-cell transcriptomics promise to learn generalizable representations of cellular states. However, recent evidence suggests they often fail to outperform simple machine learning baselines. Furthermore, their ability to generalize across unseen experimental conditions remains poorly understood, particularly in plants, where rigorous evaluation beyond cell type annotation…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}