{
  "id": 3021165,
  "title": "A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models",
  "url": "https://urgent.news/2026/08/24/a-mechanism-annotated-benchmark-reveals-limited-fidelity-to-drug",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-24T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.19.745729v1?rss=1"
  },
  "original_language": "en",
  "account": "A novel benchmark named scDrugPerturb-Bench has been introduced to assess the fidelity of single-cell drug perturbation models in predicting drug responses. This benchmark links matched control and drug-treated RNA sequencing profiles to curated evidence on gene expression changes, encompassing 181 datasets, 423 response data points, 717 key genes and 2.5 million individual cells. To evaluate the models, the authors introduce the Mechanism Fidelity Score (MFS), which assesses key-gene direction, effect-size recovery, gene-set coherence, mechanism specificity and pathway-level response polarity.\n\nWhen applied to 12 perturbation-prediction models, three baselines and ten data splits, the findings indicate that traditional expression-similarity metrics have only weak alignment with the MFS score, often leading to the selection of different model configurations. However, incorporating mechanism-awareness in model selection proved beneficial in early drug retrieval during a transcriptome-based drug design evaluation, demonstrating that the MFS offers practical insights beyond mere benchmark reporting.\n\nThe systematic benchmarking of these models across various cell lines and data integration settings revealed a limited fidelity to accurately capturing drug-response signatures. Notably, embedding single-cell foundation models yielded local, metric-dependent improvements rather than universal benefits, and the contextual information provided by the source material significantly influenced model assessment. Furthermore, hard-negative tests underscored the potential for plausible perturbation responses to arise from non-specific transcriptional shortcuts.\n\nThese results collectively underscore the insufficiency of expression reconstruction as a reliable proxy for preserving drug-response signatures when evaluating single-cell drug perturbation models. The introduction of scDrugPerturb-Bench as a benchmark for mechanism-aware model evaluation marks a significant step forward in this domain, providing researchers with a more robust framework for assessing the predictive accuracy and fidelity of their models.",
  "summary": "Single-cell drug perturbation models are increasingly used to predict how compounds remodel cellular states, but they are still largely assessed by expression reconstruction. Whether high expression similarity reflects preservation of drug-response signatures remains unclear. Here we present scDrugPerturb-Bench, a mechanism-annotated benchmark that links matched control and drug-treated…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}