Signature Recontextualization: Mapping perturbational signatures across biological contexts
Perturbational transcriptomics is a powerful tool for understanding gene function and drug effects, yet predicting how perturbations manifest across different biological contexts remains a central challenge, limiting translation from model systems to clinically relevant tissues. Despite growing interest in this problem, benchmarking efforts have been hindered by inconsistent evaluation tasks,…
Perturbational transcriptomics offers valuable insights into gene function and drug effects; however, applying these findings across various biological contexts is a significant hurdle, impeding the transfer of model system results to clinically relevant tissues. This issue has hampered benchmarking efforts due to inconsistent evaluation tasks, diverse metrics, and limited assessments across different perturbations and biological systems.
To address this, we present a benchmarking framework for cross-context perturbation-signature prediction, also known as signature recontextualization. This framework is built on precise definitions of the prediction task, target-data accessibility, and evaluation metrics centered around signature recovery. By evaluating prediction performance across three target-context data scenarios—control only, low coverage, and high coverage—we can systematically understand how prediction performance is influenced by sample size while establishing a uniform basis for method comparison.
Our evaluation encompasses state-of-the-art projection-based (projectCor) and network-based (netProp) methods, alongside deep learning-based foundation models (scGPT, STACK) and statistical baselines. The benchmark is extensive, encompassing four distinct perturbational datasets: CRISPR knockdowns and drug perturbations in cell lines, as well as in vivo chemical perturbations in rat tissues from DrugMatrix.
This broad scope extends evaluation beyond traditional cell-line models to tissue-level responses. Our findings reveal that projection and network propagation approaches exhibit remarkable adaptability across perturbation types and biological contexts, and in several instances, match or surpass the performance of deep learning and foundation models.
This suggests that model complexity may not inherently enhance cross-context generalization. Furthermore, we demonstrate that the predictability of perturbations varies significantly with pathway conservation, transcriptional response strength, and the baseline similarity between source and target contexts. All datasets, methods, and evaluation tools are made available as an open-source R package (sigRecon), providing a robust foundation for reproducible benchmarking and future method development.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.