{
  "id": 12225539,
  "title": "Explainable t-SNE for Single-Cell RNA-seq Data Analysis: A Cell-DrivenInformation Fusion Framework",
  "url": "https://urgent.news/2026/10/05/explainable-t-sne-for-single-cell-rna-seq-data-analysis-a-cell",
  "topic": "science",
  "section": "Science",
  "published": "2026-10-05T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.29.755455v1?rss=1"
  },
  "original_language": "en",
  "account": "Single-cell RNA sequencing depends on determining cell-to-cell similarities to identify unique cell states. Conventional dimensional-reduction techniques like t-SNE and UMAP use Euclidean geometry, which is not ideal for high-dimensional, sparse count matrices and results in clusters with little biological meaning. The authors introduce cell-driven information fusion, a framework that creates cell affinity directly from raw counts by combining three biological signals. These signals include gene detection patterns, expression magnitude of main transcripts, and overall profile correlation. A sparsity-adaptive operator fuses these signals into one affinity matrix, while an entry-usage diagnostic identifies which biological component influences each local cluster boundary. The framework has two versions: c-TSNE for datasets with low-to-moderate sparsity (less than 80 percent) and c-UMAP for high-sparsity situations. c-UMAP shows the best results across many datasets and protocols, outperforming Euclidean baselines. With all improvements staying significant even after adjusting for multiple comparisons, these benefits come from the algorithm's inherent strength, not specific dataset tuning. Moreover, c-UMAP can handle up to 1.14 million cells and runs up to 22 times faster than the best current method. Importantly, the method works directly with raw counts, preventing the loss of top differentially expressed genes that typically happens when using variance-based gene selection. This preserves critical lineage markers like IAPP and FOXP3. These findings show that having efficient atlas-scale performance, maintaining raw data, and providing clear biological insights can all coexist.",
  "summary": "Single-cell RNA sequencing fundamentally depends on defining cell-to-cell similarity to resolve distinct cell states. However, conventional dimensional-reduction frameworks such as t-SNE and UMAP inherently rely on Euclidean geometry, a construct ill-suited for high-dimensional, sparse count matrices, and produce cluster boundaries that lack biological interpretability. Here, we introduce…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}