Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients
Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to…
Acute myocardial infarction (AMI) is a leading cause of death globally, and peripheral blood mononuclear cells play a crucial role in post-AMI inflammation and tissue repair. Single-cell RNA sequencing (scRNAseq) data is often affected by dropout events, yielding sparse and noisy datasets that can negatively impact downstream results.
To evaluate the impact of missing data, researchers artificially introduced varying levels of dropout (10%, 20%, and 30%) into scRNAseq data using the missing completely at random (MCAR) framework, replicating the experiment 10 times. Six imputation strategies were then benchmarked: MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN).
The evaluation metrics included marker gene preservation, clustering consistency (ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The findings reveal that no single imputation method consistently outperformed the others across all metrics. Mean and KNN imputers displayed limited recovery in transcriptional data.
Conversely, GAN showed superior global transcriptional recovery, while SoftImpute excelled in maintaining biologically relevant cell-type signals. The study emphasizes the significance of selecting appropriate imputation methods during the pre-processing stage to address transcriptome recovery, marker gene detection, and cell-type-specific resolution in scRNAseq data analysis.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.