Urgent.News

What's breaking now, across thousands of outlets.

Science

MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data

Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results:…

Legacy human microarray data offers a vast record of transcriptomic biology, yet integrating such datasets for large-scale analysis presents challenges due to platform design differences, preprocessing variations, measurement scales, and gene coverage disparities. To address this issue, researchers have developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource that harmonizes and completes gene-expression profiles across heterogeneous microarray platforms.

At the core of MPGEM is the Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework. These tools enable the transformation of gene-expression profiles with varying gene coverage onto a shared quantitative scale, thereby facilitating their integration. The MPGEM Engine, a multilayer perceptron, predicts the expression of unmeasured genes from shared genes across platforms.

The MPGEM framework was tested on Affymetrix GPL570, GPL571, and GPL96 platforms. Using GPL570 as a 19,320-gene reference space, MPGEM generated 207,135 human gene-expression profiles. Evaluation of the resource using masked GPL570 profiles yielded high correlations, with a mean sample-wise Pearson correlation of 0.944 and Spearman correlation of 0.939, and a mean gene-wise correlation of 0.830. Even the lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683.

Compared to baseline mean imputation and K-nearest-neighbor methods, MPGEM demonstrated comparable or superior predictive performance. The MPGEM framework, along with the trained models and expression resource, is now made available as open-source resources, allowing for widespread utilization in large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning applications.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Tuesday 25 August →