Leveraging Targeted Gene Sets and Neural Networks for Zebrafish Transcriptome Extrapolation in High-Throughput Toxicogenomics
Background: Zebrafish (Danio rerio) are a powerful vertebrate model for developmental toxicology and chemical safety assessment, yet large-scale transcriptomics in zebrafish remains limited by cost and data heterogeneity. Targeted transcriptomics offers a cost-effective alternative, but gene extrapolation methods tailored to zebrafish have not been systematically developed or evaluated.…
Background: Zebrafish (Danio rerio) serve as a valuable vertebrate model for studying developmental toxicology and chemical safety, but large-scale transcriptomics in zebrafish faces challenges due to high costs and data inconsistencies. Targeted transcriptomics provides a more affordable alternative, yet methods for extrapolating zebrafish gene expression data have not been thoroughly developed or examined.
The S1500+ platform, commonly employed for toxicogenomics research using model organisms like rat, mouse, and human cell lines, has been constrained in zebrafish studies due to limited data availability and insufficient bioinformatics tools for analyzing such data.
To address this issue, researchers aimed to (1) create a comprehensive zebrafish transcriptomic training dataset and (2) assess various machine learning techniques for predicting unmeasured transcriptome-wide expression profiles from zebrafish-specific reduced representation gene set data (Zf S1500+). The team compiled 14,924 zebrafish RNA-Seq samples encompassing 21,930 genes from 1,246 studies.
Using the Zf S1500+ gene subset (3,062 genes), three extrapolation methods - principal components regression (PCR), a locally weighted extension of PCR (PCR+), and a neural network mixture-of-experts model (NN-MoE) - were trained and evaluated. Model performance was measured using mean absolute error (MAE), mean squared regression error (MSRE), and weighted versions of these metrics.
The findings revealed that extrapolation accuracy was greatly affected by tissue and developmental context, with models performing better within a specific tissue or between developmentally related tissues. Errors were minimal when training and testing data were collected within the same tissue or between related developmental stages.
Both PCR+ and NN-MoE outperformed the baseline PCR method, with NN-MoE reducing average MAE by approximately 20% and MSRE by around 17%. Crucially, the extrapolation remained consistent for the vast majority of genes, even when limiting predictions to high-confidence results based on an empirical MAE threshold.
In conclusion, the study demonstrates that targeted transcriptomics can be effectively extended to zebrafish, facilitating robust transcriptome-wide extrapolation while reducing expenses. The NN-MoE approach yielded the most significant improvements, underscoring the benefits of non-linear and ensemble modeling in handling heterogeneous datasets.
These results establish a scalable framework for zebrafish toxicogenomics, suggesting that accuracy will continue to improve with larger, better-annotated datasets, thereby facilitating wider use in chemical safety evaluations.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.