Detecting CYP2C19 deletions from genotyping array signals using neural networks
Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still…
The study aimed to improve the detection of copy number variations (CNVs) in pharmacogenes, which can significantly alter drug metabolism. Whole-genome sequencing, particularly long-read sequencing, is currently the gold standard for CNV detection. However, genotyping arrays remain popular due to their cost-effectiveness in biobank and clinical settings. The challenge lies in the low base pair resolution of array intensity signals when calling CNVs.
The researchers developed a neural network model named nnCNV to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. This model was compared to the widely-used PennCNV algorithm. The results showed that nnCNV achieved 100% accuracy in the test dataset, demonstrating superior performance.
To validate the predictions, the researchers used an identity-by-descent (IBD) sharing method, which also favored nnCNV. In cases where both methods disagreed, PCR analysis was performed, revealing 97% precision for nnCNV predictions compared to 23% for PennCNV.
The researchers further analyzed the gradient-based feature importance maps generated by nnCNV. They discovered that the model utilized signal intensity information not only from deletion probes but also from neighboring probes in flanking regions. This ability to leverage long-range information, which cannot be accessed by hidden Markov models, significantly improved CNV calling.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.