A simple and accurate method for inferring missing ploidy information from sequence data
Polyploidy can be a critical factor for explaining plant trait variation, niche diversification, or speciation. However, inferring ploidy from silica-dried or historical samples using chromosome counts or flow cytometry is not possible, and scaling up ploidy estimation to population-level fresh contemporary samples can be challenging as well. Thus, we present a new method for estimating ploidy…
Polyploidy's influence on plant traits and speciation is significant, yet inferring this ploidy level from samples is difficult due to the unavailability of chromosome counts or flow cytometry on silica-dried or historical specimens. To tackle this challenge, scientists have introduced the Polyploid Population Genomics Tool Kit (PPGTK), a machine learning-based method for directly estimating ploidy levels from sequencing data.
This novel approach offers several advantages, including the relaxation of assumptions present in previous probabilistic methods, per-sample probabilities, and the ability to evaluate uncertainty in the system under study.
To validate the effectiveness of PPGTK, researchers conducted simulations and empirical analyses using target enrichment data from blueberry wild relatives (Vaccinium sect. Cyanococcus) and whole-genome data from sweetpotato wild relatives (Ipomoea ser. Batatas). The simulations demonstrated a remarkable 99% accuracy, even for low-coverage data, provided that the long reads were mappable to the reference genome.
Empirically, the method achieved 99% accuracy for ploidy recovery across 70 Vaccinium individuals and 97% accuracy for 82 Ipomoea individuals.
Implementing PPGTK in collections-based research is efficient and cost-effective as it requires only a multisample VCF, which is typically generated for research purposes, and samples of known ploidy for training the classifier. Moreover, the approach can be applied to historical specimens by classifying their ploidy based on present-day observations. The method is readily accessible through a new Python package that can run on a conventional laptop, making it a promising tool for future research in the field.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.