Benchmarking long-read RNA sequencing for de novo transcriptome assembly in non-model plant species: insights from Moricandia arvensis
De novo transcriptome assembly is the standard approach for constructing a reference transcriptome in non-model plants that lack a high-quality genome, yet short-read assemblies struggle to resolve full-length isoforms. Long-read Iso-Seq (PacBio) captures full-length transcripts directly, but its use as a primary reference and the choice of downstream assembly pipeline remains poorly benchmarked.…
Constructing genome-free transcriptomes for non-model plants like Moricandia arvensis can be challenging due to the limitations of short-read assemblies. Long-read Iso-Seq (PacBio) technology can capture complete transcripts, but its effectiveness as a reference assembly and the appropriate downstream pipeline are not yet well-established.
In this study, researchers generated genome-free transcriptomes from two organs of Moricandia arvensis - flower and leaf - and compared various pipeline strategies, including Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction.
The optimal pipeline strategy varied depending on the organ being analyzed. For the flower, CD-HIT pre-filtering followed by Cogent reconstruction resulted in a high-quality reference transcriptome with 95.3% BUSCO completeness. However, for the leaf, Cogent reconstruction proved detrimental, reducing BUSCO completeness from 90.1% to 78.0% by incorrectly merging distinct genes. In this case, CD-HIT at a 95% identity threshold without reconstruction was retained.
The difference in performance between organ types can be attributed to the unique characteristics of the input data. Leaf transcripts exhibit extreme full-length-read expression skew and predominantly single-isoform gene support, which lacks the multi-isoform evidence that the Cogent graph algorithm requires to function correctly.
Interestingly, the concentration of full-length reads from the most highly expressed transcripts serves as a predictive indicator of pipeline suitability before reconstruction. Furthermore, per-transcript read depth acts as a necessary but non-discriminating lower bound.
As the leaf reference transcriptome lacked gene-level structure, the researchers further improved their results by utilizing expression-aware read-clustering (Corset) to recover gene-isoform grouping. This method preserved completeness while restoring the paralog structure expected of a paleopolyploid genome, outperforming sequence-only clustering techniques.
Overall, the study provides a robust, genome-free framework for constructing full-length reference transcriptomes in non-model plant species, highlighting the importance of evaluating the appropriate pipeline strategy per organ rather than assuming a one-size-fits-all approach.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.