{
  "id": 6533276,
  "title": "Benchmarking long-read RNA sequencing for de novo transcriptome assembly in non-model plant species: insights from Moricandia arvensis",
  "url": "https://urgent.news/2026/09/09/benchmarking-long-read-rna-sequencing-for-de-novo-transcriptome",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-09T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.08.750093v1?rss=1"
  },
  "original_language": "en",
  "account": "Constructing genome-free transcriptomes for non-model plants like Moricandia arvensis can be challenging due to the limitations of short-read assemblies. Long-read Iso-Seq (PacBio) technology can capture complete transcripts, but its effectiveness as a reference assembly and the appropriate downstream pipeline are not yet well-established. In this study, researchers generated genome-free transcriptomes from two organs of Moricandia arvensis - flower and leaf - and compared various pipeline strategies, including Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction.\n\nThe optimal pipeline strategy varied depending on the organ being analyzed. For the flower, CD-HIT pre-filtering followed by Cogent reconstruction resulted in a high-quality reference transcriptome with 95.3% BUSCO completeness. However, for the leaf, Cogent reconstruction proved detrimental, reducing BUSCO completeness from 90.1% to 78.0% by incorrectly merging distinct genes. In this case, CD-HIT at a 95% identity threshold without reconstruction was retained.\n\nThe difference in performance between organ types can be attributed to the unique characteristics of the input data. Leaf transcripts exhibit extreme full-length-read expression skew and predominantly single-isoform gene support, which lacks the multi-isoform evidence that the Cogent graph algorithm requires to function correctly. Interestingly, the concentration of full-length reads from the most highly expressed transcripts serves as a predictive indicator of pipeline suitability before reconstruction. Furthermore, per-transcript read depth acts as a necessary but non-discriminating lower bound.\n\nAs the leaf reference transcriptome lacked gene-level structure, the researchers further improved their results by utilizing expression-aware read-clustering (Corset) to recover gene-isoform grouping. This method preserved completeness while restoring the paralog structure expected of a paleopolyploid genome, outperforming sequence-only clustering techniques. Overall, the study provides a robust, genome-free framework for constructing full-length reference transcriptomes in non-model plant species, highlighting the importance of evaluating the appropriate pipeline strategy per organ rather than assuming a one-size-fits-all approach.",
  "summary": "De novo transcriptome assembly is the standard approach for constructing a reference transcriptome in non-model plants that lack a high-quality genome, yet short-read assemblies struggle to resolve full-length isoforms. Long-read Iso-Seq (PacBio) captures full-length transcripts directly, but its use as a primary reference and the choice of downstream assembly pipeline remains poorly benchmarked.…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "bioRxiv",
        "title": "Long read sequencing of retinal RNA improves killifish transcriptome annotation",
        "url": "https://urgent.news/2026/09/07/long-read-sequencing-of-retinal-rna-improves-killifish-transcriptome",
        "published": "2026-09-07T00:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}