{
  "id": 2728240,
  "title": "Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval",
  "url": "https://urgent.news/2026/08/22/sequence-derived-representations-versus-pfam-domain-content-for",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-22T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.21.746127v1?rss=1"
  },
  "original_language": "en",
  "account": "Retrieving biosynthetic gene clusters (BGCs) akin to those of a known producer is akin to a representation-learning endeavor. We speculate that sequence-derived representations, specifically those generated by ESM-2, could surpass Pfam-domain content metrics in this context. Our toolkit encompasses group-disjoint train, validation, and test assignments, validation-frozen model selection, five seeds, and family-level paired inference. We evaluated 6,953 atlas BGCs from 182 Streptomyces griseus genome accessions, with 5,325 silver-labeled BGCs partitioned into 98 training, 21 validation, and 21 test reference groups. Of these test reference groups, 16 were suitable for retrieval diagnostics. The Pfam Jaccard metric yielded a Recall@50 score of 0.8788, whereas the Pfam-augmented BGC-SetNet achieved 0.8472. When we combined ESM and Pfam-augmented BGC-SetNet, the score improved to 0.8769. A weighted Pfam Jaccard metric marginally surpassed the unweighted Pfam accard, resulting in a score of 0.8789, a difference that is insignificant in practical terms. Our findings do not corroborate the notion that sequence-derived representations can uncover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the paramount signal for this objective. Our results elucidate the curation and pathway-level validation procedures required for a more robust biological test.",
  "summary": "Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}