Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval
Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired…
Retrieving biosynthetic gene clusters (BGCs) akin to those of a known producer is akin to a representation-learning endeavor. We speculate that sequence-derived representations, specifically those generated by ESM-2, could surpass Pfam-domain content metrics in this context. Our toolkit encompasses group-disjoint train, validation, and test assignments, validation-frozen model selection, five seeds, and family-level paired inference.
We evaluated 6,953 atlas BGCs from 182 Streptomyces griseus genome accessions, with 5,325 silver-labeled BGCs partitioned into 98 training, 21 validation, and 21 test reference groups. Of these test reference groups, 16 were suitable for retrieval diagnostics. The Pfam Jaccard metric yielded a Recall@50 score of 0.8788, whereas the Pfam-augmented BGC-SetNet achieved 0.8472.
When we combined ESM and Pfam-augmented BGC-SetNet, the score improved to 0.8769. A weighted Pfam Jaccard metric marginally surpassed the unweighted Pfam accard, resulting in a score of 0.8789, a difference that is insignificant in practical terms. Our findings do not corroborate the notion that sequence-derived representations can uncover alternative biosynthetic pathways on this benchmark.
Instead, explicit Pfam remains the paramount signal for this objective. Our results elucidate the curation and pathway-level validation procedures required for a more robust biological test.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.