REFinder: Mining new enzymes from metagenomes
Microbial metagenomes encode vast catalytic diversity, but recovering enzymes of novel, yet-undescribed functionality typically requires whole-community assembly plus a means of identifying active proteins ab initio. The first step of the task is compute-intensive and misses low-abundance sequences. The second is complicated by our inability to predict novel functions. REBEAN, our DNA language…
Microbial metagenomes contain a vast array of catalytic diversity, but discovering enzymes with novel, previously unknown functionality typically demands full community assembly and a method to identify active proteins from scratch. The initial challenge involves computationally intensive processing, which often overlooks low-abundance sequences.
Additionally, predicting novel functions remains complicated by our inability to accurately predict them. REBEAN, a DNA language model, addresses the latter issue by assigning each sequencing read an Enzyme Commission (EC) class or a non-enzyme label, without the need for alignment.
To tackle both challenges, we developed REFinder, a pipeline that employs REBEAN-annotated reads. These reads are then used to assemble only those that display catalytic signatures, effectively routing them to the next stage. In our evaluation of 50 microbiome metagenomes, REFinder identified 1.2 to 2.1 times more enzymes compared to homology-based annotation of proteins derived from full assemblies.
Furthermore, REFinder proved to be up to 6.4 times more cost-effective computationally than conducting a full assembly. Across all samples, REFinder captured at least three-quarters and up to 90% of the homology-accessible enzymes discovered through full assembly of the complete metagenomes. Remarkably, a fraction of these enzymes, ranging from 22% to 45% per EC class, bore no resemblance to Swiss-Prot proteins, which signifies enzymes that homology annotation cannot identify.
For nearly two-fifths of the over thirteen thousand novel oxidoreductases obtained from two saliva samples, ESMFold successfully predicted structures with an alignment score (TM-score) of 0.7 or higher to a characterized enzyme structure in the Protein Data Bank (PDB). This substantial structural similarity was achieved without relying on sequence homology.
These findings demonstrate that targeted, alignment-free assembly can transform even well-mapped microbiomes into a reservoir of thousands of previously invisible but credible novel enzymes, ultimately expanding our expectations for environmental microbiome exploration.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.