A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction
Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address…
Enzyme Commission (EC) prediction is a complex task due to the hierarchical organization of enzyme functions, imbalanced class annotations, and the fact that sequence similarity alone may not always be enough to differentiate functional differences. To tackle these issues, researchers created ESM-ECForest, a two-stage machine learning pipeline that integrates protein embeddings from the pretrained language model ESM-2 with Random Forest classifiers.
The first stage of the pipeline distinguishes enzymes from non-enzymatic proteins, while the second stage assigns one or more EC numbers to the enzymes identified. In an external benchmark consisting of 25,778 protein sequences, ESM-ECForest demonstrated the highest weighted F1 score across all four EC levels, ranging from 0.94 at Level 1 to 0.90 at Level 4.
The improvements were particularly notable for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 were still the most challenging classes to classify. By examining the ESM-2 embedding space through UMAP visualization, the researchers observed clustering patterns that aligned with enzyme functional relationships, suggesting that the pretrained representations preserved biologically relevant information before the supervised classification process.
These findings indicate that leveraging pretrained protein language model embeddings can serve as an effective foundation for enzyme annotation. The ESM-ECForest approach, which combines large-scale sequence representations with a lightweight supervised classifier, offers a scalable solution for EC prediction and may aid in the functional annotation of protein sequences obtained from extensive genomic and metagenomic datasets.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.