{
  "id": 2741646,
  "title": "A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models",
  "url": "https://urgent.news/2026/08/22/a-novel-benchmark-dataset-for-enzyme-function-prediction-reveals-the",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-22T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.21.746242v1?rss=1"
  },
  "original_language": "en",
  "account": "A new benchmark dataset, called EnzymARC, for predicting enzyme function has unveiled the shortcomings of current models. Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is crucial for genome annotation and enzyme design. However, researchers have been uncertain if state-of-the-art predictors truly understand the structural factors that enable catalytic activity or merely rely on global sequence similarity to known homologs. To tackle this issue, EnzymARC, a new dataset of putative non-functional decoy sequences, was created by systematically disrupting the active sites of experimentally annotated enzymes. The disruption targeted catalytic residues and surrounding regions of 5, 10, and 15 Angstroms. Three prediction methods were tested against EnzymARC: homology-based annotation using DIAMOND, contrastive learning with protein language models (CLEAN), and a deep learning model that incorporates negative examples (DeepEC). The results showed that current models are highly susceptible to phylogenetic shortcuts. Both DIAMOND and CLEAN had false positive rates over 90% for low-perturbation decoys, confidently assigning original EC numbers to the altered sequences, even when the catalytic machinery was destroyed. DeepEC showed better sensitivity at higher perturbation levels, emphasizing the importance of negative training examples. Overall, modern EC predictors are not effective at distinguishing catalytically incompetent variants from functional enzymes. The study concludes that integrating structure-aware negative examples into both training and benchmarking is essential for developing functionally robust models in computational enzymology.",
  "summary": "Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap,…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}