Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation
Enzyme turnover numbers (kcat) are essential for kinetic models and enzyme-constrained genome-scale metabolic models (ecGEMs), but measured values are sparse and therefore increasingly estimated using machine learning (ML). Although these predictors are commonly evaluated by global regression metrics, their practical utility depends on how errors propagate through downstream models. We…
Turning numbers of enzymes, represented as kcat, are crucial for kinetic models and genome-scale metabolic models that consider enzyme constraints (ecGEMs). These kcat values are often estimated using machine learning (ML) techniques because they are not always readily available. While ML predictors are typically assessed using global regression metrics, their practical value hinges on how errors propagate through subsequent models.
To investigate this, we evaluated six kcat predictors on two datasets: one derived from BRENDA and another from EnzyExtract. Additionally, we compared each benchmark dataset with the training data used for each predictor.
The benchmark results were moderate on the BRENDA dataset and significantly lower on EnzyExtract, where all predictors had R2 values of 0.20 or less. This drop-off coincided with a reduced overlap between the benchmark and training datasets, ranging from 24% to 78% for BRENDA and 9% to 26% for EnzyExtract. However, this overlap did not fully account for the disparities in generalization across predictors.
Furthermore, the performance of downstream ecGEMs was not explained by the benchmark rankings. In the 19 conditions tested, none of the ecGEMs specifically trained with a particular predictor accurately replicated the observed variations in growth.
The discrepancy was traced back to localized errors at high-leverage positions within yeast's metabolic network. Specifically, underpredicted mitochondrial ADP/ATP carrier turnover numbers limited the exchange of adenine nucleotides, creating a bottleneck in cytosolic ATP supply. When this constraint was relaxed, the predicted growth aligned better with the experimental reference.
Consequently, ML-derived kcat values can influence not only growth predictions but also the phenotypes identified by mechanistic models. These findings underscore the necessity of validating biological parameter predictors in the specific systems they are intended to support, rather than solely relying on benchmark accuracy.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.