Quantifying the Provenance-to-Function Gap in Antidiabetic Peptide Prediction: Homology-Aware Evaluation and the ADP-Hybrid Baseline
Antidiabetic peptides (ADPs) are short bioactive sequences of therapeutic interest, and sequence-based classifiers prioritise experimental candidates. A classifier is useful only if its accuracy transfers to unseen sequences, which depends on how the benchmark was as sembled and partitioned. We term the distance between what such a classifier is scored on and the function it is meant to predict…
The provenance-to-function gap quantifies the disparity between a classifier's performance on a benchmark and its predictive power for actual antidiabetic peptides. This gap is particularly relevant for antidiabetic peptide (ADP) prediction, as classifier accuracy hinges on how the benchmark was assembled and divided. To address this issue, the authors re-evaluated the two-layer ADP benchmark of Basith et al. using a protocol that grouped homologous peptides instead of splitting them at random.
In this re-evaluation, 218 of the 877 ADPs traced back to a precursor protein through exact substring containment. Within this subset, the second-layer label matched the precursor identity exactly, including all 140 human-insulin fragments labeled as type-1 and all 76 bovine milk-protein fragments labeled as type-2 without exception.
Accuracy in ADP prediction was found to correlate with identity to the training set. As identity increased from below 50% to between 70% and 90%, the Matthews correlation coefficient (MCC) rose from 0.39-0.43 to 0.92-0.96, respectively. Furthermore, peptide length alone achieved an MCC of 0.619 on held-out data.
The authors also audited sixteen additional peptide benchmarks, finding that the coupling of accuracy to identity is not limited to the Basith et al. resource. In eleven out of fourteen testable datasets, fragment families shared a label more often than would be expected by chance, while in three datasets, this was not the case. This indicates that the property of accuracy tracking identity is common rather than universal.
When the published architecture of the ADP classifier was rebuilt on identical folds, its advantage over a single classifier became apparent, primarily due to the partitioning method. The improved performance was observed under random partitioning but absent once homologous peptides were separated.
To address the provenance-to-function gap in ADP prediction, the authors introduced ADP-Hybrid, a classifier consisting of one tree-ensemble classifier per layer. ADP-Hybrid demonstrated superior performance with an MCC of 0.843 on layer 1, surpassing the published MCC of 0.841 by a factor of 21-38 times the inference throughput.
On layer 2, the MCC was 0.801, outperforming the published MCC of 0.858. The authors have released the protocol and controls that expose these properties, enabling further research into quantifying the provenance-to-function gap in antidiabetic peptide prediction.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.