Position-wise Fold-Switch Prediction in Metamorphic Proteins: A Comparison of Handcrafted and Language Model Features
In this work, we focus on predicting fold-switching phenomena observed in certain protein structures. Proteins that undergo fold-switching are called metamorphic proteins. We model fold-switch prediction in proteins as a position/residue-wise classification task. We use various features extracted from a protein's sequence as well as structure to train the classifiers. We propose a novel class of…
This research paper investigates the prediction of fold-switching, a phenomenon observed in specific protein structures. Metamorphic proteins are those that undergo such fold-switching. The authors model fold-switch prediction as a position-wise classification task, utilizing various features extracted from both the protein's sequence and structure.
They propose a new class of handcrafted parametrized features that capture the structural context of a residue and also employ embeddings from existing protein language models (PLMs) to train classifiers. The study trains both binary and 3-class classifiers, with the latter providing more detailed information on the type of fold-switch.
The paper finds that binary prediction generally outperforms 3-class prediction. Handcrafted features perform comparably to PLM embeddings when tested on a standard dataset. An extensive error analysis identifies a significant data-leakage vulnerability in the metamorphic protein test data. Additionally, the analysis shows that the PLM-based classifier is prone to sequence-memorization bias and struggles to generalize to new fold-switching sequences.
Conversely, one of the domain-informed handcrafted feature-based classifiers can detect fold-switching without relying on sequence memory.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.