A discrete protein subset drives structure prediction discordance in orphan proteins
Structure and disorder predictors are increasingly used as decision-grade tools in protein engineering and in the analysis of newly emerged proteins, yet how the current state-of-the-art behaves on sequences outside the well-charted evolutionary space remains poorly characterised. We previously reported that AlphaFold2 confidence and the disorder predictor flDPnn produced discordant predictions…
A specific subset of proteins, identified by a class-specific, score-defined driver, is responsible for discordance in structure prediction predictions made by currently available tools, according to a recent study. This driver accounts for 24.5% of de novo proteins, 29.4% of random sequences, 5.1% of conserved proteins, and 1.3% of disordered proteins. When this subset is removed, the correlations between the predictions become normalized.
The discordant predictions were observed between AlphaFold2 and flDPnn, as well as between AlphaFold3 and PUNCH2, when applied to naturally evolved de novo proteins, shuffled sequences, and conserved Drosophila proteins. Despite the discordance, the predictions for these sequences are based on localised, compositionally identifiable phenotypes that current predictors handle in a non-canonical way.
The driver subset was identified through a held-out classifier trained on architectural and compositional features not previously used in the driver definition. This subset is characterized by high pLDDT, high disorder, and low strand fraction, and is predicted by factors such as helix and coil fraction, sequence length, entropy, and hydropathy.
This phenomenon highlights a potential failure mode for protein designers and researchers working with remote sequence space proteins to consider when relying on the outputs of predictor tools.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.