AI model designs potentially stable protein sequences beyond those found in nature
A protein's function is determined by its structure, and structure—the way a protein folds—is determined by its sequence of amino acids, the building blocks of proteins.
Protein function hinges on its structure, with folding into shape determined by its sequence of amino acids. Designing novel proteins typically follows a two-step approach: establishing the structure first, then utilizing a machine-learning model to generate potential sequences that could adopt that structure. Nature allows for multiple amino acid sequences to fold into identical structures, while a single sequence might adapt to various structures based on the protein's flexibility or functional cues.
When researchers employ AI to design new proteins, the challenge lies in guiding the AI to recognize the existence of numerous beneficial answers—sequences capable of adopting the same fold.
Traditionally, success in protein design has been gauged by whether a model can recreate the protein sequence that natural selection ultimately chose. However, a groundbreaking study by researchers at the Department of Biology suggests this metric may not be optimal. Introducing PottsMPNN, a machine-learning framework developed in the department, has shown promise in enhancing sequence generation and predicting protein stability.
This model incorporates physical principles governing protein structure and stability, offering a more nuanced understanding of the sequence-energy landscape—the relationship between amino acid identity and protein stability.
Foster Birnbaum, a graduate student and lead author of the study, emphasizes that when designing entirely new, designed protein structures, there are no existing native sequences for comparison. Instead, the focus shifts to the likelihood of generated sequences folding into the desired structures, the model's comprehension of the sequence-energy landscape, and its ability to predict the impact of mutations on protein stability.
Similar to how AI has catalyzed social transformations, machine learning has significantly accelerated and broadened fundamental biological research. One key factor contributing to PottsMPNN's effectiveness is its integration of "noise"—variations introduced during training—to reduce the model's tendency to mimic native sequences, thereby increasing the diversity of structures it can generate.
Another crucial component is the pairwise distribution captured by PottsMPNN, which accounts for interactions between amino acids at specific positions within the protein. This consideration, coupled with the inclusion of evolutionarily related sequence sets during training, enables the model to better understand how different sequences can adopt the same folded structure.
By moving away from rigid adherence to native sequences and incorporating evolutionary insights, PottsMPNN has demonstrated improved structural compatibility and energy prediction, particularly for novel proteins that lack natural sequence analogs.
The implications of PottsMPNN's advancements are far-reaching, hinting at the potential for unprecedented biological engineering capabilities. However, Birnbaum remains optimistic about the future of the field, anticipating further refinements and fine-tunings of the model for specific tasks, which could lead to more accurate predictions about the outcomes of particular mutations.
Ultimately, Keating underscores that these methods are steering the field towards the design of novel, non-natural proteins for a wide array of applications, laying a robust foundation for future biological discoveries.
Written by urgent.news from Phys.org's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.