A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides
Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization based…
Recent progress in artificial intelligence has expedited the identification of bioactive peptides through computational exploration of the extensive peptide sequence landscape. Nevertheless, current peptide generation techniques predominantly employ either distribution-learning models, which produce biologically realistic sequences without consistently enhancing functional activity, or optimization-focused methods, which prioritize prediction accuracy but frequently diverge from the actual distribution of experimentally confirmed peptides.
To rectify this issue, a two-step generative evolutionary framework is introduced that harmonizes distribution learning with evolutionary optimization. The initial step involves leveraging Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) to generate biologically plausible seed peptides.
Subsequently, these peptides serve as the starting point for a Hill Climbing optimization routine that iteratively refines the fitness function score. This novel framework was subjected to evaluation using a dataset containing experimentally verified IL2-inducing peptides. Independent analysis using different IL2 prediction models revealed that the combination of Autoregressive Transformer and Hill Climbing yielded the highest overall performance, attaining a mean IL2 induction confidence score of 0.96, while simultaneously decreasing the Kullback-Leibler divergence from 2.26 for the standalone Hill Climbing approach to 0.75.
A case study on an independent IL13-inducing peptide dataset yielded comparable results, with the ART-initialized Hill Climbing method achieving a mean IL13 induction score of 0.99, and reducing the KL divergence from 1.76 to 0.59. In summary, this framework offers a versatile strategy for harmonizing functional optimization and distributional realism, making it applicable to the discovery of peptides and data augmentation in imbalanced biological datasets, thereby facilitating the generation of high confidence peptides for validation in experimental settings.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.