Urgent.News

What's breaking now, across thousands of outlets.

Science

Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction

Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural…

Pep-PU-GAN introduces a novel deep learning framework for peptide classification, addressing the challenge of limited labeled data in bioinformatics. This framework leverages positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural networks (GNNs) to effectively classify peptides, despite the scarcity of verified negative examples.

Peptides are represented as residue graphs, consisting of amino acid nodes connected by edges representing adjacent residues. This graph representation allows for attention-based message passing over local neighborhoods, capturing the complex relationships between amino acid sequences and their biological functions.

The architecture of Pep-PU-GAN comprises a generator that synthesizes peptide embeddings in the encoder space, and a dual-function discriminator that simultaneously discriminates between real and synthetic embeddings while performing PU classification. Training is conducted using a custom loss function that integrates non-negative PU (nnPU) risk estimation with adversarial objectives, optimizing the model's ability to distinguish between positive and unlabeled peptides.

To further enhance the training process, a self-training mechanism is incorporated. This mechanism utilizes high-confidence synthetic positive embeddings to augment the training set, effectively improving the model's performance. The Pep-PU-GAN framework was evaluated on a dataset of neuropeptides, consisting of 4,049 verified positive neuropeptides and 8,558 unlabeled peptides.

Compared to baseline models, Pep-PU-GAN demonstrated superior performance, achieving an F1 score of 0.93 and an AUROC of 0.98 on an independent held-out benchmark.

The proposed Pep-PU-GAN approach offers a promising solution for peptide classification tasks that face challenges such as scarce labeled data and abundant unlabeled data. By harnessing the power of PU learning, GANs, and GNNs, this framework has the potential to significantly impact computational biology and drug discovery domains, where peptide functionality prediction plays a crucial role.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Friday 18 September →