Urgent.News

What's breaking now, across thousands of outlets.

Science

Multi-Peptide Prompting Enables In-Context Learning in Protein Language

Protein language models (PLMs) are trained primarily on individual protein sequences, yet many peptide-discovery problems require inference from only a small number of labeled examples. Here, we show that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification. We introduce multi-peptide example prompts…

Protein language models (PLMs) typically learn from individual protein sequences, but peptide-discovery tasks often rely on limited labeled examples. Researchers have discovered that single-sequence PLMs can effectively perform in-context learning for peptide tasks without requiring any gradient updates, retraining, or architectural changes.

The team introduces multi-peptide example prompts (MPEPs), which concatenate demonstration peptides with glycine spacers to serve as context for predicting the probability of query peptides. They tested this method on three different peptide tasks - synthetic pattern-completion, secondary-structure classification, and MHC-II binder prediction - using both encoder-only ESM-2 models and decoder-only ProGen2 models.

The results show that performance improves with more peptide examples and larger model size, suggesting that PLMs can identify shared sequence-level properties from prompted examples. A new difference score was also developed to address compositional biases in raw PLM probabilities, which further enhances classification accuracy.

On MHC-II binder prediction, classification based on MPEPs with larger ESM-2 models achieved results similar to or better than low-data classifiers trained on frozen ESM-2 embeddings, all without any additional training. This study uncovers an unforeseen capability of in-context inference in single-sequence PLMs and presents MPEP conditioning as a simple and efficient approach for dealing with limited peptide classification data.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Thursday 27 August →