Urgent.News

600+ sources. One page. See who else covered it.

Editions

Science

Protein language models and the long tail of functional diversity

Protein language model performance on downstream tasks depends on the pretraining data, motivating recent efforts to combine genomic- and metagenomic-derived protein sequences into large-scale atlases. Because these datasets are highly redundant, sequences are typically clustered by similarity and sampled during training. Sequences that do not belong to any cluster, known as ``singletons'', are…

We haven't written up this one. bioRxiv has the full story — the link below goes straight to it.

Read the original at biorxiv.org →

More in Science

More from Friday 14 August →