AtlasFold: Protein structure prediction with metagenomic-scale language models
Protein language models (PLMs) trained on evolutionary sequences learn representations that encode protein structure, enabling direct structure prediction without multiple-sequence alignments (MSAs). Here we present the Atlas model family, an open and trainable system spanning protein language modeling, monomer folding, and protein-complex prediction. AtlasLM-3B is a 3B-scale language model…
Protein language models (PLMs) trained on evolutionary sequences can predict protein structures without the need for multiple-sequence alignments (MSAs). The Atlas model family, an open and trainable system, encompasses protein language modeling, monomer folding, and protein-complex prediction. AtlasLM-3B is a 3 billion-scale language model trained with masked language modeling on around 1.56 billion sequences, including metagenomic data. This model surpasses the similarly sized ESM2-3B in unsupervised contact prediction.
AtlasFold, which builds on these representations, predicts all-atom protein structures and attains state-of-the-art accuracy among PLM-based folding models. By fine-tuning AtlasFold for protein-complex prediction, AtlasFold-M, or AtlasFold for Multimer prediction, is created. This protein-specific folding architecture allows for fast, memory-efficient inference with AtlasFold and AtlasFold-M.
The training code, data, stage checkpoints, and model weights are released under the MIT License, providing a foundation for advancing PLM-based protein structure prediction.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.