{
  "id": 12427292,
  "title": "LOAM: A family of genomic language models trained on long-read soil metagenomes",
  "url": "https://urgent.news/2026/10/02/loam-a-family-of-genomic-language-models-trained-on-long-read-soil",
  "topic": "science",
  "section": "Science",
  "published": "2026-10-02T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.28.755120v1?rss=1"
  },
  "original_language": "en",
  "account": "Soil ecosystems harbor a vast, largely uncharted microbial diversity. Recent advancements in long-read sequencing have significantly improved the recovery and quality of microbial genomes from intricate samples. However, genomic language models have largely been trained on reference genomes or short-read assemblies. This study introduces LOAM, a series of decoder-only genomic language models with parameter counts ranging from 25 million to 624 million, trained solely on long-read environmental metagenomes sourced from Oxford Nanopore technology. The performance of LOAM models increases predictably with model size and the quantity of training tokens. Despite being trained on a modest corpus of sequences, LOAM models outperformed comparably sized models on biological benchmarks and achieved performance levels comparable to significantly larger state-of-the-art models trained on considerably larger datasets. Experiments involving context intervention also revealed that LOAM models utilize genomic information over several kilobases, underscoring the potential benefits of enhanced contiguity provided by long-read metagenome-assembled genomes. In benchmark tasks utilizing probes, the representations were systematically evaluated across hidden layers, revealing that task-relevant biological information is frequently more linearly accessible from intermediate model layers rather than final layers. Notably, variation in zero-shot variant-effect prediction was strongly linked to the presence of homologous target sequences in the pre-training corpus. Overall, these findings establish long-read environmental metagenomes as a viable foundation for training competitive genomic language models and highlight the significance of both model scale and training-corpus composition for biological generalization.",
  "summary": "Soil ecosystems represent a vast, largely uncharacterised reservoir of microbial diversity. Metagenomic assembly has unlocked access to this resource, and recent advances in long-read sequencing have improved the recovery and quality of microbial genomes from complex samples. While genomic language models have proven highly effective at capturing biological concepts from large sequence datasets,…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}