{
  "id": 4341917,
  "title": "Corpusome, a cross-body-site human microbiome corpus for representation learning",
  "url": "https://urgent.news/2026/08/29/corpusome-a-cross-body-site-human-microbiome-corpus-for",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-29T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.08.28.747922v1?rss=1"
  },
  "original_language": "en",
  "account": "Corpusome, a new human microbiome database, was recently made available for machine-learning research. This resource addresses the challenge of limited cross-body-site representation and cross-study generalization in existing models of the human microbiome. The problem stems from a lack of a harmonized multi-body-site corpus that incorporates technical metadata necessary for effective modeling.\n\nCorpusome aims to overcome these limitations by releasing a harmonized two-tier cross-body-site human microbiome corpus. This collection contains 187,546 samples, integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. The database follows a two-tier design that maintains both functional depth and cross-body-site breadth.\n\nThe shotgun tier of Corpusome consists of 22,588 samples from 93 studies, providing species- and pathway-level profiles. Meanwhile, the 16S tier includes 164,958 samples from the complete pull of 708 MGnify studies, offering genus-level profiles that extend coverage to oral, skin, respiratory, and urogenital sites. The database spans six body sites and two modalities, with harmonized metadata designed for batch-aware modeling.\n\nOne of the key strengths of Corpusome is its ability to highlight body-site signal that exceeds technical/source variance. In the 16S tier, this improvement is approximately 2.4-fold. This enhanced representation of body-site-specific information is crucial for developing accurate and generalizable machine-learning models of the human microbiome.",
  "summary": "Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}