{
  "id": 12908181,
  "title": "VINTER: a generative vision-language model for annotating and interrogating spatial tissues",
  "url": "https://urgent.news/2026/10/08/vinter-a-generative-vision-language-model-for-annotating-and",
  "topic": "science",
  "section": "Science",
  "published": "2026-10-08T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.10.01.752251v1?rss=1"
  },
  "original_language": "en",
  "account": "A new generative vision-language model called VINTER has been developed to annotate and examine spatial tissues. Spatial transcriptomics links gene expression to tissue structure, but existing models learn embeddings that require post-processing or do not incorporate measured expression. Therefore, tissue decisions are often made independently without considering morphology, genes, and spatial neighbourhood together.\n\nThe VINTER model was introduced, which is a generative vision-language model aligned on 1.45 million histology-caption pairs. It reads histology, measured gene expression, and spatial neighbourhood as a single sequence. The model returns the probability of every tissue class, allowing for the editing of molecular evidence with the image held constant. By integrating all three sources of information, VINTER achieved higher accuracy than spatial-domain methods and a spatial foundation model in annotating tissue. Additionally, VINTER was able to transfer across cohorts without target labels.\n\nOne of the key features of VINTER is its ability to edit only the neighborhood of a tissue region, which resolved the follicular B-cell and plasma-cell programs behind the reading of tertiary lymphoid structure (TLS) maturation in lung and kidney cancer. Furthermore, the model's class probabilities exposed immune heterogeneity within stroma and cancer that may be hidden beneath benign labels in triple-negative breast cancer.\n\nInterestingly, VINTER was able to transfer to Visium HD and single-cell Xenium data without any retraining and resolved follicular architecture within colorectal cancer TLS. This means that the same model can be used across platforms and resolutions for spatial annotation. Overall, VINTER turns spatial annotation into an interrogable reading of tissue, allowing for a deeper understanding of tissue structure and function.",
  "summary": "Spatial transcriptomics links gene expression to tissue structure, yet current models learn embeddings whose meaning is assigned afterwards or read histology without measured expression. Tissue decisions are therefore rarely made from morphology, genes and neighbourhood together, or examined through them. Here we introduce VINTER, a generative vision-language model aligned on 1.45 million…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}