Urgent.News

What's breaking now, across thousands of outlets.

AI

End-to-end multimodal pathology foundation model with clinical dialogue

Nature Medicine, Published online: 31 July 2026; doi:10.1038/s41591-026-04521-4 Trained on 2.3 million whole-slide images and 14 million clinical question−answer pairs, the multimodal foundation model PRISM2 matched clinical-grade cancer detection performance without task-specific training, demonstrating the value of clinical dialogue supervision for computational pathology.

Recent advances in computational pathology have been facilitated by foundation models, which move beyond encoding image patches to whole-slide understanding. However, their clinical utility is still limited. The authors introduce PRISM2, a multimodal slide-level foundation model trained on a vast dataset of 2.3 million whole-slide images and 14 million question-answer pairs derived from 700,000 pathology reports.

Through clinical dialogue supervision, PRISM2 aligns histomorphology with diagnostic reasoning, producing representations that support both prompt-based inference and transferable embeddings for various tasks.

PRISM2 achieves performance comparable to clinical-grade products for cancer detection in prostate, breast, and breast lymph node samples. Its embeddings outperform previous foundation models in comprehensive diagnostic, biomarker, and survival benchmarks. Task-specific fine-tuning on survival prediction surpasses training from scratch on a large survival dataset.

This demonstrates the potential of language-supervised pretraining to create scalable, clinically grounded pathology representations that bridge human diagnostic reasoning and foundation model performance.

The field of computational pathology has seen significant progress with models like Virchow2, UNI2, and H-Optimus-1, trained on millions of histopathology tiles and employing self-supervised objectives. However, these models are limited in case-level inference, which requires aggregating information from whole-slide images. Current workflows often use weakly supervised learning for task-specific tile aggregators, but this approach lacks robustness and requires extensive data curation.

PRISM2 addresses this issue by introducing a novel dual-embedding architecture. The 'base' embedding transfers well to complex tasks like biomarker prediction, while the 'diagnostic' embedding is derived from a 4-billion-parameter language model, tuned for cancer detection and subtyping. Despite being trained to replicate clinical dialogue, PRISM2 is not intended for conversational use but rather to utilize single-turn dialogue as a rich supervisory signal for learning generalizable slide-level representations.

PRISM2 undergoes a two-stage training process. In the first stage, a Perceiver-based slide encoder aggregates tile embeddings from Virchow2. This representation is aligned with diagnostic summary text using contrastive objectives based on BioGPT text embeddings and autoregressive objectives using the Phi-3 Mini language model. The second stage involves fine-tuning the Phi-3 Mini language model, leveraging its reasoning capabilities to enable prompt-based inference.

This capability allows PRISM2 to serve as a specialized agent for an interactive assistant model, providing second opinions, pre-populating pathology reports, and answering queries during routine sign-out.

Written by urgent.news from Nature Medicine's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at nature.com →

More in AI

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more…

Amazon’s stock pops on roaring cloud growth and soaring AI demand

Amazon.com Inc. delivered a solid earnings and revenue beat as it posted its second-quarter financial results, driven by surging growth in its cloud infrastructure business. Demand for artificial intelligence was the primary factor in that growth, causing the company to boost its capital expenditure forecast once again.

More from Friday 31 July →