Benchmarking biochemical networks generated by large language models
Computational models of biochemical networks provide frameworks for predicting how molecular cues guide cell decisions. These models are typically limited by the time-intensive manual curation required to extract network mechanisms from incomplete literature. Here, we test whether general-purpose large language models (LLMs) can generate accurate models of signaling and metabolic networks. We…
The study investigates the ability of general-purpose large language models (LLMs) to generate accurate biochemical networks for signaling and metabolic processes. Researchers discovered that LLMs are capable of generating 24-65% of the reactions from literature-curated signaling networks associated with cardiomyocyte hypertrophy, myofibroblast activation, and mechanosignaling.
In terms of model performance, logic-based models constructed from these networks were found to predict responses to perturbations with a degree of accuracy ranging from 6 to 33%.
When applied to metabolic modeling, LLMs demonstrated the capacity to create 64-91% of the reactions within the core Escherichia coli metabolic network. However, the accuracy in predicting substrate utilization exhibited a high degree of variability. The study concludes that while current general-purpose LLMs generate biochemical networks with moderate accuracy, it provides a pipeline and benchmarks that could potentially guide future advancements in the field.
Written by urgent.news from eLife's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

