BELL: Biomodel Evidence and LLM-based Logic
Building accurate and predictive mechanistic models requires careful biological interaction curation and verification against existing knowledge. When done manually, these tasks become impractical, especially with massive extraction of interactions facilitated by advanced natural language processing methods and large language models (LLMs). We present BELL (Biomodel Evidence and LLM-based Logic),…
A new framework called BELL, short for Biomodel Evidence and LLM-based Logic, is revolutionizing the creation of accurate and predictive mechanistic models in biology. This innovative approach tackles the challenges of manually curating biological interaction data, which becomes unfeasible when dealing with large volumes of information extracted through advanced natural language processing techniques and large language models (LLMs).
BELL streamlines the process by automating evidence retrieval, scoring, and explanation for interaction-level verification. It employs a five-step pipeline to process each interaction. The first step involves entity grounding, where the system identifies and defines the relevant biological entities. Next, database ranking helps prioritize the most relevant data sources. Following this, evidence retrieval is carried out from seven biological databases, ensuring a comprehensive collection of relevant information.
The fourth step employs a heuristic scoring system that assigns a score to each interaction based on four dimensions. Finally, a chain-of-thought explanation is generated using LLMs, providing a logical rationale for the curator's action. This comprehensive pipeline has been tested on 210 protein-protein interactions within a curated Glioblastoma Multiforme (GBM) model.
The results of this application demonstrate the effectiveness of BELL. Forty-nine point five percent of the interactions achieved a high confidence level, indicating a strong likelihood of being accurate. Additionally, seventy-two point four percent received a positive curator recommendation, further validating the framework's accuracy. Moreover, qualitative flags were strategically placed, guiding curators' attention to evidence gaps that required further investigation.
BELL has been seamlessly integrated into the KALIMBA curation platform, enhancing its capabilities for biomodel creation. The framework is now accessible at www.boheme.pitt.edu/Kalimba, offering researchers a powerful tool to expedite and improve their biological data curation efforts.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.