Urgent.News

What's breaking now, across thousands of outlets.

Science

Validated Semantic Scoring Improves Filtering Generic Concepts from Knowledge-Graphs for Drug Repurposing Queries

When an ontology-based knowledge graph (a KG, containing information about genes, pathways, phenotypes, and molecular interactions) is analyzed using a biomedical computational reasoning system, the system can return unwanted generic concepts. Such generic concepts, in the context of a query about potential drug treatments for a disease, might include "pharmaceutical preparations" or…

In the realm of drug repurposing, knowledge graphs (KGs) serve as valuable repositories for information about genes, pathways, phenotypes, and molecular interactions. However, these KGs can sometimes yield generic concepts that dilute the relevance of search results. In an effort to refine the quality of these results, researchers have explored two primary methods: hand-curated concept block-lists and ontology Information Content (IC) thresholds.

Unfortunately, neither of these approaches has been exhaustively validated against expert assessments of generality, and IC is not universally applicable to the majority of KG nodes.

To address this limitation, researchers conducted a study involving three blinded domain experts who rated a selection of KG concepts for their level of generality. Utilizing a biomedical language model called BioLORD, the researchers developed a semantic precision score (SPS) based solely on a concept's name and description. The SPS was then compared against the hand-curated block-list and two IC thresholds when filtering the results of biomedical queries across 200 diseases.

The study yielded compelling results. The SPS achieved a Spearman correlation coefficient of 0.76 with the experts, closely matching their judgments of generality. Importantly, the SPS maintained its ordering within every IC band and remained informative for the 50 concepts lacking IC values. In comparison, the IC threshold only agreed moderately with the experts (Spearman {rho} = 0.59) and demonstrated limited resolution within IC bands.

Perhaps most striking was the performance of the SPS in removing generic concepts. At its decision boundary, the semantic filter eliminated 1,502 out of 33,161 query answers, surpassing the block-list's efficacy by nearly threefold (513). Furthermore, the SPS filtered fewer query-specific DrugCentral indications than the block-list, while preserving the integrity of approved drugs in DrugCentral.

When the comparison was repeated on an additional 200 diseases using a model frozen at its original state, the SPS continued to outperform the block-list, removing 5,135 of 91,203 answers compared to the block-list's 1,419. Crucially, the SPS did not eliminate approved drugs more frequently from the half of DrugCentral unseen by the model during training than from the half it had been trained on.

These findings suggest that semantic meaning, as captured by the SPS, more faithfully aligns with expert judgments of specificity than ontology position or IC. Consequently, the researchers argue that semantic scoring can serve as a viable alternative to IC and manually maintained block-lists for filtering drug-repurposing results. By harnessing the power of semantic analysis, researchers may be able to refine the quality of KG-derived information, ultimately leading to more accurate and relevant drug repurposing insights.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in Science

More from Tuesday 6 October →