{
  "id": 12434558,
  "title": "Validated Semantic Scoring Improves Filtering Generic Concepts from Knowledge-Graphs for Drug Repurposing Queries",
  "url": "https://urgent.news/2026/10/06/validated-semantic-scoring-improves-filtering-generic-concepts-from",
  "topic": "science",
  "section": "Science",
  "published": "2026-10-06T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.29.755063v1?rss=1"
  },
  "original_language": "en",
  "account": "In the realm of drug repurposing, knowledge graphs (KGs) serve as valuable repositories for information about genes, pathways, phenotypes, and molecular interactions. However, these KGs can sometimes yield generic concepts that dilute the relevance of search results. In an effort to refine the quality of these results, researchers have explored two primary methods: hand-curated concept block-lists and ontology Information Content (IC) thresholds. Unfortunately, neither of these approaches has been exhaustively validated against expert assessments of generality, and IC is not universally applicable to the majority of KG nodes.\n\nTo address this limitation, researchers conducted a study involving three blinded domain experts who rated a selection of KG concepts for their level of generality. Utilizing a biomedical language model called BioLORD, the researchers developed a semantic precision score (SPS) based solely on a concept's name and description. The SPS was then compared against the hand-curated block-list and two IC thresholds when filtering the results of biomedical queries across 200 diseases.\n\nThe study yielded compelling results. The SPS achieved a Spearman correlation coefficient of 0.76 with the experts, closely matching their judgments of generality. Importantly, the SPS maintained its ordering within every IC band and remained informative for the 50 concepts lacking IC values. In comparison, the IC threshold only agreed moderately with the experts (Spearman {rho} = 0.59) and demonstrated limited resolution within IC bands.\n\nPerhaps most striking was the performance of the SPS in removing generic concepts. At its decision boundary, the semantic filter eliminated 1,502 out of 33,161 query answers, surpassing the block-list's efficacy by nearly threefold (513). Furthermore, the SPS filtered fewer query-specific DrugCentral indications than the block-list, while preserving the integrity of approved drugs in DrugCentral. When the comparison was repeated on an additional 200 diseases using a model frozen at its original state, the SPS continued to outperform the block-list, removing 5,135 of 91,203 answers compared to the block-list's 1,419. Crucially, the SPS did not eliminate approved drugs more frequently from the half of DrugCentral unseen by the model during training than from the half it had been trained on.\n\nThese findings suggest that semantic meaning, as captured by the SPS, more faithfully aligns with expert judgments of specificity than ontology position or IC. Consequently, the researchers argue that semantic scoring can serve as a viable alternative to IC and manually maintained block-lists for filtering drug-repurposing results. By harnessing the power of semantic analysis, researchers may be able to refine the quality of KG-derived information, ultimately leading to more accurate and relevant drug repurposing insights.",
  "summary": "When an ontology-based knowledge graph (a KG, containing information about genes, pathways, phenotypes, and molecular interactions) is analyzed using a biomedical computational reasoning system, the system can return unwanted generic concepts. Such generic concepts, in the context of a query about potential drug treatments for a disease, might include \"pharmaceutical preparations\" or…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}