{
  "id": 10815425,
  "title": "A low-annotation-budget PubMedBERT classifier for chondrogenesis regulator discovery via active learning",
  "url": "https://urgent.news/2026/09/29/a-low-annotation-budget-pubmedbert-classifier-for-chondrogenesis",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-29T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.24.754045v1?rss=1"
  },
  "original_language": "en",
  "account": "Biomedical natural language processing (Bio-NLP) classification tasks often struggle with the high cost of manual annotation, particularly for specialized extraction problems where labelled corpora are limited. Active learning offers a solution by prioritizing informative samples, while large language models (LLMs) excel at classification tasks but at a high computational cost. To determine if an encoder trained via active learning can match LLM performance at a lower computational cost, researchers applied this approach to the identification of chondrogenesis regulators from PubMed abstracts.\n\nThe team developed an active learning framework to fine-tune PubMedBERT, a biomedical language model, to classify gene mentions as regulators of chondrocyte differentiation. This approach resulted in an AUC-ROC score of 0.93 using just 1,238 annotated sentences. When compared to open-weight LLMs (Qwen3 and Llama-3.1), PubMedBERT achieved comparable AUC-ROC scores and surpassed all models in terms of precision, reaching 0.69. Importantly, the classifier was able to process the entire corpus much faster than the LLMs.\n\nDuring the study, the model identified 1,128 candidate regulators from 67,111 gene mentions. Remarkably, it recovered 79% of Gene Ontology-annotated genes related to chondrocyte differentiation, which includes 94 genes. Additionally, the classifier proposed 981 new candidate regulators. The researchers emphasize that this pipeline is not limited to chondrogenesis and can be adapted to identify regulators for other biological processes.\n\nThe researchers have made their code available on GitHub at https://github.com/ChondroTextomics/ALRegulatorDiscovery, and the data are accessible on Zenodo at https://doi.org/10.5281/zenodo.22746090. The final model is hosted on the HuggingFace Hub under the name amav/pubmedbert-chondrogenesis-classifier.",
  "summary": "Motivation: Biomedical natural language processing (Bio-NLP) classification tasks are often limited by the cost of manual annotation, especially for specialised extraction problems where labelled corpora are scarce. Active learning can reduce this cost by prioritising informative samples, while large language models (LLMs) avoid task-specific annotation, but at high computational cost. We ask…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}