A low-annotation-budget PubMedBERT classifier for chondrogenesis regulator discovery via active learning
Motivation: Biomedical natural language processing (Bio-NLP) classification tasks are often limited by the cost of manual annotation, especially for specialised extraction problems where labelled corpora are scarce. Active learning can reduce this cost by prioritising informative samples, while large language models (LLMs) avoid task-specific annotation, but at high computational cost. We ask…
Biomedical natural language processing (Bio-NLP) classification tasks often struggle with the high cost of manual annotation, particularly for specialized extraction problems where labelled corpora are limited. Active learning offers a solution by prioritizing informative samples, while large language models (LLMs) excel at classification tasks but at a high computational cost.
To determine if an encoder trained via active learning can match LLM performance at a lower computational cost, researchers applied this approach to the identification of chondrogenesis regulators from PubMed abstracts.
The team developed an active learning framework to fine-tune PubMedBERT, a biomedical language model, to classify gene mentions as regulators of chondrocyte differentiation. This approach resulted in an AUC-ROC score of 0.93 using just 1,238 annotated sentences. When compared to open-weight LLMs (Qwen3 and Llama-3.1), PubMedBERT achieved comparable AUC-ROC scores and surpassed all models in terms of precision, reaching 0.69. Importantly, the classifier was able to process the entire corpus much faster than the LLMs.
During the study, the model identified 1,128 candidate regulators from 67,111 gene mentions. Remarkably, it recovered 79% of Gene Ontology-annotated genes related to chondrocyte differentiation, which includes 94 genes. Additionally, the classifier proposed 981 new candidate regulators. The researchers emphasize that this pipeline is not limited to chondrogenesis and can be adapted to identify regulators for other biological processes.
The researchers have made their code available on GitHub at https://github.com/ChondroTextomics/ALRegulatorDiscovery, and the data are accessible on Zenodo at https://doi.org/10.5281/zenodo.22746090. The final model is hosted on the HuggingFace Hub under the name amav/pubmedbert-chondrogenesis-classifier.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.