Automatic bioinformatic software named entity recognition from literature
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here…
Bioinformatics software and databases are crucial elements in contemporary life science research; however, their references in scientific literature are inconsistent and challenging to identify systematically on a large scale. The absence of a thorough and up-to-date catalog of bioinformatics resources hampers endeavors towards automated biomedical knowledge extraction and simplified data analysis.
The researchers introduce SNAIL, a hybrid named entity recognition framework aimed at automatically detecting bioinformatics software and database (SW/DB) references from biomedical texts. SNAIL combines lexical and semantic modeling techniques. The lexical component recognizes orthographic patterns and contextual clues suggestive of SW/DB names, while the semantic component utilizes contextual embeddings produced by transformer-based language models like SciBERT, augmented with an explicit token-masking technique to improve entity-focused representations.
A substantial training corpus was generated automatically through a hybrid pipeline that merges citation-hinted extraction with assistance from large language models. When evaluated on two separate benchmark datasets and real-world research articles, SNAIL significantly surpasses existing methods, including domain-specific techniques such as bioNerDS2 and general-purpose large language models like ChatGPT, Gemini, Grok, and Claude.
Implementing SNAIL in large-scale literature analysis uncovers distinct journal-level preferences across various bioinformatics subfields. These findings indicate that SNAIL offers an accurate and scalable solution for identifying bioinformatics resources in scientific texts, facilitating systematic meta-analysis of tool usage and research trends.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.