Detailed curation of biological samples and experimental designs for genomics using LLM-supported agentic workflows
We describe an automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource. Gemma is a hand-curated database of reprocessed transcriptomic studies, currently covering over 23,000 human, mouse and rat data sets largely drawn from the Gene Expression Omnibus (GEO). We developed a pipeline that uses both traditional…
A software tool has been developed to automate data curation tasks for the Gemma genomics data re-analysis resource. Gemma is an existing database containing over 23,000 transcriptomic studies, primarily sourced from the Gene Expression Omnibus (GEO). The new pipeline employs both traditional mechanical methods and large language models to generate precise ontology-anchored annotations at the sample and experiment levels, adhering to established curation guidelines.
The report outlines the pipeline's performance, which closely matches that of human curators. Remarkably, the system achieves this level of accuracy at only 1/20th the cost and with up to 100 times the speed of human curation. The researchers also explored triage methods to identify agent-curated samples with a higher likelihood of errors, which can be forwarded for human review.
The software, along with a benchmark set of 500 studies and an evaluation framework, is presented as potential additions to bioinformatics ecosystems. This innovative approach could significantly streamline and enhance the curation process in the field of genomics.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.


