BioSecBench-Function: A Verifiable Benchmark for Reasoning about Biological Function from Experimental Data
Inferring biological function from experimental data is central to understanding emerging pathogens and developing effective countermeasures, yet interpreting these data remains slow and expert-intensive. AI agents could help accelerate this process by reasoning across sequence, structural, and biophysical evidence. We present BioSecBench-Function, a verifiable benchmark for recovering…
Inferring biological function from experimental data is crucial for understanding emerging pathogens and devising effective countermeasures. However, interpreting experimental data can be a slow and resource-intensive process that requires expert knowledge. AI agents have the potential to expedite this procedure by reasoning across sequence, structural, and biophysical evidence.
The researchers have introduced BioSecBench-Function, a verifiable benchmark designed to assess a model's capacity to deduce biosecurity-relevant functions from real biological data. The benchmark consists of 111 evaluations that are constructed from published datasets and graded in a deterministic manner based on the ground truth.
These evaluations are organized along two dimensions: threat axis, which includes seven biosecurity-relevant question types, and biological question, which indicates whether the solution primarily relies on sequence, structure, or biophysical assay data.
In a comprehensive study involving 7,326 runs from twenty-two model-harness configurations, Opus 5 under Claude Code achieved the highest endpoint pass rate at 50.3%. Grok 4.6 under Grok Build had the highest overall pass rate at 44.1%, counting refusals as failures. The performance varied significantly across both model-harness configurations and task categories.
It was also observed that refusal rates differed sharply between model providers, and cost was not a reliable predictor of accuracy. Several configurations demonstrated a pass rate exceeding 40% even at a low cost.
BioSecBench-Function serves as a benchmark for evaluating whether AI agents can reliably interpret the functions of newly discovered pathogens or variants. As the next outbreak occurs, this benchmark will be instrumental in determining the trustworthiness of AI agents in interpreting biological data for biosecurity purposes.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.