Urgent.News

What's breaking now, across thousands of outlets.

AI

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate…

LifeSciBench is a new benchmark comprising 750 tasks meticulously crafted by experts to assess the capabilities of language models in handling realistic life science research tasks. Traditional life sciences benchmarks often lack the complexity and breadth needed to reflect real-world research scenarios, which frequently involve ambiguity and multiple dependent judgments.

What sets LifeSciBench apart is its comprehensive coverage of seven representative scientific workflows and seven life science domains, each paired with a rubric created by human experts.

When compared to five state-of-the-art models, it was found that GPT-Rosalind excelled, achieving a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1%. This means that, on average, the model was able to pass approximately one-third of the tasks it encountered. However, the data also revealed a significant gap in performance, with 171 tasks (22.8%) not being successfully passed by any of the evaluated models and another 261 tasks (34.8%) receiving a pass rate below 20% even from the best-performing model.

These findings suggest that LifeSciBench provides a detailed and rigorous evaluation of practical scientific reasoning and decision-making skills in the life sciences.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

More from Tuesday 18 August →