Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. II. Antibody Properties, Lipid-RNA Interactions
Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their utility for real-world industrial research remains insufficiently characterized. Extending the analysis presented in the first paper of this series, we evaluated the same five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and…
In the second part of a series examining the capabilities of advanced AI frameworks for real-world scientific research, a team of researchers evaluated the same five cutting-edge platforms – Kosmos, K-Dense, ToolUniverse, BioAgents and AI Scientist-v2 – on two significant projects with practical applications in biopharmaceutical development.
The first project focused on predicting antibody developability properties using pretrained protein language model embeddings, while the second examined non-covalent lipid-RNA interactions in lipid nanoparticles through all-atom molecular dynamics (MD) simulations.
The AI frameworks demonstrated genuine strengths, such as the ability to identify subtle methodological issues on their own, successfully utilize pretrained protein embeddings, and consistently report p-values and confidence intervals – data that is often absent in original research papers. However, the researchers found that no framework came close to matching the breadth and depth of the original studies.
Moreover, the AI systems suffered from multiple failures and hallucinations when attempting to reproduce the complex research conducted by pharmaceutical companies.
The results of these evaluations reinforce the conclusion of the first paper in this series – that reproducing real-world published research from pharmaceutical companies is considerably more challenging for current AI frameworks than standard benchmarks suggest. The findings highlight the limitations of these advanced AI systems when confronted with intricate, high-stakes scientific endeavors, emphasizing the significant gap between AI performance on benchmarks and their practical application in real-world research.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.