AI Models Excel at Orchestration but Falter at Biological Judgment: Findings from an Agentic Gene Annotation Study
Large language model (LLM) agents are increasingly used both to direct biological analyses and to interpret their results. The core functions of agents-workflow control and biological adjudication-are often combined within the same agent and evaluated end-to-end, making it difficult to determine whether a model that is useful in one role is also reliable in the other. In this study, Genome…
We haven't written up this one. bioRxiv has the full story — the link below goes straight to it.