From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery?
Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating…
Agentic systems are increasingly integrating molecular-design tools, but it remains uncertain which component of the system hampers results in practical projects. To address this, researchers created MAGI, an open-source modular agent that sets goals, initiates and tracks optimization, analyzes structure-activity relationships, and adjusts its approach accordingly.
MAGI can produce molecules directly via the language model (LLM) or by delegating to REINVENT4. The scoring systems can be interchanged behind a unified contract. The researchers tested MAGI across nine retrospective lead-optimization projects from three pharmaceutical firms, simulating the tasks within fixed time limits. Both approaches generated valid molecular structures: LLM proposals stayed closer to local chemistry and achieved comparable or better primary activity in fewer steps, while REINVENT explored a wider range of chemical compounds.
The success of a project depended on the predictive models rather than the generation method: completion was contingent on the model's accuracy with the proposed chemistry, which decreased once the chemistry fell outside the model's applicability domain. In a separate blinded evaluation, chemists could not distinguish MAGI's output from actual expert compounds and found the structure-activity relationship (SAR) reasoning plausible yet incomplete.
Overall, these findings indicate that MAGI can function as a flexible coordination layer within existing computational chemistry workflows. However, the ultimate limit for real-world projects currently lies in the applicability of the scoring systems, not the orchestration of tools.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.