Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design
Generative artificial intelligence is accelerating molecular design, yet the relative suitability of available models for different targets and stages of preclinical drug discovery remains unclear. Here we benchmarked 12 molecular generation and optimization methods across 176 curated protein ligand systems spanning diverse therapeutic target classes, with experimentally validated ligands…
The systematic evaluation of 12 artificial intelligence-based molecular generation and optimization methods across 176 curated protein-ligand systems has been conducted. This study aimed to determine the relative suitability of these models for different stages of preclinical drug discovery.
The methods under scrutiny included pocket conditioned 3D generation, diffusion and flow based modeling, autoregressive construction, reference conditioned optimization, and synthesis-aware design. Each approach was assessed based on several operational metrics, including robustness, chemical validity, uniqueness, diversity, drug likeness, synthetic accessibility, docking, physicochemical properties, ADMET properties, and computational resource requirements.
The findings indicated that different methods exhibit trade-offs in their performance. For instance, receptor conditioned methods leverage binding pocket geometry, flow-based approaches facilitate efficient sampling, reference-conditioned methods favor analogue generation, and synthesis-aware methods enhance chemical feasibility. However, no single method excels in all these criteria.
To further assess the functional potential of generated molecules, a state-aware functional classifier (SAFC) was developed. This classifier integrates molecular dynamics derived receptor ensembles, ensemble docking, and protein-ligand interaction graphs. The SAFC provided rankings of generated molecules for functional activity that were partly complementary to docking, drug likeness, and synthetic accessibility scores.
These results suggest that a hybrid, stage-specific deployment of generative models would be more effective than relying on any single architecture or evaluation metric. This study offers practical guidelines for integrating generative AI into preclinical drug development processes.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.