Evaluating Autonomous LLM Agents Across Molecular Prediction and Optimization Benchmarks
Large language model (LLM) agents are increasingly capable of carrying out autonomous computational research, but it remains unclear whether they can develop molecular modeling methods that compete with strong human-developed approaches. Here, we evaluate autonomous method development across four settings: Therapeutics Data Commons (TDC) ADMET tasks, the OpenADMET ExpansionRx Challenge, the…
We haven't written up this one. bioRxiv has the full story — the link below goes straight to it.