The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI
High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public…
The OpenBind project, an open-science initiative, has released its first public dataset focused on protein-ligand structures and their associated binding affinity measurements. This dataset, the largest of its kind, centers around the enteroviral 2A protease and comprises 925 crystallographic binding events from 699 compounds. Each of these 601 compounds has been associated with affinity measurements, creating a comprehensive resource for structure-based AI and molecular discovery.
This dataset was assembled from two stages of research. Initially, a fragment screen identified potential molecules, followed by a series of follow-on compounds. These data points, coupled with biophysical measurements, collectively map the protein-ligand binding modes to the antiviral discovery process.
The OpenBind dataset was employed to assess various structure-based methods in predicting protein-ligand structures and their affinity. These methods include docking and cofolding, and the evaluation revealed several challenges inherent to practical structure-based modelling. Docking accuracy is heavily influenced by the binding-pocket conformation, and poses are notoriously difficult to rank. Furthermore, predicting structure-based affinity continues to pose significant challenges.
However, the dataset also showcased the potential of OpenFold3-p2, a deep learning model. By fine-tuning this model using structures from the fragment screen, substantial improvements in pose prediction and virtual screening for related follow-on compounds were observed. This finding demonstrates how early-stage experimental structures can significantly enhance target-specific model adaptation, paving the way for more accurate structure-based AI in drug discovery.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.