MolJam: A Multidimensional Framework for Assessing Molecular Dataset Quality and Its Impact on Machine Learning
High-quality molecular datasets are essential for reliable machine learning in cheminformatics and bioinformatics, yet dataset quality is rarely assessed systematically and its relationship with downstream model performance remains poorly understood. Here, we present MolJam, an open-source framework for quantitative assessment of molecular dataset quality across five dimensions-structural…
High-quality molecular datasets are crucial for dependable machine learning in cheminformatics and bioinformatics, but assessing dataset quality systematically is uncommon, and the connection between quality and downstream model performance is unclear. The team has introduced MolJam, an open-source framework designed to evaluate molecular dataset quality across five dimensions-structural integrity, data quality, experimental information quality, chemical space coverage, and data distribution-using 12 standardized metrics.
When applied to 11 MoleculeNet and eight ChEMBL-derived datasets, MolJam uncovered significant and diverse quality problems, such as undefined stereochemistry in as much as 70.72% of molecules, inconsistent molecular representations, and conflicting labels. The team then explored whether enhancing these quality metrics would improve machine learning performance.
Refined ESOL and Lipophilicity datasets showed increased MolJam quality scores, but the impact on predictive performance was inconsistent, indicating a balancing act between reduced dataset size. Detailed ablation experiments further showed that both dataset quality and data quantity play a role in model performance and, notably, that keeping molecules with partial stereochemical information can sometimes outperform their removal when the benefit of increased data quantity counteracts the quality penalty.
Therefore, molecular dataset curation cannot merely focus on maximizing data cleanliness but must strike a balance among multiple dimensions of data quality while considering information loss. MolJam serves as a standardized tool to identify molecular dataset limitations, compare benchmark quality, and quantitatively evaluate the influence of data curation decisions on downstream machine learning.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.