Urgent.News

What's breaking now, across thousands of outlets.

AI

PG-MLD: Physics-Guided Molecular Representation Learning via Dynamic 3D Trajectory Distillation

Molecular representation learning underpins molecular property prediction and drug design by capturing molecular structure-property relationships. SMILES-based molecular language models learn chemical semantics from large-scale unlabeled data and support efficient inference. However, the one-dimensional nature of SMILES constrains their ability to capture 3D geometry and conformational evolution,…

Molecular representation learning is crucial for predicting molecular properties and drug design. SMILES-based molecular language models have learned chemical semantics from large-scale unlabeled data, enabling efficient inference. However, SMILES can only capture 1D structure, limiting their ability to represent 3D geometry and conformational changes. Conventional 3D molecular models demand conformer generation and significant computational resources.

To address this limitation, researchers have introduced PG-MLD, a dynamic 3D-to-1D physical knowledge distillation approach for molecular representation learning. PG-MLD establishes a dynamic 3D physical teacher by integrating equivariant geometric encoding with Liquid Time-Constant modeling. This enables the teacher to capture 3D molecular geometry, electronic descriptors at the atomic level, and conformational evolution.

PG-MLD then transfers this trajectory knowledge to SMILES-based students via atom- and molecule-level representation alignment and cross-modal contrastive learning. The process incorporates masked language modeling where applicable. Remarkably, the distilled students can perform downstream tasks using solely SMILES, without needing conformer generation or molecular dynamics simulations.

The effectiveness of PG-MLD was demonstrated in experiments on MoleculeNet. Across three molecular language student architectures, PG-MLD enhanced overall property prediction performance while preserving SMILES-only inference. Furthermore, the learned representations adeptly encoded 3D geometry and conformational dynamics, proving that dynamic 3D physical knowledge can be successfully transferred to SMILES-based molecular language models with varying architectures.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

Sarvam’s AI Arsenal

In June, Sarvam AI entered India’s unicorn club, raising $234 Mn (about ₹2,210 Cr) in a $300 Mn Series B…

More from Sunday 2 August →