Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis

Published in arXiv preprint, 2026

Recommended citation: Jiatong Li, Weida Wang, Changmeng Zheng, Shufei Zhang, Yatao Bian, Xiao-yong Wei, and Qing Li. (2026). Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis. arXiv:2607.01800. https://arxiv.org/abs/2607.01800

Cite this work View paper

This work introduces a Molecular Perturbation framework that creates syntax-valid structural variants under controlled graph edit distance. Even a single edit can substantially reduce performance on common molecular tasks, revealing a narrow local trust region and fragile sensitivity to structural change.

The experiments further show that in-context tuning with structurally similar molecules can partially expand that trust region, suggesting a practical direction for more robust molecular language models.

Read the paper