This dissertation advances Graph Foundation Models by grounding structure preservation as the core principle for learning representations in molecular structure analysis. It presents a series of pre-training and adaptation strategies, UGT, S-CGIB, MVC...
This dissertation advances Graph Foundation Models by grounding structure preservation as the core principle for learning representations in molecular structure analysis. It presents a series of pre-training and adaptation strategies, UGT, S-CGIB, MVCIB, and CaMol, that enable models to learn structurally consistent, transferable, and interpretable representations across chemical domains. (1) UGT preserves local and global graph structure through structural identity and a novel graph transformer, ensuring that structurally similar nodes share consistent representations. (2) S-CGIB applies a conditional information bottleneck principle to capture informative and transferable substructures, enabling compact and interpretable representations. (3) MVCIB extends this principle to multi-view molecular learning under a multi-view conditional information bottleneck framework. (4) CaMol introduces a causal adaptation framework that disentangles causal substructures from confounders, achieving interpretable transfer in few-shot molecular learning. Extensive experiments across diverse dataset benchmarks demonstrate that preserving structure substantially improves generalization, transferability, and interpretability.