Machine learning techniques can be employed to train models to learn the potential energy surface (PES) of a molecular systems, evaluated at arbitrary level of theory (LOT). Once trained, these Machine-Learning Potentials (MLPs) or Machine-Learning Interatomic Potentials (MLIPs) are able to make energy and force predictions with an accuracy comparable to the LOT of the training data, while being orders of magnitude less computationally expensive. Recent advances in MLP architecture have made it possible to simulate systems containing tens of thousands of atoms over nanosecond timescales.
Although modern foundation models have demonstrated to be great tools for dynamical studies, optimizing the accuracy and efficiency often requires training smaller, system-specific models. Training MLPs from scratch is a system-specific process and often non trivial, as the training set should sufficiently cover the regions of phase space explored during production runs. To avoid the need for expensive ab-initio molecular dynamics simulations to generate training data, active learning (AL) schemes are often used to generate accurate models from scratch. During AL, the model is iteratively refined by sampling an increasingly broad parts of phase space and extending the training set with these new configurations. Enhanced sampling is often used inside of these AL schemes to sample configurations that might otherwise not be encountered using regular dynamics, e.g. transition states of activated process. We develop new algorithms and schemes to efficiently train and deploy MLPs for rare events in nanoporous materials, fully leveraging the potential of modern supercomputing infrastructures.
If the system under investigation is too large to be evaluated in full with quantum mechanical techniques, conventional training schemes might not be directly applicable. We work on algorithms for an automated extraction of representative clusters from the original material that incorporate all of the relevant physico-chemical interactions. This cluster based learning methodology has already proven to be a data efficient way to train MLPs for large and/or disordered systems.
Computing correct dynamics and observables can sometimes require LOTs well beyond standard density functional theory, such as the random phase approximation or coupled cluster. These techniques are computationally very expensive and scale unfavourably with system size. Therefore, it might become computationally intractable to label sufficiently large training sets. To mitigate the computational cost, we develop transfer-learning techniques, in which a large dataset generated at a lower level of theory is combined with a smaller, high-level dataset, allowing the model to efficiently learn the energy correction to go from the low- to the high LOT.
Core publications
Machine learning potentials for metal-organic frameworks using an incremental learning approach. S. Vandenhaute, M. Cools-Ceuppens, S. DeKeyser, T. Verstraelen & V. Van Speybroeck (2023) npj computational materials, 9: 19. doi: 10.1038/s41524-023-00969-x
Cluster-Based Machine Learning Potentials to Describe Disordered Metal–Organic Frameworks up to the Mesoscale. P. Dobbelaere, S. Vandenhaute & V. Van Speybroeck (2025) Chemistry of Materials, 37(15): 5696-5709. doi: 10.1021/acs.chemmater.5c00821
The Operando Nature of Isobutene Adsorbed in Zeolite H−SSZ−13 Unraveled by Machine Learning Potentials Beyond DFT Accuracy. M. Bocus, S. Vandenhaute & V. Van Speybroeck (2024) Angewandte Chemie International Edition, 64(1): e202413637. doi: 10.1002/anie.202413637