Zhihui Tian, Kang Yang, Michael Tonks, Amanda R. Krause, Joel B. Harley
3D-PRIMME learns a physics-regulated local evolution rule for three-dimensional grain growth from only two consecutive time steps. After training on a 100^3 grid with 512 grains, the authors report applying the operator without retraining to 1024^3 grids with 550,000 grains while preserving coarsening kinetics and grain topology. The local rule allows grain-growth surrogates to move several orders of magnitude in system size without sacrificing the physical statistics they were trained to reproduce.
The authors introduce the Novelty-Tiered Affinity Benchmark, which partitions test data into ligand novelty tiers to control target-mirroring leakage. They report that a ChEMBL 36 meta-analysis identifies more than 6,000 such assay pairs and that ligand-only models fall to r = 0.14 in the most challenging tier. The benchmark gives affinity models a more direct test of whether they generalize beyond localized leakage and memorized training data.
The active-learning workflow uses last-layer-projection regression as a cheap per-configuration uncertainty estimate for machine-learning force fields. The authors report that LLPR-selected subsets recovered full-data accuracy across molecular, condensed-phase, and electrolyte systems using only a small fraction of the electronic-structure labels. A forward-pass uncertainty measure avoids the separate fine-tuning runs that make model committees impractical for foundation force fields.
Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Daryna Safarian, Ming Han, Juan J. de Pablo
HiPoly processes complete polymer descriptions through a three-level hierarchical graph architecture that encodes connectivity, composition, and molecular weight. The authors report state-of-the-art thermophysical-property prediction for multicomponent polymer systems, with ablations supporting each hierarchical design choice. A shared polymer representation can connect formulation data, prediction, generative design, and simulation-based validation across polymer chemistries.
Marcel Muller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garcia Carrillo, Yeonghun Kang, Juan B. Perez-Sanchez, Simone Pilon, Martin Fitzner, Timothy Noel, Frank Gu, Varinia Bernales, Alan Aspuru-Guzik
La Agente Optima constructs and supervises Bayesian optimization campaigns while maintaining persistent optimization state and separating reasoning from execution. In a multi-objective flow-chemistry campaign, the authors report increased yield over a sequence of experiments. Keeping routine loops executable and decisions auditable could make long-running optimization campaigns accessible without specialist campaign setup.
The framework audits surrogate-driven search by reserving true evaluations for every certified conclusion, then derives rank-preservation and query-complexity conditions for when a model may act as an oracle. Across 432 surrogate fits over six task-regime conditions, the authors report that the audit statistic tracked deployed search performance at Spearman rank correlation 0.80-0.99, while audited screening reduced certified oracle cost by a measured factor of 25. The result turns trust in a surrogate into an auditable design condition, separating the model's role in proposing candidates from the true evaluations required to certify them.
Britton, D., Ghose, D. A., Halpin, J. C., Birnbaum, F., Gundu, K., Raval, S., Papanastasiou, M., Carr, S. A., Keating, A. E.
The method guides diffusion-based protein-backbone generation with PDB-derived seed fragments selected for geometric complementarity to the target surface. The authors report that screening 1,402 designs produced multiple functional binders, including nanomolar-to-low-micromolar variants and one with neutralization comparable to the native antitoxin peptide. Surface-complementing seeds expand the structural and contact patterns available to generative binder design at difficult multi-site interfaces.
The workflow diversifies a parent protoglobin active site and uses active-learning-assisted directed evolution to predict promiscuous variant libraries across carbene and nitrene transfer reactions. Across 26 reactions, the authors report improved activity and selectivity for every parent-catalyzed reaction in at least one predicted-library member, plus activity on five of ten reactions absent from the parent enzyme. The strategy produces enzyme libraries that retain parent functions while adding transformations outside the parent's catalytic repertoire.
The hybrid holey-ML approach switches from machine-learning interatomic potentials to ab initio quantum chemistry when the predicted gap between electronic states becomes small. The authors report an order-of-magnitude cost reduction while reproducing the excited-state population decay obtained from fully ab initio simulations. This adaptive handoff makes nonadiabatic dynamics more practical without trusting the learned potential near the conical intersections where it fails.
MS-GPT queries a conditional molecule-language model with a calibrated band of spectrum-induced fingerprints, then ranks pooled candidates by generation-frequency consensus. The authors report new state-of-the-art exact-match accuracy on NPLIB1 and MassSpecGym at both top-1 and top-10 ranks. Treating the noisy fingerprint as a posterior rather than a thresholded answer preserves uncertainty where de novo spectral identification needs it most.
Ling Tang, Weiyi Xia, Tyler J. Slade, Paul C. Canfield, Cai-Zhuang Wang
The multi-minima iterative genetic algorithm couples an artificial-neural-network interatomic potential to a metadynamics-inspired penalty that steers search away from previously explored basins. The authors report recovery of the synthesized ground-state structure of La4Co4Pb and an exact match to the independently measured crystal structure of La5CoPb2 using composition alone. Navigating both the global minimum and nearby metastable states gives data-driven structure prediction a route beyond recombining entries from known crystal databases.
LLM4MOF assigns one language-model agent to propose chemically interpretable design hypotheses and another to translate those hypotheses into constrained metal-organic framework candidates for simulation and revision. The authors report that ten autonomous iterations concentrate searches on strong candidates across six tasks within 400 property evaluations and outperform random search and a genetic algorithm during de novo simulated design. Because each candidate remains tied to explicit choices about nodes, linkers, pores, and functional chemistry, the loop can expose why a design succeeds instead of returning only a score.
Spatium is a protein language foundation model trained on more than 51 million cells across spatial-proteomics platforms to learn protein co-expression hierarchies that remain stable across panel composition and measurement scale. The authors report that Spatium recovered cell identities with known marker patterns, separated spatial microenvironments by coherent enrichment signatures, and reconstructed missing proteins while preserving biological expression patterns. A panel-robust representation lets heterogeneous spatial-proteomics datasets support cell-identity, microenvironment, and missing-marker tasks with only lightweight adaptation.
Leonard Moracchini, Thomas Pigeon, Morgane Menz, Thibault Faney, Thomas D. Swinburne, Mihai-Cosmin Marinica
The framework combines Adaptive Multilevel Splitting with Girsanov reweighting to propagate interatomic-potential parameter uncertainty into averaged committor probabilities without resampling reactive trajectories for each parameter realization. Across a Muller-Brown potential, a solvated dimer, and a butane conformational transition, the authors report recovery of reference rare-event probabilities and, under mild basin-accuracy assumptions, uncertainty bounds on reaction rates. Path-space reweighting provides uncertainty-aware committors and rate bounds from an existing rare-event trajectory ensemble, avoiding a new simulation for every potential realization.
Serrano, L. R., Zhou, A., Wei, Z., Stocks, K.-L. K., Ektefaie, Y., Gwynne, P. J., Chen, E., Krieger, I., Sacchettini, J., Aldridge, B., Hu, L. T., Farhat, M. R.
The study compares three active-learning policies for whole-cell bacterial bioactivity, selects a balance between predicted novelty and hit rate in retrospective tuberculosis data, and then closes the loop in a Borrelia antibiotic screen. The authors report a five-fold hit-rate increase from 0.2% to 1.0% in the closed-loop screen, 53-fold enrichment with 11.0% validation in prospective selections, and intended narrow-spectrum activity for all validated hits. The acquisition strategy couples exploration and exploitation to experimental feedback, allowing a whole-cell predictor to generalize into chemically diverse, out-of-distribution screens.
A structure-aware graph neural network learns cross-functional energy residuals from 380,190 paired PBE-r2SCAN structures, converting legacy PBE formation energies onto the r2SCAN scale. The authors report a mean absolute error of 14.3 meV/atom when converting PBE energies to r2SCAN-level accuracy, compared with 18.2 meV/atom for CHGNet. Cross-functional alignment makes large legacy DFT collections usable alongside higher-fidelity calculations for materials foundation models and thermodynamic prediction.
Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
The authors formulate reaction prediction as discrete flow matching over graph-structured electron occupation vectors in a continuous-time Markov chain. On USPTO-480K and the stated out-of-distribution settings, they report competitive prediction, retained performance, mechanism-like trajectories, and side-product prediction. Electron redistribution offers a mechanistically legible alternative to product generation and direct graph edits.
Denis Blessing, Mouyang Cheng, Maximilian Schebek, Jutta Rogal, Mingda Li, Carles Domingo-Enrich,, Yuanqi Du
JANUS couples continuous and masked discrete diffusion in an equivariant graph neural network trained directly from energy evaluations, sampling both atomic identities and structure. The authors report more than three orders-of-magnitude fewer energy evaluations while reproducing reference equilibrium observables and phase behavior in benchmark systems. Coupling composition with relaxation offers a unified route to thermodynamic sampling and inverse design in chemically disordered materials.
Meng Gao, Armin Shayesteh Zadeh, Aniruddha Seal, Siva Dasetty, Siddarth K. Achar, Misko Dzamba, Benjamin K. Miller, Leif D. Jacobson, C. Lawrence Zitnick, Brandon M. Wood, Zachary W. Ulissi, Daniel S. Levine, Andrew L. Ferguson
The authors use the eSEN-omol machine-learned interatomic potential for complete all-atom enzymes in explicit solvent. They report experimental barrier trends for chorismate mutase and mechanistic alternatives for metal-activated phosphoryl transfer. The result points to a practical route for extending quantum-accurate catalytic simulations beyond system-specific hybrid setups.
Junjie Wang, Yijie Zhu, Zhongwei Zhang, Zhiyue Guo, Xudong Zhu, Lixin He, Chi Ding,, Jian Sun
The authors build HotPP-Spin with Cartesian tensor equivariant message passing and explicit axial-vector magnetic moments. For H-phase monolayer VSe2, they report a finite-size ordering crossover at 415--435 K, close to the reported experimental value. This provides one representation for connecting first-principles magnetic energetics to coupled spin and structural simulations.
Tiancheng Li, Jianming Xue, Linfeng Zhang, Duo Zhang, Han Wang
DPA4C co-designs an equivariant interatomic-potential architecture with compressed CUDA operators to pursue accuracy and throughput under deployment constraints. The authors report that its largest variant approaches MACE-Omat accuracy at roughly two orders of magnitude higher throughput, while all variants complete multimillion-atom simulations on one GPU. This moves quantum-trained universal potentials closer to the speed and scale traditionally reserved for empirical force fields.
The authors add molecular geometry and directionality to coarse-grained potentials through anisotropic descriptors and symmetry-adapted message passing. Using Gay-Berne particles and benzene, formamide, and water representations, they report improved energy, force, and torque prediction. The work identifies retained symmetry and geometry as part of the information budget in coarse-grained models.
Yann L. Muller, Claire A. Paetsch, Anirudh Raju Natarajan
pyeCE implements embedded cluster expansion, using machine learning to map many chemical species onto a smaller set of effective species. The authors report that a model spanning the full composition space of a nine-component refractory alloy resolves short-range order and order-disorder behavior. This extends a familiar thermodynamic modeling framework toward alloy spaces whose chemical complexity normally makes conventional cluster expansion impractical.
Seán R. Kavanagh, Chuin Wei Tan, Menghang Wang, Marc L. Descoteaux, Gabriel de Miranda Nascimento, Ulrik Unneberg, Laura Zichi, Francesco Libbi, Norma Rivano, Austin Glover, Vivek Bharadwaj, Anders Johansson, William C. Witt, Albert Musaelian, Boris Kozinsky
The work develops a family of equivariant foundation potentials in the NequIP and Allegro architectures, together with infrastructure for training on ultra-large datasets. The authors report leading inference speed, strong scaling, and high accuracy across materials discovery, thermal conductivity, and near-equilibrium mechanical and thermodynamic benchmarks. The analysis shifts attention from architecture alone toward the diversity and consistency of the energy surfaces used to train universal potentials.
MIRAGE benchmarks affinity and pose models across an explicit axis of historical protein-family support, using matched strata, family-disjoint controls, and temporal evaluation. The authors report that rankings reverse on novel families, where a family-disjoint random forest leads both co-folders and significantly outperforms Nesso-1. This makes training-set redundancy a measurable part of the claim of generalization, rather than a hidden feature of a pooled benchmark.
Michael Hanna, Julian Cremer, Zekiye Erarslan, Leonardo Medrano Sandonas
QALPA combines an E(3)-equivariant diffusion model, active learning, and efficient quantum-mechanical methods to explore targeted property manifolds for flexible molecules. The authors report that coupling QALPA to EquiDTB augmented alloQM, a dataset of 6,253 allosteric-drug conformers, by filling sparse regions defined by dispersion energy and the HOMO-LUMO gap. This gives generative sampling a physics-based route to extend sparse quantum datasets without treating larger flexible molecules as a separate regime.
Bilvin Varughese, Aditya Koneru, Adil Muhammad, Troy D. Loeffler, Sukriti Manna, Jan Michael Y. Carrillo, Orcun Yildiz, Thomas Peterka, Subramanian K.R.S. Sankaranarayanan
The authors train equation-learner neural networks on density-functional data and combine their symbolic forms in a weighted ensemble potential for aluminum. The authors report that all symbolic models reach sub-10 meV per atom accuracy and that their ensemble improves consistency with density-functional benchmarks for phonons, equation-of-state curvature, and melting dynamics. The result keeps a compact analytical form while using diversity among learned models to strengthen prediction beyond equilibrium structures.
NQC-PACE relabels configurations from an iron-hydrogen machine-learning potential with finite-temperature quantum mean forces. The authors report stronger hydrogen segregation at general grain boundaries and trapping behavior closer to experimental trends in Monte Carlo and molecular-dynamics simulations. This matters because quantum effects for light solutes can be incorporated into large-scale defect simulations without new density-functional calculations.
The paper defines an action-operator framework for molecular diffusion models and uses a noisy operator bridge to read out free-energy differences from endpoint ensembles. The authors report that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without force or action supervision. This provides a route for turning diffusion models from coordinate samplers into thermodynamic estimators with an explicit physical interpretation.
PANO and CANO map full-field displacement measurements and net reaction forces directly to hyperelastic strain-energy density functions, using Laplacian eigenfunctions and physically admissible output constraints. The authors report near-instantaneous constitutive inference in one forward pass and evaluate the operators on unseen, noisy, incomplete, differently discretized data and geometries of different sizes. The operator formulation makes constitutive discovery fast, discretization-independent, and constrained to physically admissible material responses.
Mikołaj J. Gawkowski, Nongnuch Artrith, Silvia Bonfanti, Abhijeet Sadashiv Gangan, Hendrik H. Heenen, Joseph Kioseoglou, Ivor Lončarić, Hemanadhan Myneni, Janosh Riebesell, Mariana Rossi, Matthias Rupp, Jonathan Schmidt, Shubham Sharma, Benjamin X. Shi, Antoni Wadowski, Lukas Hörmann, Venkat Kapil
Dyna-Mat evaluates 15 foundation interatomic potentials against finite-temperature first-principles trajectories, comparing static energy and force errors with structural and dynamical observables from model-driven simulations. The authors report that low force error usually tracks better ensemble behavior but can still coincide with qualitative structural failure, while pressure remains poorly described across most models. The benchmark makes trajectory-level physical behavior, rather than a favorable static error alone, the standard for judging a deployable potential.
Stefan Gugler, Max Eissler, Khaled Kahouli, Klaus-Robert Müller
ReactionAtlas builds reaction networks from a handful of seed molecules without hand-crafted rules, using a generative model to propose reactions and a DFT-trained machine-learned force field to filter valid transition states. The authors report roughly 47,000 reactions among roughly 12,000 compounds from eight pre-biotic seeds, with 85% of machine-learned transition states within 0.5 Å RMSD of PBE0 references. This combination makes deep network exploration more tractable when the relevant products and transition states are not known in advance.
Jacob W. Toney, Ayleen Y. Farnood, Samir Darouich, Heather J. Kulik
The authors construct fixed-length molecular fingerprints from spectral graph theory using three-dimensional interaction-weighted graphs. The authors report strong benchmark performance across organic, inorganic, biological, reticular, and reaction-chemistry datasets while distinguishing structures with identical two-dimensional connectivity. This matters because an interpretable representation of three-dimensional similarity could scale beyond pairwise descriptors without depending on learned embedding coverage.
CGMas coordinates agents for polymer topology construction, equilibration, coarse-graining, potential derivation, and validation from a natural-language specification. The authors report completion of all 27 polymer tasks, density agreement within five percent in 22 cases, and a reduction in simulation time to one minute. This matters because automated coarse-graining could make polymer models easier to build for mappings that otherwise require bespoke work.
Weichi Yao, Cameron Gruich, Bryan R. Goldsmith, Yixin Wang
The authors introduce a two-stage generative framework that relies on a fixed-dimensional molecule-level latent representation to generate variable-size molecules. On PCQM4Mv2, the authors report the highest fraction of outputs that were unique, novel, sanitized, and passed PoseBusters checks. The architecture makes molecular size a generated consequence of the representation, which matters for open-ended property-directed design.
The authors introduce AdaptNTK, which measures uncertainty as a regularized Mahalanobis distance in empirical neural tangent kernel feature space. On held-out rMD17 data, they report force-error correlations of 0.68 and 0.71 and a 2.6-fold speedup per Transition-1X cycle. Sequential uncertainty updates offer a route to less redundant acquisition without retraining after each selection.
The authors develop excited-state machine-learning molecular dynamics calibrated against real-time time-dependent density functional theory benchmarks. They report that large-scale simulations resolve phonon competition in bismuth phase transition and structural rearrangement in selenium photoamorphization. This extends learned dynamics toward complex photoexcited materials where electron-nuclear evolution matters.
The method couples a grain graph neural-network surrogate for stress prediction with conditional reverse diffusion over alpha-phase fraction, elastic modulus, and yield-stress targets. Across four target regimes, the authors report that finite-element checks of the five best candidates reached a maximum absolute relative error of 1.0%, while diffusion used 32 evaluations per input graph rather than approximately 40,000 for the search baselines. Generating microstructures directly at prescribed properties reduces the search burden while retaining topology and crystallographic checks for physical consistency.
The autonomous agent selects active spaces, submits ORCA calculations, analyzes outputs, and records its choices in an auditable reasoning log. With a structured decision ladder, the authors report increased coverage of vertical transition energies alongside a lower mean absolute error. It clarifies that reliable autonomy in multireference chemistry depends on making expert workflow decisions explicit and checkable.
Yiting Zheng, Cheng Fang, Anthony Donofrio, Haote Li
RxnCLF pretrains a reaction foundation model on condensed graphs that unify reactant and product information into an explicit transformation representation. The authors report improved yield-prediction performance over graph and sequence baselines across several reaction benchmarks. This matters because reaction representations that retain both centers and surrounding context may transfer to a wider range of reaction-informatics tasks.
Danish Khan, Maurice D. Hanisch, Nikolai Argatoff, Evan Xie, Sandeep Sharma, Anima Anandkumar
A domain-invariant SE(3)-equivariant Fourier neural operator learns the Kohn-Sham map from potential to electron density on real-space grids, enabling self-consistent-field calculations without explicit orbital construction. The authors report that one model trained on 8,504 molecules and solids generalized to out-of-distribution molecules, insulators, and metals, and converged a magnesium dislocation calculation with 82,500 valence electrons on one GPU. Learning the map rather than an ill-conditioned functional offers a path toward orbital-free calculations that retain Kohn-Sham-level observables at larger scales.
Juniper frames compositional backmapping as discrete denoising diffusion over molecular graphs conditioned on octanol-water partition free energy. The authors report valid and unique molecules for two-bead targets whose partition-free-energy distributions track the coarse-grained targets linearly. The method turns a lossy coarse-grained screening result back into candidate molecules for atomistic study or synthesis.
Sergey A. Shteingolts, Salman N. Salman, Ron Levie, Dan Mendels
This inverse-design framework uses a graph neural network molecular-dynamics simulator, a short dynamical initialization, and physics-based refinement during simulation. A simulator trained on non-auxetic systems designed strongly auxetic networks, and the framework generalizes across system size. The combination of learned dynamics and physical refinement offers a way to optimize mechanical response beyond the distribution that supplied the training data.
GALAXI assigns each candidate crystalline phase an independent binary classifier, then uses Rietveld refinement to choose the combination explaining the full pattern. On curated experimental patterns, the authors report correct phase identification and robustness to common diffraction artifacts. Independent phase models make a large reference library an incremental engineering problem rather than a reason to retrain a monolithic classifier.
Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao
The authors introduce Hyper-Fold, a rank-K separable convolutional backbone that organizes each radius neighborhood into sequence and contact hyperedges. Across enzyme function, fold classification, and binding-site tasks, they report leading protein-encoder results and lower parameter count and latency for Hyper-Fold-Pocket. The work argues that expressive content-geometry interactions can recover information often attributed to large evolutionary pretraining.
SE(3)-MeanFlow extends MeanFlow to protein-frame Lie-group geometry and derives closed-form average-velocity targets for simulation-free training. The authors report matched or better backbone generation than flow-matching baselines using several times fewer sampling steps, with the largest gains in the few-step regime. Few-step generation directly targets the network-evaluation bottleneck in high-throughput protein design while making the associated diversity tradeoff explicit.
The authors introduce a Fourier Neural Operator crystal-field solver that maps chemical formulae and lattice parameters to periodic density fields. They report novel structures across 104 chemical formulae with competitive reconstruction accuracy, generative diversity, and structural validity. This couples composition-conditioned generation to a reconstruction and screening path for crystal discovery.
Jia-Wen Li, Sheng Meng, Xinghua Shi, Jin Zhang,, Wei-Hai Fang
EFR-GNN predicts energies, forces, Born effective charge tensors, atom-resolved charges and magnetic moments, and supports long-time field-driven molecular dynamics that track localized electronic states. In hole-doped MgO, GaAs, and superionic alpha-AgI, the authors report field-driven polaron, phonon, and ionic dynamics that agree with a nearest-neighbor model, experiment, and temperature-dependent mobility. A single field-aware potential can therefore connect atomistic motion, electronic response, and localized-carrier dynamics over the long trajectories needed for finite-temperature materials simulations.
Liang Shuang, Haocheng Wang, Jiayi Song, Shuquan Ye, Ben Fei
ED-DiT learns transferable molecular representations by reconstructing corrupted and masked electron-density fields while using an electron-number consistency constraint to preserve total electronic mass. The authors report that electron-density prediction RMSE fell from 2.2474 to 1.3753 and exceeded the available baseline. Pretraining on a physical field gives one encoder access to molecular properties, spin state, retrieval, and density generation rather than a single downstream label.
PolyLatentFlow uses continuous-time flow matching in latent space for conditional polymer generation and pairs it with the sequence-and-structure representation LlamaUni. For carbon-dioxide and nitrogen conditioning, the authors report that PolyLatentFlow with LlamaUni achieved the largest per-attempt yield of nonreplayed target hits among the evaluated representations. The comparison shows that representation determines how a polymer generator balances target control against exploration beyond its labeled chemistry.