Ran Elgedawy, Sanjay Das, Ethan Seefried, Ryan Burchfield, Gavin Wiggins, Kimberly Jeskie, Chrissi Schnell, Susan Fiscor, Subhamay Pramanik, Billey Thomas, Sudarshan Srinivasan, Tirthankar Ghosal
SafeChem combines a regulatory-grounded dataset of 32,211 substances and 30 hazard labels with separate tests of structure-based prediction and LLM safety reliability. The authors report macro-AUPRC values of 0.45–0.55 for high-prevalence labels and omission-hallucination rates above 0.21 for all eight LLMs in a 500-substance stress test. The benchmark supplies a common test bed for molecular models and general-purpose LLMs operating in chemical-safety workflows.
Multi4D combines a latent-space Diffusion Transformer with a rotation-invariant convolutional network for diffraction datasets. The authors report high classification accuracy and apply the method to superconducting heterostructures, corroded alloy surfaces, and degraded solid-state battery interfaces at single-nanometer resolution. It joins automated diffraction interpretation to a local measure of structural ambiguity at interfaces where degradation begins.
RiemannMol adds an invertible, isometry-regularized metric head to a frozen molecular latent space so that distance tracks a chosen chemical property. The authors report more efficient property-guided sampling and optimization than decode-then-filter, with behavior validated against exact LogP ground truth. Learning a chemically purposeful latent metric gives molecular generators a clearer geometry for controllable sampling and optimization.
The study trains a mixture-of-experts reranker under mixed spectroscopic conditions, including modality-specific perturbations and chemically informed spectrum replacements. Across a large test set spanning predefined conditions, the authors report that mixed-condition training raised the mixture-of-experts mean reciprocal rank. The contribution is a more realistic definition of robustness for structure identification when the available spectra are incomplete or inconsistent.
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
The authors train C3LM on roughly 45.6 million verified reactions and use Top-K prompting with ChemCensor-based and novelty-oriented rewards to generate diverse retrosynthetic predictions. They report state-of-the-art performance on the URSA-expert-2026 benchmark and complementary reaction-space exploration by language and conventional models. Plausibility-aware multiple predictions better reflect the one-to-many character of retrosynthesis than a single-answer target.
Dahvyd Wing, Mihail Bogojeski, Szabolcs Goger, Klaus-Robert Muller, Alexandre Tkatchenko
DensIP uses machine-learned electron densities and four universal parameters to model intermolecular interactions, trained and tested on CCSD(T)/CBS dimer energies. The authors report sub-kcal/mol errors for dimers containing molecules absent from training, including non-equilibrium conformations, and better long-range-interaction performance than general-purpose machine-learned force fields. Accurate, inexpensive synthetic reference data could ease the ab initio-data bottleneck in training more general force fields.
Mohamed S. Abdallah, Zhuotao Jin, Boris Kozinsky, Kyle Bystrom
CIDER26SS combines machine learning with explicitly nonlocal, physically informed descriptors in a transferability-regularized exchange-correlation functional. The authors report that it resolves the CO/Pt(111) binding-site puzzle with accurate adsorption, lattice, and surface-energy predictions, including when Pt bulk and surface data are excluded from training. The result points to a functional-design route for heterogeneous catalysis that is not confined to systems represented in its training data.
Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow
The authors train molecular-design language models through a curriculum of synthetic tasks that gradually becomes more challenging. They report that this curriculum surpassed much larger frontier models on structure-based lead optimization. Synthetic tasks could let post-training teach useful search strategies when direct chemical scoring is too costly for online optimization.
TDAP-eML couples machine-learned electronic structure with atomistic propagation to simulate photoexcitation-induced lattice dynamics. The authors report reproduction of key photoexcited lattice responses in silicon and FeSe, with nearly three orders of magnitude lower cost for the large systems examined. The framework connects nonequilibrium electronic evolution to extended structural dynamics at a computational scale useful for experimentally accessible observables.
TSBench asks an LLM agent to use structure-editing tools to build three-dimensional transition-state guesses checked by an automated quantum-chemical pipeline. Across 546 evaluations, the authors report that diagnosis-driven revision raised the aggregate success rate. The benchmark ties mechanistic language to a reaction path, giving claims about chemical reasoning a physical pass-or-fail test.
The authors approximate the finite-time transition kernel of a stochastic reaction-network Markov chain with a conditional normalizing flow. Their numerical examples report statistically consistent coarse-step trajectories with reduced computational cost. A learned stochastic propagator can decouple useful simulation steps from microscopic reaction-event resolution.
Shixu Liu, Xingding Li, Haozhe Li, Yang Zhong, Hongjun Xiang, Xin-Gao Gong, Ji-Hui Yang
The authors combine a first-principles electron-magnon formalism with machine-learned spinful Hamiltonians to calculate transport effects in collinear magnetic systems. They report recovery of the full T2 resistivity component in ferromagnetic α-Fe with a coefficient agreeing quantitatively with measurement, and an ARPES-observed magnon kink in K-doped BaMn2As2. The framework supplies a route to assess electron-magnon contributions without reducing magnetic transport to electron-phonon effects alone.
Wei-Jian Jiang, Ye-Nan Sha, Hui Guo, Jie Chen, Yu-Cai Liang, Ke Zhou, Qi-Long Gao, Dong-Lin Han, Xin-Gao Gong, Wan-Jian Yin
The authors develop PIRAG-LM, which retrieves synthesis precedents by chemical, structural, and thermodynamic similarity before structured route reasoning. They report 91.4% synthesis-method prediction accuracy, compared with 72.1% for the language model alone, and experimental synthesis of five new compounds. Retrieval gives the planning system an interpretable path from materials discovery to experimental realization.
Changquan Zhao, Yuxiang Sun, Ruihao Zhu, Cheng Hua, Yulian He
DASH decouples surrogate selection from acquisition control by scoring surrogates for predictive reliability, uncertainty calibration, and ranking consistency, then reallocating acquisition-function quotas before an LLM chooses from the resulting shortlist. Across four chemical optimization tasks, the authors report 12.51% better trajectory-level Acceleration Factor and 5.00% better endpoint Enhancement Factor than the strongest AutoBO baseline, with ablations supporting complementary contributions from all components. The separation gives automated optimization a way to adapt model reliability and search behavior independently as campaign feedback accumulates.
Jaehwan Choi, Kunik Jang, Seongmin Kim, Shuan Chen, Kyungju Nam, Seung Hyo Noh, Donghwi Kim, Yousung Jung
The authors introduce CRISP, which samples chemical rules, consolidates them, and compiles each into an executable scalar descriptor for a conventional learner. For inorganic-crystal synthesizability, they report stronger results than expert-curated and generic representations under a shared learner, especially under structural-size and chemical-family shifts. The method makes chemical heuristics auditable computational features without relying on structures, labels, or data splits during rule construction.
Tianyu Gao, Zhikai Su, Jiashu Li, Wenjun Gao, Zichuan Ying, Zhe Zhao, Fei Zhang, Ye Wei
The authors use a language-informed flow-matching framework in which target-aware SMILES supply semantic priors to geometric molecular generation. On Cross-Docked2020, they report competitive distribution matching with improved medicinal chemistry metrics and competitive structural validity under task steering. The approach tests whether language-derived chemical context can guide three-dimensional design without further generator fine-tuning.
FS-TPS and LaS-TPS introduce stochastic control policies for molecular transition-path sampling during explicit molecular-dynamics rollouts. The authors report better transition success and path quality than deterministic-policy baselines across three biomolecular systems, with lower sensitivity to initialization. This matters because stochasticity can improve rare-event sampling without abandoning the physical trajectory dynamics.
Judith Bernett, Anton Spannagl, Joel \AAs, Markus List, David B. Blumenthal
The authors audit protein-protein interaction datasets with similarity-aware splitting and bias-minimizing negative sampling, both posed as integer linear programs. They report that removing train-test protein overlap still leaves shortcuts from self-interactions, taxonomic identity, and functional relatedness, with prevalence that depends on the data source. The study supplies a general way to quantify and reduce dataset shortcuts before a model’s apparent biological signal is taken at face value.
The authors introduce S3C-LLM, which retrieves spectroscopy skills, runs analysis code, and integrates peak-level evidence before generating SMILES. Across diverse benchmarks, they report that S3C-LLM outperforms general and spectrum-specific models while using less than 1/10th of SpectraLLM training data. The workflow places analytical constraints inside spectrum-to-structure prediction rather than treating it as direct generation.
Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr., Joseph A. Morrone
GRACE adapts a pretrained geometry encoder through residual learning against an adduct-aware physical baseline and encoder-level adduct conditioning. Across random, scaffold, and adduct-sensitive splits, the authors report the best mean percentage difference among the evaluated learned models. It identifies residual learning and early adduct integration as concrete design choices for collision-cross-section prediction that must generalize beyond familiar scaffolds.
Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Alan Aspuru-Guzik
El Agente Potente combines typed execution graphs for standardized interatomic-potential workflows with a coding agent for procedures that require more flexibility. The authors demonstrate the system across materials discovery, molecular energy landscapes, adsorption, and catalytic reaction workflows, alongside reproducibility and token-cost benchmarks. The division of planning from deterministic computation makes agentic atomistic simulation easier to audit while leaving room for custom workflows.
Ross Irwin, Alessandro Tibo, Jon Paul Janet, Simon Olsson
The authors introduce ensemble-conditioned guidance, jointly conditioning a 3D molecular generator on molecular modes and ensemble properties. The authors report that, in dual-target binder and active-state-selective agonist design, adding the additional state improved the desired outcome over single-state conditioning. That makes the conformational ensemble, rather than an isolated conformer, an explicit variable in molecular design.
Longfei Lv, Li Fu, Si Zhou, Lingzhi Zhao, Jijun Zhao
The authors couple a structure generator, property predictor, and multi-criteria validation workflow for inverse design of singlet-fission molecules. The authors report a high success rate for candidates after evaluating a large generated set and validating a random subset with time-dependent density functional theory. The study offers a route through a sparse excited-state search space while retaining a quantum-chemical check on generated candidates.
Siyuan Ma, Yi Chai, Yi Wu, Qixin Zhang, Yajing Yuan, Kanglu Zhao, Zhikang Chen, Haowei Wang, Shuying Cao, Xiaolei Yu, Xiangfei Han, Yun Liu, Yang Liu, Tingting Zhu, Dacheng Tao
The authors couple RegimeFormer with RegimeAtlas to model protein perturbation regimes across a harmonized global sequence collection. Across deep mutational scanning and molecular benchmarks, they report reproducible regimes and improved substitution-specific prediction under unseen-protein, unseen-family, and low-homology evaluation. The framework offers a scalable way to map and query mutation response across protein sequence space.
Isaac Malsky, Xi Zhang, Tiffany Kataria, Matthew Graham, Ziyu Huang, Boris Bonev, Shang-Min Tsai,, Elspeth K.H. Lee
The authors develop a residual flow-map neural solver for local-box chemical kinetics in exoplanet atmospheres. They report microsecond-scale inference, percent-level accuracy, and robustness under the extreme stiffness of atmospheric chemistry. A fast surrogate could let atmospheric models retain kinetic effects that equilibrium approximations miss.
The authors use a symmetric dual-path architecture that iteratively combines protein language and multimodal protein language knowledge for sequence generation. Across standard inverse-folding benchmarks, they report state-of-the-art performance and ablations supporting the symmetric design. The design directly addresses the limits of post-hoc sequence refinement in inverse folding.
The authors introduce TFMat, which conditions a CrystalFlow generator on structured materials language as a semantic prior. Across the stated crystal benchmarks, they report a 92.04% MP-20 match rate with 20 candidates and improved alignment in de novo generation. The result makes materials language an inspectable control layer before downstream simulation and validation.
The authors combine density matrix renormalization group calculations as a complete active-space solver with machine learning to evaluate electronic excitations. They demonstrate transferability on challenging polycyclic aromatic examples containing up to 34 pi electrons. This supplies a practical route to excited-state estimates for correlated materials that are expensive for conventional many-body calculations.
Marco G. Barnfield, Sergei N. Yurchenko, Jonathan Tennyson
The pipeline trains a transductive GraphSAGE network on empirical energy levels to propagate assignment information to unlabelled calculated states. The authors report assignments for previously unlabelled states across carbon-dioxide isotopologues, increasing coverage at low energy. The workflow points to automated annotation of large line lists when symmetry-like constraints can be stated directly in the assignment problem.
HiMatGen organizes language-model agents into discussion pods and domain representatives that exchange evidence and unresolved questions during crystal design. Across six chemical systems, the authors report more final crystal candidates and more stable, unique, and novel structures than a single-agent baseline. Its value lies in carrying scientific disagreement and computational evidence through an iterative materials-design process.
Boltz-Perturb perturbs model-conditioning signals during inference through Token Bias Perturbation and Token Conditioning Perturbation, increasing exploration of alternative protein-ligand binding poses. Across diverse protein-ligand systems, the authors report that token conditioning improved top-20 oracle success rates by 2.6- to 7.8-fold, while Boltz-Perturb required over 75% less compute than the Boltz-2 high diffusion temperature variant. The method makes latent binding-mode diversity accessible without retraining, so a sampling deficiency can be addressed directly during co-folding inference.
Rishabh Arora, Lisa Scheunemann, Tim Brepols, Shahed Rezaei
The framework learns a material operator from full strain histories to stress trajectories, with causally masked attention restricting access to past states. Across rate-independent material models, the authors report accurate and robust predictions of irreversible deformation with resolution invariance and parallel efficiency. A full-path operator offers a direct representation of material history when the internal variables needed by classical constitutive models are unavailable.
Dasol Yoon, Poompol Buathong, Chia-Hao Lee, Yujia Zhang, David A. Muller, Peter I. Frazier
SBOCF exploits the composite image-matching objective and intermediate simulated-image information for inverse problems in electron microscopy. Using patch-level summaries and correction terms, the authors preserve the pixel-wise objective while reducing the number of modeled outputs. The approach shows how image structure can reduce simulator demands without substituting a broad pretrained network for the scientific objective.
Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel Angel Bautista
SimpleDesign trains sequence and structure directly in data space with a single end-to-end objective combining sequence cross-entropy and structural regression. Trained on over 2M sequence-structure pairs, the model achieved strong performance across co-design and unconditional generation benchmarks, the authors report. The result argues that protein co-design need not pass through separately trained latent representations to obtain broad generative performance.
The authors introduce isotropic and normal-mode-weighted displacement schemes that derive conformation augmentation from Hessian information through Taylor expansions. Across equilibrium and non-equilibrium datasets, they report improved accuracy where the reference forces are large. The augmentation targets force-rich regimes without changing a potential architecture or its training objective, making Hessian-aware training more portable.
Christopher Tosh Antoine de Mathelin, Wesley Tansey
ScreenShot is a hierarchical transformer pretrained on drug-screening datasets that predicts combination-therapy responses from few-shot functional observations without fine-tuning or molecular profiling. The authors report that it outperformed all baselines on four held-out datasets and matched uniform-screening hit detection with a reduced budget through active learning. This approach could prioritize combination experiments when molecular profiling or cohort-specific model training is impractical.
Roxane Axel Jacob, Daniel Rose, Thierry Langer, Johannes Kirchmair
NEAT-POCKET generates molecules atom by atom in protein-pocket environments while preserving atom permutation invariance and modeling hydrogens explicitly. On CrossDocked and SPINDR, the authors report competitive structure-based generation with substantially faster sampling than existing baselines. A pocket-conditioned generator that also completes fragments brings the same representation close to lead optimization and scaffold elaboration.
MolWeaver uses a frozen language-model backbone, explicit constraint grounding, graph-level control, and scaffold-aware decoding to turn natural-language prompts into valid molecular graphs. The authors report evaluations on ZINC22-derived prompts across ring-size, Lipinski-guided, and phosphine-ligand generation tasks with specified chemical constraints. The framework makes qualitative chemical intent directly usable for molecular generation, reducing reliance on structured numerical or categorical inputs.
Peng Kang, Da Wan, Shulin Bai, Vincent Michaud-Rioux, Zhen Li, Yu Liu, Lei Zheng, Li-Dong Zhao
The study defines a directional residual-work coefficient that combines curvature mismatch with the spatial distribution of residual force response. At a lithium-electrolyte interface, predictions fixed before future reference evaluations differ from measurements of predicted work. By resolving force error along motion, the framework gives potential assessment and adaptive reference allocation an energetic basis.
Raghutheja Bollampally, Soumya Sankar, Yuqi Qin, Berthold Jack
The authors combine a random-forest surrogate with thermodynamic constraints and expected improvement in a sequential model-based active-learning loop for molecular beam epitaxy. They report that, after a small initial training set, four active-learning iterations halved the absolute predictive error to about 10% for Fe3Sn growth. The framework targets abrupt phase boundaries and narrow growth windows that make continuous optimization models a poor fit for closed-loop thin-film synthesis.
The authors introduce GRAS, which lowers guided-proposal variance and uses an adaptive resampling temperature for training-free discrete diffusion steering. Across regulatory DNA and protein design, they report the best training-free reward and performance that matches or exceeds a reward-fine-tuned model. The analysis isolates a compact change to inference-time search that remains effective for non-differentiable rewards.
The authors introduce GITIII-scale, a hierarchical interpretable spatial transcriptomics model that decomposes cell state-niche associations and ligand-receptor pathways. On cancer types unseen in training, they report embeddings that recover niche-associated state changes more accurately than existing spatial transcriptomics foundation models. The model ties pan-cancer representation learning to biological mechanisms that can be inspected for drug-target hypotheses.
Manasa Kaniselvan, Mauro Dossena, Denghui Lu, Alexander Maeder, Nicolas Vetsch, Alexandros Nikolaos Ziogas,, Mathieu Luisier
The authors integrate machine-learned electronic structures with a quantum transport solver for semiconductor device simulation. They report 10,000X speedups over density functional theory for devices above 20,000 atoms and identify an effect of undercoordinated Hf or Al atoms on MoS2 current. Explicit oxide layers can therefore enter transport calculations at device-relevant scales.
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
The study derives a group-theoretical parity-gap criterion for determining when a model's feature parity labels allow exact symmetry-forced zero predictions. The authors report that parity-labelled architectures stayed at the floating-point floor on centrosymmetric crystals while rotation-only models predicted forbidden piezoelectric responses on 90-96% of cases. Exposing this design choice makes exact physical constraints testable before costly training and benchmarking.
ReCurveflow learns transition-state geometries from curved reference paths derived from full NEB bands, with off-path correction during inference rollout. The authors report best results on most split-metric combinations against seven baselines and trajectories whose energy profiles closely track the reference NEB paths. Curved-path supervision and corrective fields address a mismatch between idealized straight interpolation and reaction trajectories.
Ethan Ma, Zihan Wang, Chris Siu Yeung Chow, Xinguo Feng, Qingqing Li, Rui Jiang, Naipeng Dong, Guangdong Bai
Embedded Graph Flows learns continuous node and unordered-edge embeddings, then transports Gaussian noise toward them with a permutation-equivariant graph transformer. On QM9, the authors report the best result among the compared methods on every reported metric. Learning category geometry instead of fixing one-hot distances gives graph generation a representation that can respect order-invariant molecular structure.
Ryuhei Okuno, Nontawat Charoenphakdee, Kaoru Hisama, Yuta Tsuboi
MLIP Detective turns benchmark evidence into falsifiable physics-informed failure hypotheses, then screens inexpensive simulations before escalation to experts. Without issue-specific prompting, the authors report a systematic anomaly for some oxygen- or fluorine-containing adsorbates on surfaces. That makes evaluation an active search for consequential failures, rather than a scorecard limited to the configurations already in a benchmark.
Agent-MD reserves language-model reasoning for campaign setup and event-triggered review while a persistent rule-based agent manages routine molecular-simulation operations. The authors report 120 segmented simulation cycles across 15 system-humidity states without live reasoning during production, with replay identifying workflow problems at review boundaries. This separation could make long-running simulations more auditable without turning every routine control action into a model decision.
MBTA maintains modality-specific latent spaces and connects them via flow matching, preserving the structure of each modality while enabling cross-modal translation. Across multimodal single-cell benchmarks, the authors report that MBTA outperformed existing methods most strongly where structural mismatch was pronounced and identified transcriptomic lineage relationships corroborated by genomic variation in breast cancer profiles. By linking multiple molecular readouts without erasing their individual structure, the framework makes those structural differences usable in layered descriptions of cellular identity.
ChemDIRT evaluates chemical reasoning in large language models by varying instructions and molecular representations across eight task categories, measuring both accuracy and consistency. The authors report substantial prompt sensitivity, representation dependence, and uneven performance across task families among benchmarked open- and closed-source models. The framework helps distinguish robust chemical reasoning from performance that depends on a favorable problem formulation.