Preprint scanner project / Issue 2026-08-19

Useful Chemistry

A weekly review of arXiv, ChemRxiv, and bioRxiv for ai-enabled methodological advances. Up to twenty papers reviewed every week; a quiet week may publish none.

20 selected / 114 reviewed Aug 12, 2026–Aug 18, 2026
Groups
Group roll call
96 followed groups Research-group roll call
  • Milad Abolhasani Abolhasani Lab North Carolina State University Site ↗
  • Surl-Hee (Shirley) Ahn Ahn Lab University of California, Davis Site ↗
  • Mohammed AlQuraishi AlQuraishi Laboratory Columbia University Site ↗
  • Sayan Banerjee Banerjee Group University of Tennessee, Knoxville Site ↗
  • Christopher J. Bartel Bartel Research Group University of Minnesota Twin Cities Site ↗
  • Jan-Niklas Boyn Boyn Research Group University of Minnesota Twin Cities Site ↗
  • Ting Cao Cao Group University of Washington Site ↗
  • Timothy Cernak Cernak Lab University of Michigan Site ↗
  • Ming Chen Ming Chen Group Purdue University Site ↗
  • Po-Yen Chen Chen Research Group University of Maryland, College Park Site ↗
  • Bingqing Cheng Cheng Group University of California, Berkeley Site ↗
  • Brian Cleary Algorithmic Lens on Experimental Biology Laboratory Boston University Site ↗
  • Connor W. Coley Coley Research Group Massachusetts Institute of Technology Site ↗
  • Yamil J. Colón Colón Group University of Notre Dame Site ↗
  • César de la Fuente Machine Biology Group University of Pennsylvania Site ↗
  • Julia Dshemuchadse Cape Crystal Cornell University Site ↗
  • Thomas P. Fay Fay Group University of California, Los Angeles Site ↗
  • Kara D. Fong Fong Lab California Institute of Technology Site ↗
  • Thomas E. Gartner III Gartner Group Lehigh University Site ↗
  • Rafael Gómez-Bombarelli Learning Matter Massachusetts Institute of Technology Site ↗
  • Jonathan Gootenberg Gootenberg Laboratory Harvard University Site ↗
  • Prashun Gorai 3D Materials Lab Rensselaer Polytechnic Institute Site ↗
  • Diptarka Hait Hait Group Columbia University Site ↗
  • Brian Hie Laboratory of Evolutionary Design Stanford University Site ↗
  • Guoxiang Hu Hu Group Georgia Institute of Technology Site ↗
  • Yong-Jie Hu Materials Computation and Informatics Group Drexel University Site ↗
  • Yifei Huang Huang Lab Pennsylvania State University Site ↗
  • Yunha Hwang Hwang Lab Massachusetts Institute of Technology Site ↗
  • Nicholas E. Jackson AI for Materials Group University of Illinois Urbana-Champaign Site ↗
  • William M. Jacobs Jacobs Group Princeton University Site ↗
  • Dipti Jasrasaria Jasrasaria Group University of Chicago Site ↗
  • Anupama Jha Jha Lab Yale University Site ↗
  • Kaiyi Jiang Jiang Lab Princeton University Site ↗
  • Adrian Jinich Jinich Lab University of California, San Diego Site ↗
  • Felipe Jornada Jornada Research Group Stanford University Site ↗
  • Kalli Kappel Kappel Lab University of California, Los Angeles Site ↗
  • Joshua Kretchmer Kretchmer Research Group Georgia Institute of Technology Site ↗
  • Aditi S. Krishnapriyan Krishnapriyan Research Group University of California, Berkeley Site ↗
  • Sebastian Kube Kube Lab University of Wisconsin–Madison Site ↗
  • Heather J. Kulik Kulik Research Group Massachusetts Institute of Technology Site ↗
  • Ambarish R. Kulkarni Kulkarni Research Group University of California, Davis Site ↗
  • Joseph S. Kwon Kwon Research Group The Ohio State University Site ↗
  • Joonho Lee Lee Group Harvard University Site ↗
  • Can Li Li Research Group Purdue University Site ↗
  • Wanlu Li Wanlu Li Research Group University of California, San Diego Site ↗
  • Rebecca K. Lindsey Lindsey Lab University of Michigan Site ↗
  • Ge Liu Ge Liu Group University of Illinois Urbana-Champaign Site ↗
  • Yuanyue Liu Yuanyue Liu Group The University of Texas at Austin Site ↗
  • Yang Lu Lu Lab University of Wisconsin–Madison Site ↗
  • Jiankun Lyu Evnin Family Laboratory of Computational Molecular Discovery The Rockefeller University Site ↗
  • Cong Ma Cong Ma Lab University of Michigan Site ↗
  • Arkajit Mandal Mandal Group Texas A&M University Site ↗
  • Andrew J. Medford Medford Research Group Georgia Institute of Technology Site ↗
  • Ilias Mitrai Systems and AI Lab The University of Texas at Austin Site ↗
  • Jeetain Mittal Mittal Group Texas A&M University Site ↗
  • Joel A. Paulson Paulson Lab University of Wisconsin–Madison Site ↗
  • Elisa Pieri Pieri Lab University of North Carolina at Chapel Hill Site ↗
  • Doran Raccah MesoScience Lab The University of Texas at Austin Site ↗
  • Phillip Rauscher Rauscher Group New York University Site ↗
  • Wesley Reinhart Reinhart Group Pennsylvania State University Site ↗
  • Gabriel J. Rocklin Rocklin Lab Northwestern University Site ↗
  • Andrew S. Rosen Rosen Research Group Princeton University Site ↗
  • Grant M. Rotskoff Rotskoff Group Stanford University Site ↗
  • Janani Sampath Sampath Research Group University of Florida Site ↗
  • Elvira Sayfutyarova Sayfutyarova Group Pennsylvania State University Site ↗
  • Martin Seifrid Seifrid Group North Carolina State University Site ↗
  • Thomas P. Senftle Senftle Group Rice University Site ↗
  • Karthik Shekhar Shekhar Lab University of California, Berkeley Site ↗
  • Zachary M. Sherman Z Lab University of Washington Site ↗
  • Krishna Shrinivas Shrinivas Lab Northwestern University Site ↗
  • Rohit Singh Singh Lab Duke University Site ↗
  • Micheline Soley Soley Group University of Wisconsin–Madison Site ↗
  • Kevin V. Solomon Solomon Laboratory University of Delaware Site ↗
  • Jeff Spence Spence Lab University of California, San Francisco Site ↗
  • Kayla G. Sprenger Rational Design of Interfaces Lab University of Colorado Boulder Site ↗
  • Chong Sun Sun Lab Rutgers University–New Brunswick Site ↗
  • Yidan Sun Sun Lab Washington University in St. Louis Site ↗
  • Daniel Tabor Tabor Research Group Texas A&M University Site ↗
  • Ming Tang Mesoscale Materials Science Group Rice University Site ↗
  • Roel Tempelaar Tempelaar Team Northwestern University Site ↗
  • Erik Thiede Thiede Lab Cornell University Site ↗
  • Pratyush Tiwary Artificial Chemical Intelligence@Maryland University of Maryland, College Park Site ↗
  • Brian Trippe Trippe Lab Stanford University Site ↗
  • Alexander Urban Urban Research Group Columbia University Site ↗
  • David Van Valen Van Valen Lab California Institute of Technology Site ↗
  • Vojtech Vlcek Vlcek Group University of California, Santa Barbara Site ↗
  • Allon Wagner Wagner Lab University of California, Berkeley Site ↗
  • Shunzhi Wang Wang Lab New York University Site ↗
  • Michael A. Webb Webb Research Group Princeton University Site ↗
  • Mingjian Wen Wen Research Group University of Houston Site ↗
  • Hong-Zhou Ye Ye Group University of Maryland, College Park Site ↗
  • Shuwen Yue Yue Research Group Cornell University Site ↗
  • Daiwei (David) Zhang Daiwei Zhang Lab University of North Carolina at Chapel Hill Site ↗
  • Hongbo Zhao Zhao Research Group University of California, San Diego Site ↗
  • Jian Zhou Zhou Lab University of Chicago Site ↗
  • Tianyu Zhu Zhu Group Yale University Site ↗

Issue 2026-08-19

Highlights from this week

Preprint—not peer reviewed

Autonomous mechanism discovery from minimal experiments

MiMEDAL combines adaptive experimentation, symbolic regression, LLM reasoning, and first-principles falsification to refine physical mechanisms from sparse data. The authors report that the autonomous system stopped after 25 experiments, reached 96.1% accuracy across 103 unseen COFs, and guided synthesis of a COF with a 61% solid-state photoluminescence quantum yield. The approach matters because it joins data-efficient experimentation to physical falsification, offering a route from statistical prediction to transferable mechanism discovery.

AI–materials Autonomous labsScientific agents
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

MACROS uses multiple agents to automate iterative hypothesis testing for molecular structure elucidation from combinations of routine spectra. The authors report zero-shot identification of diverse real-world samples above 500 daltons with one-dimensional NMR, as well as faster and more accurate elucidation through chemist collaboration. This matters because scalable spectral reasoning could move automated structure elucidation beyond fixed reference-library matching.

AI–chemistry Autonomous labsFoundation modelsMolecular representationScientific agents
Abstract brief CC0 Preprint

Preprint—not peer reviewed

High-throughput Molecular Dynamics Simulation on an AIpowered Platform

The platform uses a message-passing neural network to predict a complete polarizable-force-field parameter set from molecular structure and couples that model to automated system building and trajectory analysis. The authors report density errors below 0.02 g/cm3, ionic conductivity near 1.5 mS/cm in agreement with measurement, and solubility predictions within 10%. Reducing force-field parameterization from weeks of expert work to a single forward pass could make high-throughput molecular dynamics practical across solvents, salts, and additives.

AI–chemistryAI–materials Atomistic modelingSurrogate modeling
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

UltraIR pretrains a large infrared foundation model on simulated spectra and adapts it to multiple chemical-sensing tasks. The authors report stronger performance than conventional and task-specific learning across analytical settings, including limited-label and cross-instrument tests. This matters because a shared spectral representation could reduce the data burden of deploying infrared analysis across laboratories and sample types.

AI–chemistry Foundation modelsMolecular representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation

TransTS learns atom-level reaction transformations together with aligned reactant, transition-state, and product geometries to generate transition-state starting structures. The authors report more frequent convergence to validated saddle points and intended elementary reactions on challenging out-of-distribution benchmarks. This matters because chemically informed initial guesses can lower the quantum-chemical cost of mechanistic modeling.

AI–chemistry Generative designMolecular representationReaction modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

PROBE: An Executable Physics-Validation Benchmark for LLM-Generated Process Model Code

PROBE evaluates LLM-generated process-model code with executable checks for runnability, units, reference structure, physical invariants, and numerical agreement. The authors report no failures in 150 generations from the three Claude models across ten tasks, with an upper 95% failure-rate bound of 2% under the pooled interpretation. Replacing an LLM judge with executable physics tests makes process-model evaluation reproducible and exposes whether generated code obeys the specified scientific contract.

AI–chemistry Datasets + benchmarks
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Coupled-cluster molecular properties across the main group that extrapolate beyond training size

MEHnet-MG predicts an effective one-electron Hamiltonian from an inexpensive density-functional calculation and derives multiple molecular properties from it. The authors report three-point-eight- to 230-fold lower errors than several density-functional baselines while adding about 25 milliseconds per molecule. This matters because an architecture with the right size-scaling can extend coupled-cluster-quality trends beyond the sizes used for training.

AI–chemistry Electronic structureMolecular representationSurrogate modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Neural Networks Accelerate Ab Initio Multiple Spawning Simulations: A Case Study of Using Machine Learning Potentials for Excited State Dynamics

The hybrid holey-ML approach switches from machine-learning interatomic potentials to ab initio quantum chemistry when the predicted gap between electronic states becomes small. The authors report an order-of-magnitude cost reduction while reproducing the excited-state population decay obtained from fully ab initio simulations. This adaptive handoff makes nonadiabatic dynamics more practical without trusting the learned potential near the conical intersections where it fails.

AI–chemistry Atomistic modelingElectronic structureNeural potentials
Abstract brief CC BY Preprint

Preprint—not peer reviewed

A nuclear-quantum-corrected machine-learning potential reveals quantum-enhanced hydrogen segregation at general grain boundaries in alpha-iron

NQC-PACE relabels configurations from an iron-hydrogen machine-learning potential with finite-temperature quantum mean forces. The authors report stronger hydrogen segregation at general grain boundaries and trapping behavior closer to experimental trends in Monte Carlo and molecular-dynamics simulations. This matters because quantum effects for light solutes can be incorporated into large-scale defect simulations without new density-functional calculations.

AI–materials Atomistic modelingMultiscale modelingNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

SafeChem: A Benchmark Dataset for Multi-Label Chemical Hazard Prediction and LLM Safety Hallucination Evaluation

SafeChem combines a regulatory-grounded dataset of 32,211 substances and 30 hazard labels with separate tests of structure-based prediction and LLM safety reliability. The authors report macro-AUPRC values of 0.45–0.55 for high-prevalence labels and omission-hallucination rates above 0.21 for all eight LLMs in a 500-substance stress test. The benchmark supplies a common test bed for molecular models and general-purpose LLMs operating in chemical-safety workflows.

AI–chemistry Datasets + benchmarks
Abstract brief CC BY Preprint

Preprint—not peer reviewed

RiemannMol I: Molecular Generation with Learned Latent Metrics

RiemannMol adds an invertible, isometry-regularized metric head to a frozen molecular latent space so that distance tracks a chosen chemical property. The authors report more efficient property-guided sampling and optimization than decode-then-filter, with behavior validated against exact LogP ground truth. Learning a chemically purposeful latent metric gives molecular generators a clearer geometry for controllable sampling and optimization.

AI–chemistry Generative designMolecular representation
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Stochastic Control Policies for Robust Molecular Transition Path Sampling

FS-TPS and LaS-TPS introduce stochastic control policies for molecular transition-path sampling during explicit molecular-dynamics rollouts. The authors report better transition success and path quality than deterministic-policy baselines across three biomolecular systems, with lower sensitivity to initialization. This matters because stochasticity can improve rare-event sampling without abandoning the physical trajectory dynamics.

AI–biochemistryAI–chemistry Atomistic modelingBiomolecular structure
Abstract brief CC0 Preprint

Preprint—not peer reviewed

ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening

ScreenShot is a hierarchical transformer pretrained on drug-screening datasets that predicts combination-therapy responses from few-shot functional observations without fine-tuning or molecular profiling. The authors report that it outperformed all baselines on four held-out datasets and matched uniform-screening hit detection with a reduced budget through active learning. This approach could prioritize combination experiments when molecular profiling or cohort-specific model training is impractical.

AI–biochemistry Foundation models
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Controllable molecular graph generation from natural-language chemical constraints

MolWeaver uses a frozen language-model backbone, explicit constraint grounding, graph-level control, and scaffold-aware decoding to turn natural-language prompts into valid molecular graphs. The authors report evaluations on ZINC22-derived prompts across ring-size, Lipinski-guided, and phosphine-ligand generation tasks with specified chemical constraints. The framework makes qualitative chemical intent directly usable for molecular generation, reducing reliance on structured numerical or categorical inputs.

AI–chemistry Foundation modelsGenerative designMolecular representation
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Active learning molecular beam epitaxy of complex quantum materials

The authors combine a random-forest surrogate with thermodynamic constraints and expected improvement in a sequential model-based active-learning loop for molecular beam epitaxy. They report that, after a small initial training set, four active-learning iterations halved the absolute predictive error to about 10% for Fe3Sn growth. The framework targets abrupt phase boundaries and narrow growth windows that make continuous optimization models a poor fit for closed-loop thin-film synthesis.

AI–materials Autonomous labsSurrogate modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Agentic AI in Process Analytical Technology: LLM Assistants for Chemometric Workflows

The PAT agents combine LLMs, chemometric software, and a multi-agent orchestrator in a workflow that translates user requests into reviewable analyses. The authors report 93–96% pass rates across nine Raman tasks while grounding numerical claims in the outputs of the analytical tools. This capability could make advanced process analytics more accessible while retaining structured logs and the expert oversight needed for industrial use.

AI–chemistry Scientific agents
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Electrostatic Phenomenology Benchmarks for Machine-Learned Interatomic Potentials in Electrochemistry: Beyond the Energy-Force Metric

EPhEct is a focused test suite that probes machine-learned interatomic potentials for electrochemically relevant phenomena beyond aggregate energy and force errors. The authors report that its cases test image-charge attraction, screening through optical-phonon splitting, interfacial-water dipoles, and Fermi-level pinning during ion discharge. Such diagnostics can expose physically consequential failures that an otherwise favorable error summary would leave hidden.

AI–materials Datasets + benchmarksNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Crystal-structure design by agentic AI in a language of motifs

MatEvolve represents crystals as human-readable motif profiles that an agent edits to propose candidates, which are then tested by first-principles calculation. The authors report that its rare-earth-lean magnet designs reached new structural prototypes more than three times as often as generative models under an equal validation budget. Linking proposals to editable structural motifs offers a route to design that is both generative and inspectable at the level of recurring geometry.

AI–materials Generative designMaterials representationScientific agents
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Unlocking Multi-Component Bulk-Materials Molecular Dynamics with a Small-Footprint Machine Learning Interatomic Potential

The proposed machine-learned interatomic potential reduces feature-vector dimensionality with physical and chemical knowledge and eliminates intermediate tensors by kernel fusion. The authors report molecular dynamics of a six-component bulk system using 144 NVIDIA A100 GPUs. A substantially smaller memory footprint could make chemically heterogeneous bulk simulations accessible without the supercomputer scale previously associated with unary systems.

AI–materials Atomistic modelingNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

The study benchmarks four molecular language models across six virtual libraries, then fine-tunes their encoders on structures drawn from each target library. The authors report that domain adaptation consistently improves sample efficiency and produces several top-performing representations across the benchmark tasks. This result makes the target library itself a practical part of representation design for adaptive screening and self-driving experimentation.

AI–chemistry Foundation modelsMolecular representation
Abstract brief CC0 Preprint