Preprint scanner project / Issue 2026-07-08

Useful Chemistry

A weekly review of arXiv, ChemRxiv, and bioRxiv for ai-enabled methodological advances. Up to twenty papers reviewed every week; a quiet week may publish none.

10 selected / 225 reviewed Jul 1, 2026–Jul 7, 2026
Groups
Group roll call
96 followed groups Research-group roll call
  • Milad Abolhasani Abolhasani Lab North Carolina State University Site ↗
  • Surl-Hee (Shirley) Ahn Ahn Lab University of California, Davis Site ↗
  • Mohammed AlQuraishi AlQuraishi Laboratory Columbia University Site ↗
  • Sayan Banerjee Banerjee Group University of Tennessee, Knoxville Site ↗
  • Christopher J. Bartel Bartel Research Group University of Minnesota Twin Cities Site ↗
  • Jan-Niklas Boyn Boyn Research Group University of Minnesota Twin Cities Site ↗
  • Ting Cao Cao Group University of Washington Site ↗
  • Timothy Cernak Cernak Lab University of Michigan Site ↗
  • Ming Chen Ming Chen Group Purdue University Site ↗
  • Po-Yen Chen Chen Research Group University of Maryland, College Park Site ↗
  • Bingqing Cheng Cheng Group University of California, Berkeley Site ↗
  • Brian Cleary Algorithmic Lens on Experimental Biology Laboratory Boston University Site ↗
  • Connor W. Coley Coley Research Group Massachusetts Institute of Technology Site ↗
  • Yamil J. Colón Colón Group University of Notre Dame Site ↗
  • César de la Fuente Machine Biology Group University of Pennsylvania Site ↗
  • Julia Dshemuchadse Cape Crystal Cornell University Site ↗
  • Thomas P. Fay Fay Group University of California, Los Angeles Site ↗
  • Kara D. Fong Fong Lab California Institute of Technology Site ↗
  • Thomas E. Gartner III Gartner Group Lehigh University Site ↗
  • Rafael Gómez-Bombarelli Learning Matter Massachusetts Institute of Technology Site ↗
  • Jonathan Gootenberg Gootenberg Laboratory Harvard University Site ↗
  • Prashun Gorai 3D Materials Lab Rensselaer Polytechnic Institute Site ↗
  • Diptarka Hait Hait Group Columbia University Site ↗
  • Brian Hie Laboratory of Evolutionary Design Stanford University Site ↗
  • Guoxiang Hu Hu Group Georgia Institute of Technology Site ↗
  • Yong-Jie Hu Materials Computation and Informatics Group Drexel University Site ↗
  • Yifei Huang Huang Lab Pennsylvania State University Site ↗
  • Yunha Hwang Hwang Lab Massachusetts Institute of Technology Site ↗
  • Nicholas E. Jackson AI for Materials Group University of Illinois Urbana-Champaign Site ↗
  • William M. Jacobs Jacobs Group Princeton University Site ↗
  • Dipti Jasrasaria Jasrasaria Group University of Chicago Site ↗
  • Anupama Jha Jha Lab Yale University Site ↗
  • Kaiyi Jiang Jiang Lab Princeton University Site ↗
  • Adrian Jinich Jinich Lab University of California, San Diego Site ↗
  • Felipe Jornada Jornada Research Group Stanford University Site ↗
  • Kalli Kappel Kappel Lab University of California, Los Angeles Site ↗
  • Joshua Kretchmer Kretchmer Research Group Georgia Institute of Technology Site ↗
  • Aditi S. Krishnapriyan Krishnapriyan Research Group University of California, Berkeley Site ↗
  • Sebastian Kube Kube Lab University of Wisconsin–Madison Site ↗
  • Heather J. Kulik Kulik Research Group Massachusetts Institute of Technology Site ↗
  • Ambarish R. Kulkarni Kulkarni Research Group University of California, Davis Site ↗
  • Joseph S. Kwon Kwon Research Group The Ohio State University Site ↗
  • Joonho Lee Lee Group Harvard University Site ↗
  • Can Li Li Research Group Purdue University Site ↗
  • Wanlu Li Wanlu Li Research Group University of California, San Diego Site ↗
  • Rebecca K. Lindsey Lindsey Lab University of Michigan Site ↗
  • Ge Liu Ge Liu Group University of Illinois Urbana-Champaign Site ↗
  • Yuanyue Liu Yuanyue Liu Group The University of Texas at Austin Site ↗
  • Yang Lu Lu Lab University of Wisconsin–Madison Site ↗
  • Jiankun Lyu Evnin Family Laboratory of Computational Molecular Discovery The Rockefeller University Site ↗
  • Cong Ma Cong Ma Lab University of Michigan Site ↗
  • Arkajit Mandal Mandal Group Texas A&M University Site ↗
  • Andrew J. Medford Medford Research Group Georgia Institute of Technology Site ↗
  • Ilias Mitrai Systems and AI Lab The University of Texas at Austin Site ↗
  • Jeetain Mittal Mittal Group Texas A&M University Site ↗
  • Joel A. Paulson Paulson Lab University of Wisconsin–Madison Site ↗
  • Elisa Pieri Pieri Lab University of North Carolina at Chapel Hill Site ↗
  • Doran Raccah MesoScience Lab The University of Texas at Austin Site ↗
  • Phillip Rauscher Rauscher Group New York University Site ↗
  • Wesley Reinhart Reinhart Group Pennsylvania State University Site ↗
  • Gabriel J. Rocklin Rocklin Lab Northwestern University Site ↗
  • Andrew S. Rosen Rosen Research Group Princeton University Site ↗
  • Grant M. Rotskoff Rotskoff Group Stanford University Site ↗
  • Janani Sampath Sampath Research Group University of Florida Site ↗
  • Elvira Sayfutyarova Sayfutyarova Group Pennsylvania State University Site ↗
  • Martin Seifrid Seifrid Group North Carolina State University Site ↗
  • Thomas P. Senftle Senftle Group Rice University Site ↗
  • Karthik Shekhar Shekhar Lab University of California, Berkeley Site ↗
  • Zachary M. Sherman Z Lab University of Washington Site ↗
  • Krishna Shrinivas Shrinivas Lab Northwestern University Site ↗
  • Rohit Singh Singh Lab Duke University Site ↗
  • Micheline Soley Soley Group University of Wisconsin–Madison Site ↗
  • Kevin V. Solomon Solomon Laboratory University of Delaware Site ↗
  • Jeff Spence Spence Lab University of California, San Francisco Site ↗
  • Kayla G. Sprenger Rational Design of Interfaces Lab University of Colorado Boulder Site ↗
  • Chong Sun Sun Lab Rutgers University–New Brunswick Site ↗
  • Yidan Sun Sun Lab Washington University in St. Louis Site ↗
  • Daniel Tabor Tabor Research Group Texas A&M University Site ↗
  • Ming Tang Mesoscale Materials Science Group Rice University Site ↗
  • Roel Tempelaar Tempelaar Team Northwestern University Site ↗
  • Erik Thiede Thiede Lab Cornell University Site ↗
  • Pratyush Tiwary Artificial Chemical Intelligence@Maryland University of Maryland, College Park Site ↗
  • Brian Trippe Trippe Lab Stanford University Site ↗
  • Alexander Urban Urban Research Group Columbia University Site ↗
  • David Van Valen Van Valen Lab California Institute of Technology Site ↗
  • Vojtech Vlcek Vlcek Group University of California, Santa Barbara Site ↗
  • Allon Wagner Wagner Lab University of California, Berkeley Site ↗
  • Shunzhi Wang Wang Lab New York University Site ↗
  • Michael A. Webb Webb Research Group Princeton University Site ↗
  • Mingjian Wen Wen Research Group University of Houston Site ↗
  • Hong-Zhou Ye Ye Group University of Maryland, College Park Site ↗
  • Shuwen Yue Yue Research Group Cornell University Site ↗
  • Daiwei (David) Zhang Daiwei Zhang Lab University of North Carolina at Chapel Hill Site ↗
  • Hongbo Zhao Zhao Research Group University of California, San Diego Site ↗
  • Jian Zhou Zhou Lab University of Chicago Site ↗
  • Tianyu Zhu Zhu Group Yale University Site ↗

Issue 2026-07-08

Highlights from this week

Preprint—not peer reviewed

Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

A multi-agent language-model pipeline classifies patent reactions, writes deterministic rules, and tests every proposed rule against a corpus of 665,901 transformations. The authors report that the resulting system expands a 68-class taxonomy to 14,073 classes without human curation. A verified rule-writing loop turns reaction classification from a fixed taxonomy into a symbolic system that can extend itself when new chemistry appears.

AI–chemistry Reaction modelingScientific agents
Abstract brief CC0 Preprint

Preprint—not peer reviewed

AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms

AquaGen samples all-atom molecular configurations with explicit solvent and periodic boundaries from a learned Boltzmann distribution, preserving compatibility with force-field evaluation and molecular dynamics refinement. For absolute hydration free energies, the authors report estimates with accuracy comparable to standard graphics-processor molecular dynamics at four- to tenfold lower cost. High-resolution ensemble generation retains the physical observables and post-processing hooks that are often lost when generative models remove solvent or coarse-grain the system.

AI–biochemistryAI–chemistryAI–materials Atomistic modelingGenerative design
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks

SpinGTP generalizes the Gaunt tensor product from scalar functions to spin-weighted spherical harmonics, restoring antisymmetric interactions while retaining its asymptotic efficiency. Across Tetris, 3BPA, SPICE-MACE-OFF, and OC20 benchmarks, the authors report accuracy comparable to full CGTP and better performance on chiral materials and non-centrosymmetric geometries. The construction provides a complete scalable equivariant basis for parity-sensitive interactions in large three-dimensional atomistic simulations.

AI–chemistryAI–materials Molecular representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

EquiFiLM: Charge-Conditioned Equivariant Force Fields via Feature-wise Linear Modulation

EquiFiLM adds continuous external-state conditioning to an equivariant foundation force field by modulating only scalar channels in each interaction layer. The authors report stable molecular dynamics across all tested charge states and prediction of the charge-dependent first-shell response measured by ultrafast electron diffraction. The lightweight conditioning axis offers a practical way to adapt ground-state foundation potentials to driven electronic processes without rebuilding them from scratch.

AI–chemistryAI–materials Atomistic modelingNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

MARLIN: De Novo Molecular Structure Elucidation from Tandem Mass Spectra without a Ground-Truth Formula

MARLIN predicts a molecular fingerprint directly from tandem mass-spectrum peaks and uses a block-diffusion language model with an exact mass-shell constraint to generate structures without assuming a molecular formula. On NPLIB1, the authors report the strongest results among methods denied the ground-truth formula across exact match, structural distance, and fingerprint similarity, while recovering the correct formula about as often as a dedicated predictor. Removing the formula oracle places de novo structure elucidation closer to the untargeted setting where genuinely novel metabolites are first encountered.

AI–chemistry Generative designMolecular representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Breaking the Synthesis Barrier for AI-Designed DNA Libraries

Policy Gradients for Library Design optimizes a synthesis-aware parameterization of stochastic DNA libraries against a chosen sequence objective. The authors report that the method supports multi-round lab-in-the-loop design and was used to synthesize a large influenza-antibody sequence library for about 700 dollars. This formulation lets generative sequence models exploit the scale of multiplexed experiments without requiring every proposed sequence to be synthesized separately.

AI–biochemistry Generative designProtein engineering
Abstract brief CC BY Preprint

Preprint—not peer reviewed

A Physics-Regulated Neural Framework for Learning 3D Grain Growth Dynamics

3D-PRIMME learns a physics-regulated local evolution rule for three-dimensional grain growth from only two consecutive time steps. After training on a 100^3 grid with 512 grains, the authors report applying the operator without retraining to 1024^3 grids with 550,000 grains while preserving coarsening kinetics and grain topology. The local rule allows grain-growth surrogates to move several orders of magnitude in system size without sacrificing the physical statistics they were trained to reproduce.

AI–materials Multiscale modelingSurrogate modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Complex crystal structure prediction using ML-enhanced multi-minima iterative genetic algorithm

The multi-minima iterative genetic algorithm couples an artificial-neural-network interatomic potential to a metadynamics-inspired penalty that steers search away from previously explored basins. The authors report recovery of the synthesized ground-state structure of La4Co4Pb and an exact match to the independently measured crystal structure of La5CoPb2 using composition alone. Navigating both the global minimum and nearby metastable states gives data-driven structure prediction a route beyond recombining entries from known crystal databases.

AI–materials Atomistic modelingGenerative designNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Improving Generalizability in Whole-Cell Antibiotic Discovery Through Active Learning

The study compares three active-learning policies for whole-cell bacterial bioactivity, selects a balance between predicted novelty and hit rate in retrospective tuberculosis data, and then closes the loop in a Borrelia antibiotic screen. The authors report a five-fold hit-rate increase from 0.2% to 1.0% in the closed-loop screen, 53-fold enrichment with 11.0% validation in prospective selections, and intended narrow-spectrum activity for all validated hits. The acquisition strategy couples exploration and exploitation to experimental feedback, allowing a whole-cell predictor to generalize into chemically diverse, out-of-distribution screens.

AI–biochemistryAI–chemistry Molecular representation
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Dyna-Mat: End-to-end benchmarking of foundation machine learning interatomic potentials in finite-temperature ensembles

Dyna-Mat evaluates 15 foundation interatomic potentials against finite-temperature first-principles trajectories, comparing static energy and force errors with structural and dynamical observables from model-driven simulations. The authors report that low force error usually tracks better ensemble behavior but can still coincide with qualitative structural failure, while pressure remains poorly described across most models. The benchmark makes trajectory-level physical behavior, rather than a favorable static error alone, the standard for judging a deployable potential.

AI–materials Atomistic modelingDatasets + benchmarksNeural potentials
Abstract brief CC0 Preprint