Preprint scanner project / Issue 2026-08-26

Useful Chemistry

A weekly review of arXiv, ChemRxiv, and bioRxiv for ai-enabled methodological advances. Up to twenty papers reviewed every week; a quiet week may publish none.

20 selected / 117 reviewed Aug 19, 2026–Aug 25, 2026
Groups
Group roll call
96 followed groups Research-group roll call
  • Milad Abolhasani Abolhasani Lab North Carolina State University Site ↗
  • Surl-Hee (Shirley) Ahn Ahn Lab University of California, Davis Site ↗
  • Mohammed AlQuraishi AlQuraishi Laboratory Columbia University Site ↗
  • Sayan Banerjee Banerjee Group University of Tennessee, Knoxville Site ↗
  • Christopher J. Bartel Bartel Research Group University of Minnesota Twin Cities Site ↗
  • Jan-Niklas Boyn Boyn Research Group University of Minnesota Twin Cities Site ↗
  • Ting Cao Cao Group University of Washington Site ↗
  • Timothy Cernak Cernak Lab University of Michigan Site ↗
  • Ming Chen Ming Chen Group Purdue University Site ↗
  • Po-Yen Chen Chen Research Group University of Maryland, College Park Site ↗
  • Bingqing Cheng Cheng Group University of California, Berkeley Site ↗
  • Brian Cleary Algorithmic Lens on Experimental Biology Laboratory Boston University Site ↗
  • Connor W. Coley Coley Research Group Massachusetts Institute of Technology Site ↗
  • Yamil J. Colón Colón Group University of Notre Dame Site ↗
  • César de la Fuente Machine Biology Group University of Pennsylvania Site ↗
  • Julia Dshemuchadse Cape Crystal Cornell University Site ↗
  • Thomas P. Fay Fay Group University of California, Los Angeles Site ↗
  • Kara D. Fong Fong Lab California Institute of Technology Site ↗
  • Thomas E. Gartner III Gartner Group Lehigh University Site ↗
  • Rafael Gómez-Bombarelli Learning Matter Massachusetts Institute of Technology Site ↗
  • Jonathan Gootenberg Gootenberg Laboratory Harvard University Site ↗
  • Prashun Gorai 3D Materials Lab Rensselaer Polytechnic Institute Site ↗
  • Diptarka Hait Hait Group Columbia University Site ↗
  • Brian Hie Laboratory of Evolutionary Design Stanford University Site ↗
  • Guoxiang Hu Hu Group Georgia Institute of Technology Site ↗
  • Yong-Jie Hu Materials Computation and Informatics Group Drexel University Site ↗
  • Yifei Huang Huang Lab Pennsylvania State University Site ↗
  • Yunha Hwang Hwang Lab Massachusetts Institute of Technology Site ↗
  • Nicholas E. Jackson AI for Materials Group University of Illinois Urbana-Champaign Site ↗
  • William M. Jacobs Jacobs Group Princeton University Site ↗
  • Dipti Jasrasaria Jasrasaria Group University of Chicago Site ↗
  • Anupama Jha Jha Lab Yale University Site ↗
  • Kaiyi Jiang Jiang Lab Princeton University Site ↗
  • Adrian Jinich Jinich Lab University of California, San Diego Site ↗
  • Felipe Jornada Jornada Research Group Stanford University Site ↗
  • Kalli Kappel Kappel Lab University of California, Los Angeles Site ↗
  • Joshua Kretchmer Kretchmer Research Group Georgia Institute of Technology Site ↗
  • Aditi S. Krishnapriyan Krishnapriyan Research Group University of California, Berkeley Site ↗
  • Sebastian Kube Kube Lab University of Wisconsin–Madison Site ↗
  • Heather J. Kulik Kulik Research Group Massachusetts Institute of Technology Site ↗
  • Ambarish R. Kulkarni Kulkarni Research Group University of California, Davis Site ↗
  • Joseph S. Kwon Kwon Research Group The Ohio State University Site ↗
  • Joonho Lee Lee Group Harvard University Site ↗
  • Can Li Li Research Group Purdue University Site ↗
  • Wanlu Li Wanlu Li Research Group University of California, San Diego Site ↗
  • Rebecca K. Lindsey Lindsey Lab University of Michigan Site ↗
  • Ge Liu Ge Liu Group University of Illinois Urbana-Champaign Site ↗
  • Yuanyue Liu Yuanyue Liu Group The University of Texas at Austin Site ↗
  • Yang Lu Lu Lab University of Wisconsin–Madison Site ↗
  • Jiankun Lyu Evnin Family Laboratory of Computational Molecular Discovery The Rockefeller University Site ↗
  • Cong Ma Cong Ma Lab University of Michigan Site ↗
  • Arkajit Mandal Mandal Group Texas A&M University Site ↗
  • Andrew J. Medford Medford Research Group Georgia Institute of Technology Site ↗
  • Ilias Mitrai Systems and AI Lab The University of Texas at Austin Site ↗
  • Jeetain Mittal Mittal Group Texas A&M University Site ↗
  • Joel A. Paulson Paulson Lab University of Wisconsin–Madison Site ↗
  • Elisa Pieri Pieri Lab University of North Carolina at Chapel Hill Site ↗
  • Doran Raccah MesoScience Lab The University of Texas at Austin Site ↗
  • Phillip Rauscher Rauscher Group New York University Site ↗
  • Wesley Reinhart Reinhart Group Pennsylvania State University Site ↗
  • Gabriel J. Rocklin Rocklin Lab Northwestern University Site ↗
  • Andrew S. Rosen Rosen Research Group Princeton University Site ↗
  • Grant M. Rotskoff Rotskoff Group Stanford University Site ↗
  • Janani Sampath Sampath Research Group University of Florida Site ↗
  • Elvira Sayfutyarova Sayfutyarova Group Pennsylvania State University Site ↗
  • Martin Seifrid Seifrid Group North Carolina State University Site ↗
  • Thomas P. Senftle Senftle Group Rice University Site ↗
  • Karthik Shekhar Shekhar Lab University of California, Berkeley Site ↗
  • Zachary M. Sherman Z Lab University of Washington Site ↗
  • Krishna Shrinivas Shrinivas Lab Northwestern University Site ↗
  • Rohit Singh Singh Lab Duke University Site ↗
  • Micheline Soley Soley Group University of Wisconsin–Madison Site ↗
  • Kevin V. Solomon Solomon Laboratory University of Delaware Site ↗
  • Jeff Spence Spence Lab University of California, San Francisco Site ↗
  • Kayla G. Sprenger Rational Design of Interfaces Lab University of Colorado Boulder Site ↗
  • Chong Sun Sun Lab Rutgers University–New Brunswick Site ↗
  • Yidan Sun Sun Lab Washington University in St. Louis Site ↗
  • Daniel Tabor Tabor Research Group Texas A&M University Site ↗
  • Ming Tang Mesoscale Materials Science Group Rice University Site ↗
  • Roel Tempelaar Tempelaar Team Northwestern University Site ↗
  • Erik Thiede Thiede Lab Cornell University Site ↗
  • Pratyush Tiwary Artificial Chemical Intelligence@Maryland University of Maryland, College Park Site ↗
  • Brian Trippe Trippe Lab Stanford University Site ↗
  • Alexander Urban Urban Research Group Columbia University Site ↗
  • David Van Valen Van Valen Lab California Institute of Technology Site ↗
  • Vojtech Vlcek Vlcek Group University of California, Santa Barbara Site ↗
  • Allon Wagner Wagner Lab University of California, Berkeley Site ↗
  • Shunzhi Wang Wang Lab New York University Site ↗
  • Michael A. Webb Webb Research Group Princeton University Site ↗
  • Mingjian Wen Wen Research Group University of Houston Site ↗
  • Hong-Zhou Ye Ye Group University of Maryland, College Park Site ↗
  • Shuwen Yue Yue Research Group Cornell University Site ↗
  • Daiwei (David) Zhang Daiwei Zhang Lab University of North Carolina at Chapel Hill Site ↗
  • Hongbo Zhao Zhao Research Group University of California, San Diego Site ↗
  • Jian Zhou Zhou Lab University of Chicago Site ↗
  • Tianyu Zhu Zhu Group Yale University Site ↗

Issue 2026-08-26

Highlights from this week

Preprint—not peer reviewed

UBio-MolFM: Enabling Biomolecular Dynamics at DFT Accuracy and $10^5$ Atoms with One Untuned Potential

UBio-MolFM is a foundation model trained on 160 million quantum-chemical labels, designed with a receptive field spanning noncovalent distances at near-linear cost. The authors report force errors near 20 meV/Å past a thousand atoms and a 108,964-atom KcsA channel simulation that formed the anhydrous knock-on geometry in four of five replicas. Extending quantum-trained modeling to such system sizes can make electronic structure accessible where fixed-charge approximations obscure the mechanism.

AI–biochemistry Atomistic modelingFoundation modelsNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

JANUS: A Multi-modal Foundation Neural Sampler for Disordered Materials

JANUS couples continuous and masked discrete diffusion in an equivariant graph neural network trained directly from energy evaluations, sampling both atomic identities and structure. The authors report more than three orders-of-magnitude fewer energy evaluations while reproducing reference equilibrium observables and phase behavior in benchmark systems. Coupling composition with relaxation offers a unified route to thermodynamic sampling and inverse design in chemically disordered materials.

AI–materials Generative designMaterials representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Universal Machine-learning Molecular Dynamics at the Speed of Empirical Potentials

DPA4C co-designs an equivariant interatomic-potential architecture with compressed CUDA operators to pursue accuracy and throughput under deployment constraints. The authors report that its largest variant approaches MACE-Omat accuracy at roughly two orders of magnitude higher throughput, while all variants complete multimillion-atom simulations on one GPU. This moves quantum-trained universal potentials closer to the speed and scale traditionally reserved for empirical force fields.

AI–materials Atomistic modelingNeural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory

A domain-invariant SE(3)-equivariant Fourier neural operator learns the Kohn-Sham map from potential to electron density on real-space grids, enabling self-consistent-field calculations without explicit orbital construction. The authors report that one model trained on 8,504 molecules and solids generalized to out-of-distribution molecules, insulators, and metals, and converged a magnesium dislocation calculation with 82,500 valence electrons on one GPU. Learning the map rather than an ill-conditioned functional offers a path toward orbital-free calculations that retain Kohn-Sham-level observables at larger scales.

AI–materials Electronic structureSurrogate modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

The authors train C3LM on roughly 45.6 million verified reactions and use Top-K prompting with ChemCensor-based and novelty-oriented rewards to generate diverse retrosynthetic predictions. They report state-of-the-art performance on the URSA-expert-2026 benchmark and complementary reaction-space exploration by language and conventional models. Plausibility-aware multiple predictions better reflect the one-to-many character of retrosynthesis than a single-answer target.

AI–chemistry Foundation modelsSynthesis planning
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Accurate and Transferable Intermolecular Potential Based on Machine-Learned Molecular Electron Density

DensIP uses machine-learned electron densities and four universal parameters to model intermolecular interactions, trained and tested on CCSD(T)/CBS dimer energies. The authors report sub-kcal/mol errors for dimers containing molecules absent from training, including non-equilibrium conformations, and better long-range-interaction performance than general-purpose machine-learned force fields. Accurate, inexpensive synthetic reference data could ease the ab initio-data bottleneck in training more general force fields.

AI–chemistry Neural potentials
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Machine-learned exchange-correlation functionals for molecules, solids, and reactive surfaces

CIDER26SS combines machine learning with explicitly nonlocal, physically informed descriptors in a transferability-regularized exchange-correlation functional. The authors report that it resolves the CO/Pt(111) binding-site puzzle with accurate adsorption, lattice, and surface-energy predictions, including when Pt bulk and surface data are excluded from training. The result points to a functional-design route for heterogeneous catalysis that is not confined to systems represented in its training data.

AI–chemistryAI–materials Electronic structure
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Scalable photoexcitation-induced molecular dynamics with machine-learned Hamiltonians

TDAP-eML couples machine-learned electronic structure with atomistic propagation to simulate photoexcitation-induced lattice dynamics. The authors report reproduction of key photoexcited lattice responses in silicon and FeSe, with nearly three orders of magnitude lower cost for the large systems examined. The framework connects nonequilibrium electronic evolution to extended structural dynamics at a computational scale useful for experimentally accessible observables.

AI–materials Atomistic modelingElectronic structure
Abstract brief CC0 Preprint

Preprint—not peer reviewed

First-Principles Electron-Magnon Coupling with Machine-Learning Hamiltonians: From Band Renormalization to Transport

The authors combine a first-principles electron-magnon formalism with machine-learned spinful Hamiltonians to calculate transport effects in collinear magnetic systems. They report recovery of the full T2 resistivity component in ferromagnetic α-Fe with a coefficient agreeing quantitatively with measurement, and an ARPES-observed magnon kink in K-doped BaMn2As2. The framework supplies a route to assess electron-magnon contributions without reducing magnetic transport to electron-phonon effects alone.

AI–materials Electronic structure
Abstract brief CC0 Preprint

Preprint—not peer reviewed

A single design choice determines whether machine learning models of materials make physically impossible predictions

The study derives a group-theoretical parity-gap criterion for determining when a model's feature parity labels allow exact symmetry-forced zero predictions. The authors report that parity-labelled architectures stayed at the floating-point floor on centrosymmetric crystals while rotation-only models predicted forbidden piezoelectric responses on 90-96% of cases. Exposing this design choice makes exact physical constraints testable before costly training and benchmarking.

AI–materials Materials representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries

ReCurveflow learns transition-state geometries from curved reference paths derived from full NEB bands, with off-path correction during inference rollout. The authors report best results on most split-metric combinations against seven baselines and trajectories whose energy profiles closely track the reference NEB paths. Curved-path supervision and corrective fields address a mismatch between idealized straight interpolation and reaction trajectories.

AI–chemistry Generative designReaction modeling
Abstract brief CC0 Preprint

Preprint—not peer reviewed

ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation

ChemDIRT evaluates chemical reasoning in large language models by varying instructions and molecular representations across eight task categories, measuring both accuracy and consistency. The authors report substantial prompt sensitivity, representation dependence, and uneven performance across task families among benchmarked open- and closed-source models. The framework helps distinguish robust chemical reasoning from performance that depends on a favorable problem formulation.

AI–chemistry Datasets + benchmarksFoundation models
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Diagnosing and narrowing the simulation-to-real gap in powder X-ray diffraction with a wet-dry agentic loop

Xtalyst combines real-spectrum fine-tuning, peak-aligned reranking, and recalibration in an agent-orchestrated PXRD analysis system for phase identification, refinement, and property prediction. The authors report that correcting a small peak-position drift more than doubled median retrieval correlation and that a wet-dry recommend-rescan-reanalyze loop flipped a blinded silicon standard to a gated PASS. The work treats the simulation-to-real gap as a structural calibration problem inside the experimental loop rather than a generic denoising exercise.

AI–materials Scientific agents
Abstract brief CC0 Preprint

Preprint—not peer reviewed

A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models

scDrugPerturb-Bench links matched control and drug-treated single-cell profiles to literature-curated directional key-gene evidence, then measures mechanism fidelity with a composite score. The authors report that expression-similarity metrics aligned weakly with this score across 12 models, 3 baselines, and 10 data splits, while mechanism-aware selection improved early drug retrieval. The benchmark directs evaluation toward whether a model preserves drug-response signatures rather than only reconstructing expression.

AI–biochemistry Datasets + benchmarks
Abstract brief CC BY Preprint

Preprint—not peer reviewed

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a molecular foundation model pretrained on more than 81 million PM6 molecules, with stereochemistry-aware graph representation, multi-task learning, and TabPFN-based downstream prediction. The authors report that it matches or exceeds existing molecular foundation models on Polaris benchmarks and significantly improves on activity-cliff benchmarks. The combination of pretraining and adaptable downstream prediction may be especially useful where structure–activity changes are sharp and data remain limited.

AI–chemistry Foundation modelsMolecular representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Mol-JEPA learns molecular representations with modality masking across structures, cellular phenotypes, binding affinities, ADMET profiles, quantum-chemistry simulations, and other drug-discovery data. The authors report strong representation performance across multiple benchmarks. Incorporating biochemical context through latent-space prediction could make molecular models less dependent on a single data modality.

AI–chemistry Foundation modelsMolecular representation
Abstract brief CC0 Preprint

Preprint—not peer reviewed

Structure-Agnostic Prediction of the Electronic Density of States with a Chemical Language Model

DOSSIER is a chemical language model that predicts electronic density of states directly from elemental composition, with an encoder pretrained by distillation from an interatomic potential. The authors report a mean absolute error of 3.76 states eV−1 on Mat2Spec, close to the best structure-aware model at 3.64, and favorable placement of known oxygen-reduction electrocatalysts in a composition screen. Predicting spectra before crystal structures are known could broaden electronic-structure screening across unsynthesized compositions.

AI–materials Electronic structure
Abstract brief CC0 Preprint

Preprint—not peer reviewed

FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design

FAR-DPO steers cyclic-peptide generators with feasibility-gated preference pairs and difficulty-aware group-robust optimization. The authors report fixed-budget success-rate gains from 46.89% to 57.79% on PepGLAD and from 47.96% to 49.57% on PepFlow. Putting feasibility directly into optimization could improve cyclic-peptide yield without relying primarily on post hoc filtering.

AI–biochemistry Biomolecular designGenerative design
Abstract brief CC0 Preprint

Preprint—not peer reviewed

De novo Design of Macrocyclic Molecular Glues

EvoBind-multimer designs macrocyclic peptide molecular glues directly from protein sequences, bridging specified protein pairs without prior interface knowledge or existing ligands. The authors report NanoBRET evidence of design-induced proximity for VHL with KRAS and BRD4, along with ternary complexes that drove ligase-dependent degradation and signaling shutdown. Sequence-only glue design could expand induced-proximity programs beyond retrospective optimization of serendipitous binders.

AI–biochemistry Biomolecular designGenerative design
Abstract brief CC BY Preprint

Preprint—not peer reviewed

PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases

PhageLysData integrates dispersed sequence and annotation records through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The authors report 807,366 source observations consolidated into 759,105 unique sequences, including an evidence-supported Core of 11,867 entities. Its explicit separation of evidence-supported and prediction-only records gives machine-learning workflows a traceable starting point for phage-enzyme retrieval and comparison.

AI–biochemistry Datasets + benchmarks
Abstract brief CC BY Preprint