<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Useful Chemistry — Brett M. Savoie</title><description>Weekly preprint highlights for AI-enabled methodological advances in chemistry, biochemistry, and materials modeling.</description><link>https://brettsavoie.com/projects/useful-chemistry/</link><item><title>arXiv: RS-CIDER: A non-local machine learning model for approximating screened hybrid functionals</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13139</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13139</guid><description>RS-CIDER fits a non-local exchange functional to ground-state energies and single-particle levels, replacing the short-range Hartree-Fock term of HSE06 during self-consistent evaluation. The authors report agreement with HSE06 for reaction energies and band gaps, with a measured per-self-consistent-field-step wall time more than an order of magnitude lower. The work puts a more costly hybrid-functional reference within reach for larger periodic calculations without dropping self-consistency. Preprint—not peer reviewed.</description><pubDate>Fri, 11 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.13139</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.13139v1</dc:source><dc:creator>Zhuotao Jin</dc:creator><dc:creator>Mohamed S. Abdallah</dc:creator><dc:creator>Boris Kozinsky</dc:creator><dc:creator>Kyle Bystrom</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: uFlowCSP: Crystal Structure Prediction using Mean flow generative models</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-09799</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-09799</guid><description>uFlowCSP learns an average probability-flow velocity to generate crystal structures in a small number of evaluations. On MP-20, the authors report that one step matches CrystalFlow accuracy with far fewer evaluations, while additional steps improve accuracy further. The result shifts the emphasis in crystal generation from peak accuracy alone to useful accuracy per network evaluation. Preprint—not peer reviewed.</description><pubDate>Wed, 09 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.09799</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.09799v1</dc:source><dc:creator>Sourin Dey</dc:creator><dc:creator>Dipannoy Das Gupta</dc:creator><dc:creator>Lai Wei</dc:creator><dc:creator>Sadman Sadeed Omee</dc:creator><dc:creator>Jianjun Hu</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category></item><item><title>arXiv: Dynamic language model representations for multi-objective reaction optimisation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-11790</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-11790</guid><description>The method learns a reaction representation from textual condition descriptions with a fine-tuned language model and Gaussian-process surrogates in multi-objective Bayesian optimization. In prospective studies, the authors report that high-throughput experiments produced conditions translating directly to gram scale with high isolated yields and enantiomeric excess. The approach treats the description of a reaction system as a learned experimental representation rather than a fixed descriptor library. Preprint—not peer reviewed.</description><pubDate>Thu, 10 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.11790</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.11790v1</dc:source><dc:creator>Joshua W. Sin</dc:creator><dc:creator>David Ming Segura</dc:creator><dc:creator>Bojana Rankovic</dc:creator><dc:creator>Siu Lun Chau</dc:creator><dc:creator>Marius D. R. Lutz</dc:creator><dc:creator>Andrea Anelli</dc:creator><dc:creator>Ryan P. Burwood</dc:creator><dc:creator>Kurt Puntener</dc:creator><dc:creator>Maximilian J. Notheis</dc:creator><dc:creator>Raphael Bigler</dc:creator><dc:creator>Philippe Schwaller</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>reaction-modeling</category></item><item><title>arXiv: Neural-Network Solutions to Real-Space Charge Density and Generalization</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14906</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14906</guid><description>AIDEN separates element-dependent one-center charge density from environment-induced redistribution. The authors report state-of-the-art accuracy on periodic crystal benchmarks, competitive molecular performance, and zero-shot transfer across structurally distinct out-of-distribution cases. It advances charge-density surrogates by separating transferable local structure from the spatial grid used to evaluate the field. Preprint—not peer reviewed.</description><pubDate>Sun, 13 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.14906</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.14906v1</dc:source><dc:creator>Yuxuan Zeng</dc:creator><dc:creator>Taoyuze Lv</dc:creator><dc:creator>Zhicheng Zhong</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: pyeCE: A Python Implementation of the Embedded Cluster Expansion</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-10190</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-10190</guid><description>pyeCE implements embedded cluster expansion, using machine learning to map many chemical species onto a smaller set of effective species. The authors report that a model spanning the full composition space of a nine-component refractory alloy resolves short-range order and order-disorder behavior. This extends a familiar thermodynamic modeling framework toward alloy spaces whose chemical complexity normally makes conventional cluster expansion impractical. Preprint—not peer reviewed.</description><pubDate>Wed, 09 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.10190</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.10190v1</dc:source><dc:creator>Yann L. Muller</dc:creator><dc:creator>Claire A. Paetsch</dc:creator><dc:creator>Anirudh Raju Natarajan</dc:creator><category>arXiv</category><category>ai-materials</category><category>materials-representation</category></item><item><title>arXiv: MIRAGE: Measuring Interpolation and Redundancy in Affinity GEneralization</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14491</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14491</guid><description>MIRAGE benchmarks affinity and pose models across an explicit axis of historical protein-family support, using matched strata, family-disjoint controls, and temporal evaluation. The authors report that rankings reverse on novel families, where a family-disjoint random forest leads both co-folders and significantly outperforms Nesso-1. This makes training-set redundancy a measurable part of the claim of generalization, rather than a hidden feature of a pooled benchmark. Preprint—not peer reviewed.</description><pubDate>Sun, 13 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.14491</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.14491v1</dc:source><dc:creator>Mehdi Yazdani-Jahromi</dc:creator><dc:creator>Sanjay Padhi</dc:creator><dc:creator>Ivan Garibay</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>datasets-benchmarks</category></item><item><title>arXiv: QALPA: Property-guided diffusion modeling for efficient exploration of chemical spaces of flexible molecules</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16527</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16527</guid><description>QALPA combines an E(3)-equivariant diffusion model, active learning, and efficient quantum-mechanical methods to explore targeted property manifolds for flexible molecules. The authors report that coupling QALPA to EquiDTB augmented alloQM, a dataset of 6,253 allosteric-drug conformers, by filling sparse regions defined by dispersion energy and the HOMO-LUMO gap. This gives generative sampling a physics-based route to extend sparse quantum datasets without treating larger flexible molecules as a separate regime. Preprint—not peer reviewed.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.16527</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.16527v1</dc:source><dc:creator>Michael Hanna</dc:creator><dc:creator>Julian Cremer</dc:creator><dc:creator>Zekiye Erarslan</dc:creator><dc:creator>Leonardo Medrano Sandonas</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category></item><item><title>arXiv: Symbolic Ensemble Learning Enables Discovery of Fast Accurate Physics-Based Interatomic Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16526</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16526</guid><description>The authors train equation-learner neural networks on density-functional data and combine their symbolic forms in a weighted ensemble potential for aluminum. The authors report that all symbolic models reach sub-10 meV per atom accuracy and that their ensemble improves consistency with density-functional benchmarks for phonons, equation-of-state curvature, and melting dynamics. The result keeps a compact analytical form while using diversity among learned models to strengthen prediction beyond equilibrium structures. Preprint—not peer reviewed.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.16526</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.16526v1</dc:source><dc:creator>Bilvin Varughese, Aditya Koneru, Adil Muhammad, Troy D. Loeffler, Sukriti Manna, Jan Michael Y. Carrillo, Orcun Yildiz, Thomas Peterka</dc:creator><dc:creator>Subramanian K.R.S. Sankaranarayanan</dc:creator><category>arXiv</category><category>ai-materials</category><category>neural-potentials</category></item><item><title>arXiv: Can Autonomous LLM Agents Execute Multireference Quantum Chemistry Calculations?</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13357</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13357</guid><description>The autonomous agent selects active spaces, submits ORCA calculations, analyzes outputs, and records its choices in an auditable reasoning log. With a structured decision ladder, the authors report increased coverage of vertical transition energies alongside a lower mean absolute error. It clarifies that reliable autonomy in multireference chemistry depends on making expert workflow decisions explicit and checkable. Preprint—not peer reviewed.</description><pubDate>Fri, 11 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.13357</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.13357v1</dc:source><dc:creator>Victor Chang Lee</dc:creator><dc:creator>James M. Rondinelli</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>electronic-structure</category><category>scientific-agents</category></item><item><title>arXiv: Molecular representation shapes the balance between target fidelity and exploration in flow based polymer generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16028</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16028</guid><description>PolyLatentFlow uses continuous-time flow matching in latent space for conditional polymer generation and pairs it with the sequence-and-structure representation LlamaUni. For carbon-dioxide and nitrogen conditioning, the authors report that PolyLatentFlow with LlamaUni achieved the largest per-attempt yield of nonreplayed target hits among the evaluated representations. The comparison shows that representation determines how a polymer generator balances target control against exploration beyond its labeled chemistry. Preprint—not peer reviewed.</description><pubDate>Thu, 10 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.16028</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.16028v1</dc:source><dc:creator>Tianren Zhang</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category></item><item><title>arXiv: Multi4D: an end-to-end neural network for structural determination at complex material interfaces</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14348</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14348</guid><description>Multi4D combines a latent-space Diffusion Transformer with a rotation-invariant convolutional network for diffraction datasets. The authors report high classification accuracy and apply the method to superconducting heterostructures, corroded alloy surfaces, and degraded solid-state battery interfaces at single-nanometer resolution. It joins automated diffraction interpretation to a local measure of structural ambiguity at interfaces where degradation begins. Preprint—not peer reviewed.</description><pubDate>Sun, 13 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.14348</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.14348v1</dc:source><dc:creator>Haoran Zhang</dc:creator><dc:creator>Zian Mao</dc:creator><dc:creator>Shufen Chu</dc:creator><dc:creator>Xiaoya He</dc:creator><dc:creator>Yuyan Guan</dc:creator><dc:creator>Antong Yang</dc:creator><dc:creator>Mingze Li</dc:creator><dc:creator>Xiaoqin Zeng</dc:creator><dc:creator>Yujun Xie</dc:creator><category>arXiv</category><category>ai-materials</category><category>materials-representation</category></item><item><title>arXiv: Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14360</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14360</guid><description>The study trains a mixture-of-experts reranker under mixed spectroscopic conditions, including modality-specific perturbations and chemically informed spectrum replacements. Across a large test set spanning predefined conditions, the authors report that mixed-condition training raised the mixture-of-experts mean reciprocal rank. The contribution is a more realistic definition of robustness for structure identification when the available spectra are incomplete or inconsistent. Preprint—not peer reviewed.</description><pubDate>Sun, 13 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.14360</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.14360v1</dc:source><dc:creator>Bowen Gao</dc:creator><dc:creator>Lei Zhu</dc:creator><dc:creator>Yiying Wang</dc:creator><dc:creator>Wenjie Yu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>arXiv: Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-10193</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-10193</guid><description>The authors audit protein-protein interaction datasets with similarity-aware splitting and bias-minimizing negative sampling, both posed as integer linear programs. They report that removing train-test protein overlap still leaves shortcuts from self-interactions, taxonomic identity, and functional relatedness, with prevalence that depends on the data source. The study supplies a general way to quantify and reduce dataset shortcuts before a model’s apparent biological signal is taken at face value. Preprint—not peer reviewed.</description><pubDate>Wed, 09 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.10193</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.10193v1</dc:source><dc:creator>Judith Bernett</dc:creator><dc:creator>Anton Spannagl</dc:creator><dc:creator>Joel \AAs</dc:creator><dc:creator>Markus List</dc:creator><dc:creator>David B. Blumenthal</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>datasets-benchmarks</category></item><item><title>arXiv: Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-12223</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-12223</guid><description>GRACE adapts a pretrained geometry encoder through residual learning against an adduct-aware physical baseline and encoder-level adduct conditioning. Across random, scaffold, and adduct-sensitive splits, the authors report the best mean percentage difference among the evaluated learned models. It identifies residual learning and early adduct integration as concrete design choices for collision-cross-section prediction that must generalize beyond familiar scaffolds. Preprint—not peer reviewed.</description><pubDate>Thu, 10 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.12223</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.12223v1</dc:source><dc:creator>Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr.</dc:creator><dc:creator>Joseph A. Morrone</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>arXiv: El Agente Potente: High-Throughput Agentic Atomistic Simulations</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14840</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-14840</guid><description>El Agente Potente combines typed execution graphs for standardized interatomic-potential workflows with a coding agent for procedures that require more flexibility. The authors demonstrate the system across materials discovery, molecular energy landscapes, adsorption, and catalytic reaction workflows, alongside reproducibility and token-cost benchmarks. The division of planning from deterministic computation makes agentic atomistic simulation easier to audit while leaving room for custom workflows. Preprint—not peer reviewed.</description><pubDate>Sun, 13 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.14840</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.14840v1</dc:source><dc:creator>Tsz Wai Ko</dc:creator><dc:creator>Jiaru Bai</dc:creator><dc:creator>Thomas Swanick</dc:creator><dc:creator>Yeonghun Kang</dc:creator><dc:creator>Changhyeok Choi</dc:creator><dc:creator>Angelina Qihong Jiang</dc:creator><dc:creator>Aiwei Yin</dc:creator><dc:creator>Varinia Bernales</dc:creator><dc:creator>Alan Aspuru-Guzik</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>scientific-agents</category></item><item><title>arXiv: Ensemble-Conditioned Molecular Design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-15077</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-15077</guid><description>The authors introduce ensemble-conditioned guidance, jointly conditioning a 3D molecular generator on molecular modes and ensemble properties. The authors report that, in dual-target binder and active-state-selective agonist design, adding the additional state improved the desired outcome over single-state conditioning. That makes the conformational ensemble, rather than an isolated conformer, an explicit variable in molecular design. Preprint—not peer reviewed.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.15077</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.15077v1</dc:source><dc:creator>Ross Irwin</dc:creator><dc:creator>Alessandro Tibo</dc:creator><dc:creator>Jon Paul Janet</dc:creator><dc:creator>Simon Olsson</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category></item><item><title>arXiv: Navigating Sparse Singlet Fission Chemical Space: An Intelligent Generative-Predictive Paradigm</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-15136</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-15136</guid><description>The authors couple a structure generator, property predictor, and multi-criteria validation workflow for inverse design of singlet-fission molecules. The authors report a high success rate for candidates after evaluating a large generated set and validating a random subset with time-dependent density functional theory. The study offers a route through a sparse excited-state search space while retaining a quantum-chemical check on generated candidates. Preprint—not peer reviewed.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.15136</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.15136v1</dc:source><dc:creator>Longfei Lv</dc:creator><dc:creator>Li Fu</dc:creator><dc:creator>Si Zhou</dc:creator><dc:creator>Lingzhi Zhao</dc:creator><dc:creator>Jijun Zhao</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category></item><item><title>arXiv: Fast and Accurate Excitation Energies from Density Matrix Renormalization Group Calculations Improved by Machine Learning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-12616</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-12616</guid><description>The authors combine density matrix renormalization group calculations as a complete active-space solver with machine learning to evaluate electronic excitations. They demonstrate transferability on challenging polycyclic aromatic examples containing up to 34 pi electrons. This supplies a practical route to excited-state estimates for correlated materials that are expensive for conventional many-body calculations. Preprint—not peer reviewed.</description><pubDate>Fri, 11 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.12616</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.12616v1</dc:source><dc:creator>Pavlo Golub</dc:creator><dc:creator>Libor Veis</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: Automated AFGL quantum number assignment for CO$_2$ isotopologues using a graph neural network</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13803</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-13803</guid><description>The pipeline trains a transductive GraphSAGE network on empirical energy levels to propagate assignment information to unlabelled calculated states. The authors report assignments for previously unlabelled states across carbon-dioxide isotopologues, increasing coverage at low energy. The workflow points to automated annotation of large line lists when symmetry-like constraints can be stated directly in the assignment problem. Preprint—not peer reviewed.</description><pubDate>Sat, 12 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.13803</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.13803v1</dc:source><dc:creator>Marco G. Barnfield</dc:creator><dc:creator>Sergei N. Yurchenko</dc:creator><dc:creator>Jonathan Tennyson</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>arXiv: Scaling LLM Agents for Materials Design through Hierarchical Collective Reasoning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16466</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-16/#arxiv-2609-16466</guid><description>HiMatGen organizes language-model agents into discussion pods and domain representatives that exchange evidence and unresolved questions during crystal design. Across six chemical systems, the authors report more final crystal candidates and more stable, unique, and novel structures than a single-agent baseline. Its value lies in carrying scientific disagreement and computational evidence through an iterative materials-design process. Preprint—not peer reviewed.</description><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.16466</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.16466v1</dc:source><dc:creator>Jaehwan Choi</dc:creator><dc:creator>Yousung Jung</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>scientific-agents</category></item><item><title>arXiv: HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02746</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02746</guid><description>HiPoly processes complete polymer descriptions through a three-level hierarchical graph architecture that encodes connectivity, composition, and molecular weight. The authors report state-of-the-art thermophysical-property prediction for multicomponent polymer systems, with ablations supporting each hierarchical design choice. A shared polymer representation can connect formulation data, prediction, generative design, and simulation-based validation across polymer chemistries. Preprint—not peer reviewed.</description><pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.02746</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.02746v1</dc:source><dc:creator>Ge Sun</dc:creator><dc:creator>Gervasio Zaldivar</dc:creator><dc:creator>Yuan Tian</dc:creator><dc:creator>Gustavo Perez Lemus</dc:creator><dc:creator>Juhae Park</dc:creator><dc:creator>Daryna Safarian</dc:creator><dc:creator>Ming Han</dc:creator><dc:creator>Juan J. de Pablo</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category><category>multiscale-modeling</category></item><item><title>arXiv: La Agente \&apos;Optima: Towards Agentic Self-Driving Laboratories</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04564</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04564</guid><description>La Agente Optima constructs and supervises Bayesian optimization campaigns while maintaining persistent optimization state and separating reasoning from execution. In a multi-objective flow-chemistry campaign, the authors report increased yield over a sequence of experiments. Keeping routine loops executable and decisions auditable could make long-running optimization campaigns accessible without specialist campaign setup. Preprint—not peer reviewed.</description><pubDate>Thu, 03 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.04564</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.04564v1</dc:source><dc:creator>Marcel Muller</dc:creator><dc:creator>Jiaru Bai</dc:creator><dc:creator>Willi Gottstein</dc:creator><dc:creator>Abhijoy Mandal</dc:creator><dc:creator>Mohammad Nazeri</dc:creator><dc:creator>Elia Savino</dc:creator><dc:creator>Yanlin Fang</dc:creator><dc:creator>Sujoy Das</dc:creator><dc:creator>Sergio Pablo Garcia Carrillo</dc:creator><dc:creator>Yeonghun Kang</dc:creator><dc:creator>Juan B. Perez-Sanchez</dc:creator><dc:creator>Simone Pilon</dc:creator><dc:creator>Martin Fitzner</dc:creator><dc:creator>Timothy Noel</dc:creator><dc:creator>Frank Gu</dc:creator><dc:creator>Varinia Bernales</dc:creator><dc:creator>Alan Aspuru-Guzik</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>autonomous-labs</category><category>scientific-agents</category></item><item><title>arXiv: Quantum-accurate atomistic modeling of enzyme catalysis using a machine learned potential</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-09293</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-09293</guid><description>The authors use the eSEN-omol machine-learned interatomic potential for complete all-atom enzymes in explicit solvent. They report experimental barrier trends for chorismate mutase and mechanistic alternatives for metal-activated phosphoryl transfer. The result points to a practical route for extending quantum-accurate catalytic simulations beyond system-specific hybrid setups. Preprint—not peer reviewed.</description><pubDate>Tue, 15 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.09293</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.09293v2</dc:source><dc:creator>Meng Gao</dc:creator><dc:creator>Armin Shayesteh Zadeh</dc:creator><dc:creator>Aniruddha Seal</dc:creator><dc:creator>Siva Dasetty</dc:creator><dc:creator>Siddarth K. Achar</dc:creator><dc:creator>Misko Dzamba</dc:creator><dc:creator>Benjamin K. Miller</dc:creator><dc:creator>Leif D. Jacobson</dc:creator><dc:creator>C. Lawrence Zitnick</dc:creator><dc:creator>Brandon M. Wood</dc:creator><dc:creator>Zachary W. Ulissi</dc:creator><dc:creator>Daniel S. Levine</dc:creator><dc:creator>Andrew L. Ferguson</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08333</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08333</guid><description>The authors introduce a two-stage generative framework that relies on a fixed-dimensional molecule-level latent representation to generate variable-size molecules. On PCQM4Mv2, the authors report the highest fraction of outputs that were unique, novel, sanitized, and passed PoseBusters checks. The architecture makes molecular size a generated consequence of the representation, which matters for open-ended property-directed design. Preprint—not peer reviewed.</description><pubDate>Thu, 10 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.08333</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.08333v2</dc:source><dc:creator>Weichi Yao, Cameron Gruich, Bryan R. Goldsmith</dc:creator><dc:creator>Yixin Wang</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: Recovering molecules from coarse-grained beads: free-energy-conditioned generative backmapping across chemical space</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04432</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04432</guid><description>Juniper frames compositional backmapping as discrete denoising diffusion over molecular graphs conditioned on octanol-water partition free energy. The authors report valid and unique molecules for two-bead targets whose partition-free-energy distributions track the coarse-grained targets linearly. The method turns a lossy coarse-grained screening result back into candidate molecules for atomistic study or synthesis. Preprint—not peer reviewed.</description><pubDate>Thu, 03 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.04432</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.04432v1</dc:source><dc:creator>Luis Itza Vazquez-Salazar</dc:creator><dc:creator>Tristan Bereau</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: Out-of-Distribution Inverse Design of Elastic Networks with Differentiable Graph Neural Network Molecular Dynamics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-06655</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-06655</guid><description>This inverse-design framework uses a graph neural network molecular-dynamics simulator, a short dynamical initialization, and physics-based refinement during simulation. A simulator trained on non-auxetic systems designed strongly auxetic networks, and the framework generalizes across system size. The combination of learned dynamics and physical refinement offers a way to optimize mechanical response beyond the distribution that supplied the training data. Preprint—not peer reviewed.</description><pubDate>Sun, 06 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.06655</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.06655v1</dc:source><dc:creator>Sergey A. Shteingolts, Salman N. Salman, Ron Levie</dc:creator><dc:creator>Dan Mendels</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>surrogate-modeling</category></item><item><title>arXiv: Scalable machine learning framework for multiphase identification from powder X-ray diffraction</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-06908</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-06908</guid><description>GALAXI assigns each candidate crystalline phase an independent binary classifier, then uses Rietveld refinement to choose the combination explaining the full pattern. On curated experimental patterns, the authors report correct phase identification and robustness to common diffraction artifacts. Independent phase models make a large reference library an incremental engineering problem rather than a reason to retrain a monolithic classifier. Preprint—not peer reviewed.</description><pubDate>Sun, 06 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.06908</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.06908v1</dc:source><dc:creator>Xinyang Tong, Ethan Jin, Jiahan Xu, Aditya Rao, Pengcen Jiang,</dc:creator><dc:creator>Nathan J. Szymanski</dc:creator><category>arXiv</category><category>ai-materials</category><category>materials-representation</category></item><item><title>arXiv: Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04735</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04735</guid><description>The authors train molecular-design language models through a curriculum of synthetic tasks that gradually becomes more challenging. They report that this curriculum surpassed much larger frontier models on structure-based lead optimization. Synthetic tasks could let post-training teach useful search strategies when direct chemical scoring is too costly for online optimization. Preprint—not peer reviewed.</description><pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.04735</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.04735v1</dc:source><dc:creator>Frank Hu</dc:creator><dc:creator>Shriram Chennakesavalu</dc:creator><dc:creator>Zichen Wang</dc:creator><dc:creator>Patricia Suriana</dc:creator><dc:creator>Bodhi Vani</dc:creator><dc:creator>Kirill Shmilovich</dc:creator><dc:creator>Kangway Chuang</dc:creator><dc:creator>Colin Grambow</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>generative-design</category></item><item><title>arXiv: TSBench: A physics-grounded benchmark for evaluating LLM understanding of chemical reaction mechanisms</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08503</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08503</guid><description>TSBench asks an LLM agent to use structure-editing tools to build three-dimensional transition-state guesses checked by an automated quantum-chemical pipeline. Across 546 evaluations, the authors report that diagnosis-driven revision raised the aggregate success rate. The benchmark ties mechanistic language to a reaction path, giving claims about chemical reasoning a physical pass-or-fail test. Preprint—not peer reviewed.</description><pubDate>Tue, 08 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.08503</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.08503v1</dc:source><dc:creator>Xiaohu Xu</dc:creator><dc:creator>Tong Zhu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>datasets-benchmarks</category><category>reaction-modeling</category><category>scientific-agents</category></item><item><title>arXiv: Learning the Constitutive Behavior of Materials via Neural Operators and Causal Attention: Case Studies in Plasticity and Damage</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02194</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02194</guid><description>The framework learns a material operator from full strain histories to stress trajectories, with causally masked attention restricting access to past states. Across rate-independent material models, the authors report accurate and robust predictions of irreversible deformation with resolution invariance and parallel efficiency. A full-path operator offers a direct representation of material history when the internal variables needed by classical constitutive models are unavailable. Preprint—not peer reviewed.</description><pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.02194</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.02194v1</dc:source><dc:creator>Rishabh Arora</dc:creator><dc:creator>Lisa Scheunemann</dc:creator><dc:creator>Tim Brepols</dc:creator><dc:creator>Shahed Rezaei</dc:creator><category>arXiv</category><category>ai-materials</category><category>multiscale-modeling</category><category>surrogate-modeling</category></item><item><title>arXiv: Scalable Bayesian Optimization of Composite Functions for Image-Based Inverse Problems in Materials Characterization</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02126</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-02126</guid><description>SBOCF exploits the composite image-matching objective and intermediate simulated-image information for inverse problems in electron microscopy. Using patch-level summaries and correction terms, the authors preserve the pixel-wise objective while reducing the number of modeled outputs. The approach shows how image structure can reduce simulator demands without substituting a broad pretrained network for the scientific objective. Preprint—not peer reviewed.</description><pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.02126</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.02126v1</dc:source><dc:creator>Dasol Yoon</dc:creator><dc:creator>Poompol Buathong</dc:creator><dc:creator>Chia-Hao Lee</dc:creator><dc:creator>Yujia Zhang</dc:creator><dc:creator>David A. Muller</dc:creator><dc:creator>Peter I. Frazier</dc:creator><category>arXiv</category><category>ai-materials</category><category>materials-representation</category><category>surrogate-modeling</category></item><item><title>arXiv: SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-03377</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-03377</guid><description>SimpleDesign trains sequence and structure directly in data space with a single end-to-end objective combining sequence cross-entropy and structural regression. Trained on over 2M sequence-structure pairs, the model achieved strong performance across co-design and unconditional generation benchmarks, the authors report. The result argues that protein co-design need not pass through separately trained latent representations to obtain broad generative performance. Preprint—not peer reviewed.</description><pubDate>Thu, 03 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.03377</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.03377v1</dc:source><dc:creator>Jiarui Lu</dc:creator><dc:creator>Yuyang Wang</dc:creator><dc:creator>Yizhe Zhang</dc:creator><dc:creator>Jiatao Gu</dc:creator><dc:creator>Navdeep Jaitly</dc:creator><dc:creator>Joshua M. Susskind</dc:creator><dc:creator>Miguel Angel Bautista</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05233</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05233</guid><description>The authors introduce isotropic and normal-mode-weighted displacement schemes that derive conformation augmentation from Hessian information through Taylor expansions. Across equilibrium and non-equilibrium datasets, they report improved accuracy where the reference forces are large. The augmentation targets force-rich regimes without changing a potential architecture or its training objective, making Hessian-aware training more portable. Preprint—not peer reviewed.</description><pubDate>Tue, 15 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.05233</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.05233v2</dc:source><dc:creator>Bumju Kwak</dc:creator><dc:creator>Jeonghee Jo</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05097</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05097</guid><description>NEAT-POCKET generates molecules atom by atom in protein-pocket environments while preserving atom permutation invariance and modeling hydrogens explicitly. On CrossDocked and SPINDR, the authors report competitive structure-based generation with substantially faster sampling than existing baselines. A pocket-conditioned generator that also completes fragments brings the same representation close to lead optimization and scaffold elaboration. Preprint—not peer reviewed.</description><pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.05097</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.05097v1</dc:source><dc:creator>Roxane Axel Jacob</dc:creator><dc:creator>Daniel Rose</dc:creator><dc:creator>Thierry Langer</dc:creator><dc:creator>Johannes Kirchmair</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: The energetics of force errors in machine-learned molecular dynamics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-09251</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-09251</guid><description>The study defines a directional residual-work coefficient that combines curvature mismatch with the spatial distribution of residual force response. At a lithium-electrolyte interface, predictions fixed before future reference evaluations differ from measurements of predicted work. By resolving force error along motion, the framework gives potential assessment and adaptive reference allocation an energetic basis. Preprint—not peer reviewed.</description><pubDate>Tue, 08 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.09251</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.09251v1</dc:source><dc:creator>Peng Kang</dc:creator><dc:creator>Da Wan</dc:creator><dc:creator>Shulin Bai</dc:creator><dc:creator>Vincent Michaud-Rioux</dc:creator><dc:creator>Zhen Li</dc:creator><dc:creator>Yu Liu</dc:creator><dc:creator>Lei Zheng</dc:creator><dc:creator>Li-Dong Zhao</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: Embedded Graph Flows for Categorical Graph Generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05328</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05328</guid><description>Embedded Graph Flows learns continuous node and unordered-edge embeddings, then transports Gaussian noise toward them with a permutation-equivariant graph transformer. On QM9, the authors report the best result among the compared methods on every reported metric. Learning category geometry instead of fixing one-hot distances gives graph generation a representation that can respect order-invariant molecular structure. Preprint—not peer reviewed.</description><pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.05328</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.05328v1</dc:source><dc:creator>Ethan Ma</dc:creator><dc:creator>Zihan Wang</dc:creator><dc:creator>Chris Siu Yeung Chow</dc:creator><dc:creator>Xinguo Feng</dc:creator><dc:creator>Qingqing Li</dc:creator><dc:creator>Rui Jiang</dc:creator><dc:creator>Naipeng Dong</dc:creator><dc:creator>Guangdong Bai</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08399</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08399</guid><description>MLIP Detective turns benchmark evidence into falsifiable physics-informed failure hypotheses, then screens inexpensive simulations before escalation to experts. Without issue-specific prompting, the authors report a systematic anomaly for some oxygen- or fluorine-containing adsorbates on surfaces. That makes evaluation an active search for consequential failures, rather than a scorecard limited to the configurations already in a benchmark. Preprint—not peer reviewed.</description><pubDate>Tue, 08 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.08399</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.08399v1</dc:source><dc:creator>Ryuhei Okuno</dc:creator><dc:creator>Nontawat Charoenphakdee</dc:creator><dc:creator>Kaoru Hisama</dc:creator><dc:creator>Yuta Tsuboi</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>scientific-agents</category></item><item><title>arXiv: Molecular D\&apos;ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05381</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-05381</guid><description>The study audits frontier language models on molecular regression benchmarks for verbatim retrieval of published values. The authors report that changing the reasoning level changes retrieval even for the same molecules and prompts. This separates apparent property prediction from memorized numerical recall, a necessary distinction when using language models as scientific regressors. Preprint—not peer reviewed.</description><pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.05381</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.05381v1</dc:source><dc:creator>Matthias Busch</dc:creator><dc:creator>Marius Tacke</dc:creator><dc:creator>Sviatlana V. Lamaka</dc:creator><dc:creator>Mikhail L. Zheludkevich</dc:creator><dc:creator>Christian J. Cyron</dc:creator><dc:creator>Roland C. Aydin</dc:creator><dc:creator>Christian Feiler</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>datasets-benchmarks</category><category>foundation-models</category></item><item><title>arXiv: ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04793</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-04793</guid><description>ProtLingo augments a pretrained single-sequence protein backbone with conditional local memory and sparse expert routing. Experiments across fitness, FLIP, and contact-prediction tests report competitive performance with a 150M-scale backbone. The design tests whether reusing local sequence context and selectively activating parameters can preserve mutation-sensitive and long-range representations at modest scale. Preprint—not peer reviewed.</description><pubDate>Fri, 04 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.04793</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.04793v1</dc:source><dc:creator>Mingrui Li</dc:creator><dc:creator>Sixian Shen</dc:creator><dc:creator>Minzhang Li</dc:creator><dc:creator>Ruiyi Zhang</dc:creator><dc:creator>Kexin Zhang</dc:creator><dc:creator>Jiakai Zhang</dc:creator><dc:creator>Jingyi Yu</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>foundation-models</category><category>protein-engineering</category></item><item><title>arXiv: Learning Metamaterial Eigenmodes with Wavelet-Encoded Fourier Neural Operators</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08102</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-09/#arxiv-2609-08102</guid><description>The work combines Fourier neural operators with wavelet encodings to learn multiple eigenmodes of the elastic wave equation. For metamaterial design, the authors report a three-order-of-magnitude simulation acceleration relative to finite element analysis while retaining high fidelity. It identifies input encoding as part of the solution to mode selection in spectral neural operators, a recurring obstacle in multi-mode PDE problems. Preprint—not peer reviewed.</description><pubDate>Mon, 07 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.08102</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.08102v1</dc:source><dc:creator>Han Zhang</dc:creator><dc:creator>Alexander Ogren</dc:creator><dc:creator>Cynthia Rudin</dc:creator><dc:creator>Johann Guilleminot</dc:creator><dc:creator>L. Catherine Brinson</dc:creator><category>arXiv</category><category>ai-materials</category><category>surrogate-modeling</category></item><item><title>arXiv: uMOF: A Universal Database, Benchmark, and Machine Learning Interatomic Potentials for Metal-Organic Frameworks</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-28100</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-28100</guid><description>The authors release uMOF, combining a density functional theory dataset, a literature-mined benchmark, and universal metal-organic framework potentials. They report that uMOF models outperform tested baselines for dynamics-sensitive adsorption properties and reduce error by more than 80% to within experimental uncertainty. The package links broad physically diverse training data to experimental benchmarks for transferable metal-organic framework simulation. Preprint—not peer reviewed.</description><pubDate>Tue, 08 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.28100</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.28100v2</dc:source><dc:creator>Theo Jaffrelot Inizan</dc:creator><dc:creator>Prathami Divakar Kamath</dc:creator><dc:creator>Alin Marin Elena</dc:creator><dc:creator>Kristin A. Persson</dc:creator><category>arXiv</category><category>ai-materials</category><category>datasets-benchmarks</category><category>neural-potentials</category></item><item><title>arXiv: Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27429</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27429</guid><description>The authors formulate reaction prediction as discrete flow matching over graph-structured electron occupation vectors in a continuous-time Markov chain. On USPTO-480K and the stated out-of-distribution settings, they report competitive prediction, retained performance, mechanism-like trajectories, and side-product prediction. Electron redistribution offers a mechanistically legible alternative to product generation and direct graph edits. Preprint—not peer reviewed.</description><pubDate>Fri, 28 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.27429</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.27429v2</dc:source><dc:creator>Nguyen Xuan-Vu</dc:creator><dc:creator>Octavian Susanu</dc:creator><dc:creator>Daniel Armstrong</dc:creator><dc:creator>Philippe Schwaller</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>reaction-modeling</category></item><item><title>arXiv: Cartesian tensor equivariant machine-learning force field for spin-dependent atomistic simulations</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-30338</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-30338</guid><description>The authors build HotPP-Spin with Cartesian tensor equivariant message passing and explicit axial-vector magnetic moments. For H-phase monolayer VSe2, they report a finite-size ordering crossover at 415--435 K, close to the reported experimental value. This provides one representation for connecting first-principles magnetic energetics to coupled spin and structural simulations. Preprint—not peer reviewed.</description><pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.30338</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.30338v1</dc:source><dc:creator>Junjie Wang, Yijie Zhu, Zhongwei Zhang, Zhiyue Guo, Xudong Zhu, Lixin He, Chi Ding,</dc:creator><dc:creator>Jian Sun</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: A Generalized Approach for Incorporating Geometry and Directionality into Coarse-Grained Machine-Learned Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01911</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01911</guid><description>The authors add molecular geometry and directionality to coarse-grained potentials through anisotropic descriptors and symmetry-adapted message passing. Using Gay-Berne particles and benzene, formamide, and water representations, they report improved energy, force, and torque prediction. The work identifies retained symmetry and geometry as part of the information budget in coarse-grained models. Preprint—not peer reviewed.</description><pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.01911</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.01911v1</dc:source><dc:creator>Arthur Y. Lin</dc:creator><dc:creator>Tejas Dahiya</dc:creator><dc:creator>Rose K. Cersonsky</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00488</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00488</guid><description>The authors introduce AdaptNTK, which measures uncertainty as a regularized Mahalanobis distance in empirical neural tangent kernel feature space. On held-out rMD17 data, they report force-error correlations of 0.68 and 0.71 and a 2.6-fold speedup per Transition-1X cycle. Sequential uncertainty updates offer a route to less redundant acquisition without retraining after each selection. Preprint—not peer reviewed.</description><pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.00488</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.00488v1</dc:source><dc:creator>Prajwal Ananth</dc:creator><dc:creator>Shuwen Yue</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>neural-potentials</category></item><item><title>arXiv: Accelerating dynamic simulations of photoexcited materials and their evolution by electron-informed machine learning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01492</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01492</guid><description>The authors develop excited-state machine-learning molecular dynamics calibrated against real-time time-dependent density functional theory benchmarks. They report that large-scale simulations resolve phonon competition in bismuth phase transition and structural rearrangement in selenium photoamorphization. This extends learned dynamics toward complex photoexcited materials where electron-nuclear evolution matters. Preprint—not peer reviewed.</description><pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.01492</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.01492v1</dc:source><dc:creator>Yunzhe Jia</dc:creator><dc:creator>Fankai Xie</dc:creator><dc:creator>Yunfei Bai</dc:creator><dc:creator>Miao Liu</dc:creator><dc:creator>Cui Zhang</dc:creator><dc:creator>Sheng Meng</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category></item><item><title>arXiv: Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-29207</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-29207</guid><description>The authors introduce Hyper-Fold, a rank-K separable convolutional backbone that organizes each radius neighborhood into sequence and contact hyperedges. Across enzyme function, fold classification, and binding-site tasks, they report leading protein-encoder results and lower parameter count and latency for Hyper-Fold-Pocket. The work argues that expressive content-geometry interactions can recover information often attributed to large evolutionary pretraining. Preprint—not peer reviewed.</description><pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.29207</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.29207v2</dc:source><dc:creator>Yifan Feng</dc:creator><dc:creator>Guanjie Cheng</dc:creator><dc:creator>Shihui Ying</dc:creator><dc:creator>Shaoyi Du</dc:creator><dc:creator>Yue Gao</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-structure</category></item><item><title>arXiv: Fourier Neural Operators for Composition-Driven Crystal Structure Discovery</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00900</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00900</guid><description>The authors introduce a Fourier Neural Operator crystal-field solver that maps chemical formulae and lattice parameters to periodic density fields. They report novel structures across 104 chemical formulae with competitive reconstruction accuracy, generative diversity, and structural validity. This couples composition-conditioned generation to a reconstruction and screening path for crystal discovery. Preprint—not peer reviewed.</description><pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.00900</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.00900v1</dc:source><dc:creator>Zhijie Yu</dc:creator><dc:creator>Jingyu Li</dc:creator><dc:creator>Yang Huang</dc:creator><dc:creator>Jingrun Chen</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category></item><item><title>arXiv: Data-driven Effective Modeling of Stochastic Chemical Reaction Networks</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-25421</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-25421</guid><description>The authors approximate the finite-time transition kernel of a stochastic reaction-network Markov chain with a conditional normalizing flow. Their numerical examples report statistically consistent coarse-step trajectories with reduced computational cost. A learned stochastic propagator can decouple useful simulation steps from microscopic reaction-event resolution. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.25421</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.25421v1</dc:source><dc:creator>Weize Mao Yuan Chen</dc:creator><dc:creator>Dongbin Xiu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>surrogate-modeling</category></item><item><title>arXiv: Interpretable physics-informed retrieval-augmented generation language model for end-to-end inorganic crystal synthesis planning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-25392</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-25392</guid><description>The authors develop PIRAG-LM, which retrieves synthesis precedents by chemical, structural, and thermodynamic similarity before structured route reasoning. They report 91.4% synthesis-method prediction accuracy, compared with 72.1% for the language model alone, and experimental synthesis of five new compounds. Retrieval gives the planning system an interpretable path from materials discovery to experimental realization. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.25392</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.25392v1</dc:source><dc:creator>Wei-Jian Jiang</dc:creator><dc:creator>Ye-Nan Sha</dc:creator><dc:creator>Hui Guo</dc:creator><dc:creator>Jie Chen</dc:creator><dc:creator>Yu-Cai Liang</dc:creator><dc:creator>Ke Zhou</dc:creator><dc:creator>Qi-Long Gao</dc:creator><dc:creator>Dong-Lin Han</dc:creator><dc:creator>Xin-Gao Gong</dc:creator><dc:creator>Wan-Jian Yin</dc:creator><category>arXiv</category><category>ai-materials</category><category>scientific-agents</category><category>synthesis-planning</category></item><item><title>arXiv: Compiling Chemical Knowledge into Executable Descriptors for Materials Prediction</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27587</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27587</guid><description>The authors introduce CRISP, which samples chemical rules, consolidates them, and compiles each into an executable scalar descriptor for a conventional learner. For inorganic-crystal synthesizability, they report stronger results than expert-curated and generic representations under a shared learner, especially under structural-size and chemical-family shifts. The method makes chemical heuristics auditable computational features without relying on structures, labels, or data splits during rule construction. Preprint—not peer reviewed.</description><pubDate>Thu, 27 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.27587</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.27587v1</dc:source><dc:creator>Jaehwan Choi</dc:creator><dc:creator>Kunik Jang</dc:creator><dc:creator>Seongmin Kim</dc:creator><dc:creator>Shuan Chen</dc:creator><dc:creator>Kyungju Nam</dc:creator><dc:creator>Seung Hyo Noh</dc:creator><dc:creator>Donghwi Kim</dc:creator><dc:creator>Yousung Jung</dc:creator><category>arXiv</category><category>ai-materials</category><category>foundation-models</category><category>materials-representation</category></item><item><title>arXiv: Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-31009</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-31009</guid><description>The authors use a language-informed flow-matching framework in which target-aware SMILES supply semantic priors to geometric molecular generation. On Cross-Docked2020, they report competitive distribution matching with improved medicinal chemistry metrics and competitive structural validity under task steering. The approach tests whether language-derived chemical context can guide three-dimensional design without further generator fine-tuning. Preprint—not peer reviewed.</description><pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.31009</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.31009v1</dc:source><dc:creator>Tianyu Gao</dc:creator><dc:creator>Zhikai Su</dc:creator><dc:creator>Jiashu Li</dc:creator><dc:creator>Wenjun Gao</dc:creator><dc:creator>Zichuan Ying</dc:creator><dc:creator>Zhe Zhao</dc:creator><dc:creator>Fei Zhang</dc:creator><dc:creator>Ye Wei</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>generative-design</category></item><item><title>arXiv: S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-30910</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-30910</guid><description>The authors introduce S3C-LLM, which retrieves spectroscopy skills, runs analysis code, and integrates peak-level evidence before generating SMILES. Across diverse benchmarks, they report that S3C-LLM outperforms general and spectrum-specific models while using less than 1/10th of SpectraLLM training data. The workflow places analytical constraints inside spectrum-to-structure prediction rather than treating it as direct generation. Preprint—not peer reviewed.</description><pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.30910</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.30910v1</dc:source><dc:creator>Xuanle Zhao</dc:creator><dc:creator>Xinyuan Cai</dc:creator><dc:creator>Xiang Cheng</dc:creator><dc:creator>Bo Xu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category><category>scientific-agents</category></item><item><title>arXiv: RegimeFormer: A Large Protein Model of Global Perturbation Regimes</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26586</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26586</guid><description>The authors couple RegimeFormer with RegimeAtlas to model protein perturbation regimes across a harmonized global sequence collection. Across deep mutational scanning and molecular benchmarks, they report reproducible regimes and improved substitution-specific prediction under unseen-protein, unseen-family, and low-homology evaluation. The framework offers a scalable way to map and query mutation response across protein sequence space. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.26586</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.26586v1</dc:source><dc:creator>Siyuan Ma</dc:creator><dc:creator>Yi Chai</dc:creator><dc:creator>Yi Wu</dc:creator><dc:creator>Qixin Zhang</dc:creator><dc:creator>Yajing Yuan</dc:creator><dc:creator>Kanglu Zhao</dc:creator><dc:creator>Zhikang Chen</dc:creator><dc:creator>Haowei Wang</dc:creator><dc:creator>Shuying Cao</dc:creator><dc:creator>Xiaolei Yu</dc:creator><dc:creator>Xiangfei Han</dc:creator><dc:creator>Yun Liu</dc:creator><dc:creator>Yang Liu</dc:creator><dc:creator>Tingting Zhu</dc:creator><dc:creator>Dacheng Tao</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>foundation-models</category><category>protein-engineering</category></item><item><title>arXiv: Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00428</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-00428</guid><description>The authors develop a residual flow-map neural solver for local-box chemical kinetics in exoplanet atmospheres. They report microsecond-scale inference, percent-level accuracy, and robustness under the extreme stiffness of atmospheric chemistry. A fast surrogate could let atmospheric models retain kinetic effects that equilibrium approximations miss. Preprint—not peer reviewed.</description><pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.00428</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.00428v1</dc:source><dc:creator>Isaac Malsky, Xi Zhang, Tiffany Kataria, Matthew Graham, Ziyu Huang, Boris Bonev, Shang-Min Tsai,</dc:creator><dc:creator>Elspeth K.H. Lee</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>surrogate-modeling</category></item><item><title>arXiv: SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01353</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01353</guid><description>The authors use a symmetric dual-path architecture that iteratively combines protein language and multimodal protein language knowledge for sequence generation. Across standard inverse-folding benchmarks, they report state-of-the-art performance and ablations supporting the symmetric design. The design directly addresses the limits of post-hoc sequence refinement in inverse folding. Preprint—not peer reviewed.</description><pubDate>Tue, 01 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.01353</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.01353v1</dc:source><dc:creator>Handong Wang</dc:creator><dc:creator>Jiaxin Qi</dc:creator><dc:creator>Baisheng Lai</dc:creator><dc:creator>Jianqiang Huang</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>protein-engineering</category></item><item><title>arXiv: Text-guided flow matching enables sample-efficient crystal structure generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01076</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2609-01076</guid><description>The authors introduce TFMat, which conditions a CrystalFlow generator on structured materials language as a semantic prior. Across the stated crystal benchmarks, they report a 92.04% MP-20 match rate with 20 candidates and improved alignment in de novo generation. The result makes materials language an inspectable control layer before downstream simulation and validation. Preprint—not peer reviewed.</description><pubDate>Sat, 05 Sep 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2609.01076</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2609.01076v2</dc:source><dc:creator>Wentao Li</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category></item><item><title>arXiv: GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26585</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26585</guid><description>The authors introduce GRAS, which lowers guided-proposal variance and uses an adaptive resampling temperature for training-free discrete diffusion steering. Across regulatory DNA and protein design, they report the best training-free reward and performance that matches or exceeds a reward-fine-tuned model. The analysis isolates a compact change to inference-time search that remains effective for non-differentiable rewards. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.26585</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.26585v1</dc:source><dc:creator>Kwanyoung Kim</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>protein-engineering</category></item><item><title>arXiv: Learning Interpretable Tumor Microenvironment Representations by Fitting Pan-Cancer Cell State-Niche Correlation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26208</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-26208</guid><description>The authors introduce GITIII-scale, a hierarchical interpretable spatial transcriptomics model that decomposes cell state-niche associations and ligand-receptor pathways. On cancer types unseen in training, they report embeddings that recover niche-associated state changes more accurately than existing spatial transcriptomics foundation models. The model ties pan-cancer representation learning to biological mechanisms that can be inspected for drug-target hypotheses. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.26208</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.26208v1</dc:source><dc:creator>Xiao Xiao</dc:creator><dc:creator>Jiashu He</dc:creator><dc:creator>Shiyang Zhang</dc:creator><dc:creator>Meiyi Mao</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>foundation-models</category></item><item><title>arXiv: Ab initio Modeling of MoS2/Oxide Device Interfaces with Machine Learned Electronic Structures</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27533</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-09-02/#arxiv-2608-27533</guid><description>The authors integrate machine-learned electronic structures with a quantum transport solver for semiconductor device simulation. They report 10,000X speedups over density functional theory for devices above 20,000 atoms and identify an effect of undercoordinated Hf or Al atoms on MoS2 current. Explicit oxide layers can therefore enter transport calculations at device-relevant scales. Preprint—not peer reviewed.</description><pubDate>Thu, 27 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.27533</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.27533v1</dc:source><dc:creator>Manasa Kaniselvan, Mauro Dossena, Denghui Lu, Alexander Maeder, Nicolas Vetsch, Alexandros Nikolaos Ziogas,</dc:creator><dc:creator>Mathieu Luisier</dc:creator><category>arXiv</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: UBio-MolFM: Enabling Biomolecular Dynamics at DFT Accuracy and $10^5$ Atoms with One Untuned Potential</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18623</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18623</guid><description>UBio-MolFM is a foundation model trained on 160 million quantum-chemical labels, designed with a receptive field spanning noncovalent distances at near-linear cost. The authors report force errors near 20 meV/Å past a thousand atoms and a 108,964-atom KcsA channel simulation that formed the anhydrous knock-on geometry in four of five replicas. Extending quantum-trained modeling to such system sizes can make electronic structure accessible where fixed-charge approximations obscure the mechanism. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.18623</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.18623v1</dc:source><dc:creator>Lin Huang, Frank Peng, JiaJun Cheng, Zion Wang, Hao Yin, Hao Li, Ji Zhang, Jack Jia, Junping Zhao, Arthur Jiang,</dc:creator><dc:creator>Jia Zhang</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>atomistic-modeling</category><category>foundation-models</category><category>neural-potentials</category></item><item><title>arXiv: JANUS: A Multi-modal Foundation Neural Sampler for Disordered Materials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19116</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19116</guid><description>JANUS couples continuous and masked discrete diffusion in an equivariant graph neural network trained directly from energy evaluations, sampling both atomic identities and structure. The authors report more than three orders-of-magnitude fewer energy evaluations while reproducing reference equilibrium observables and phase behavior in benchmark systems. Coupling composition with relaxation offers a unified route to thermodynamic sampling and inverse design in chemically disordered materials. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.19116</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.19116v1</dc:source><dc:creator>Denis Blessing, Mouyang Cheng, Maximilian Schebek, Jutta Rogal, Mingda Li, Carles Domingo-Enrich,</dc:creator><dc:creator>Yuanqi Du</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category></item><item><title>arXiv: Universal Machine-learning Molecular Dynamics at the Speed of Empirical Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19041</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19041</guid><description>DPA4C co-designs an equivariant interatomic-potential architecture with compressed CUDA operators to pursue accuracy and throughput under deployment constraints. The authors report that its largest variant approaches MACE-Omat accuracy at roughly two orders of magnitude higher throughput, while all variants complete multimillion-atom simulations on one GPU. This moves quantum-trained universal potentials closer to the speed and scale traditionally reserved for empirical force fields. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.19041</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.19041v2</dc:source><dc:creator>Tiancheng Li, Jianming Xue, Linfeng Zhang, Duo Zhang</dc:creator><dc:creator>Han Wang</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-23895</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-23895</guid><description>A domain-invariant SE(3)-equivariant Fourier neural operator learns the Kohn-Sham map from potential to electron density on real-space grids, enabling self-consistent-field calculations without explicit orbital construction. The authors report that one model trained on 8,504 molecules and solids generalized to out-of-distribution molecules, insulators, and metals, and converged a magnesium dislocation calculation with 82,500 valence electrons on one GPU. Learning the map rather than an ill-conditioned functional offers a path toward orbital-free calculations that retain Kohn-Sham-level observables at larger scales. Preprint—not peer reviewed.</description><pubDate>Mon, 24 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.23895</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.23895v1</dc:source><dc:creator>Danish Khan</dc:creator><dc:creator>Maurice D. Hanisch</dc:creator><dc:creator>Nikolai Argatoff</dc:creator><dc:creator>Evan Xie</dc:creator><dc:creator>Sandeep Sharma</dc:creator><dc:creator>Anima Anandkumar</dc:creator><category>arXiv</category><category>ai-materials</category><category>electronic-structure</category><category>surrogate-modeling</category></item><item><title>arXiv: Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18940</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18940</guid><description>The authors train C3LM on roughly 45.6 million verified reactions and use Top-K prompting with ChemCensor-based and novelty-oriented rewards to generate diverse retrosynthetic predictions. They report state-of-the-art performance on the URSA-expert-2026 benchmark and complementary reaction-space exploration by language and conventional models. Plausibility-aware multiple predictions better reflect the one-to-many character of retrosynthesis than a single-answer target. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.18940</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.18940v1</dc:source><dc:creator>Bogdan Zagribelnyy</dc:creator><dc:creator>Ivan Ilin</dc:creator><dc:creator>Nikita Bondarev</dc:creator><dc:creator>Maksim Kuznetsov</dc:creator><dc:creator>Mathieu Reymond</dc:creator><dc:creator>Vladimir Aladinskiy</dc:creator><dc:creator>Alex Aliper</dc:creator><dc:creator>Alex Zhavoronkov</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>synthesis-planning</category></item><item><title>arXiv: Accurate and Transferable Intermolecular Potential Based on Machine-Learned Molecular Electron Density</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20753</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20753</guid><description>DensIP uses machine-learned electron densities and four universal parameters to model intermolecular interactions, trained and tested on CCSD(T)/CBS dimer energies. The authors report sub-kcal/mol errors for dimers containing molecules absent from training, including non-equilibrium conformations, and better long-range-interaction performance than general-purpose machine-learned force fields. Accurate, inexpensive synthetic reference data could ease the ab initio-data bottleneck in training more general force fields. Preprint—not peer reviewed.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.20753</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.20753v1</dc:source><dc:creator>Dahvyd Wing, Mihail Bogojeski, Szabolcs Goger, Klaus-Robert Muller</dc:creator><dc:creator>Alexandre Tkatchenko</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>neural-potentials</category></item><item><title>arXiv: Machine-learned exchange-correlation functionals for molecules, solids, and reactive surfaces</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-21525</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-21525</guid><description>CIDER26SS combines machine learning with explicitly nonlocal, physically informed descriptors in a transferability-regularized exchange-correlation functional. The authors report that it resolves the CO/Pt(111) binding-site puzzle with accurate adsorption, lattice, and surface-energy predictions, including when Pt bulk and surface data are excluded from training. The result points to a functional-design route for heterogeneous catalysis that is not confined to systems represented in its training data. Preprint—not peer reviewed.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.21525</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.21525v1</dc:source><dc:creator>Mohamed S. Abdallah</dc:creator><dc:creator>Zhuotao Jin</dc:creator><dc:creator>Boris Kozinsky</dc:creator><dc:creator>Kyle Bystrom</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: Scalable photoexcitation-induced molecular dynamics with machine-learned Hamiltonians</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20994</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20994</guid><description>TDAP-eML couples machine-learned electronic structure with atomistic propagation to simulate photoexcitation-induced lattice dynamics. The authors report reproduction of key photoexcited lattice responses in silicon and FeSe, with nearly three orders of magnitude lower cost for the large systems examined. The framework connects nonequilibrium electronic evolution to extended structural dynamics at a computational scale useful for experimentally accessible observables. Preprint—not peer reviewed.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.20994</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.20994v1</dc:source><dc:creator>Leyu Cai</dc:creator><dc:creator>Yunzhe Jia</dc:creator><dc:creator>Daqiang Chen</dc:creator><dc:creator>Sheng Meng</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>electronic-structure</category></item><item><title>arXiv: First-Principles Electron-Magnon Coupling with Machine-Learning Hamiltonians: From Band Renormalization to Transport</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-23333</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-23333</guid><description>The authors combine a first-principles electron-magnon formalism with machine-learned spinful Hamiltonians to calculate transport effects in collinear magnetic systems. They report recovery of the full T2 resistivity component in ferromagnetic α-Fe with a coefficient agreeing quantitatively with measurement, and an ARPES-observed magnon kink in K-doped BaMn2As2. The framework supplies a route to assess electron-magnon contributions without reducing magnetic transport to electron-phonon effects alone. Preprint—not peer reviewed.</description><pubDate>Mon, 24 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.23333</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.23333v1</dc:source><dc:creator>Shixu Liu</dc:creator><dc:creator>Xingding Li</dc:creator><dc:creator>Haozhe Li</dc:creator><dc:creator>Yang Zhong</dc:creator><dc:creator>Hongjun Xiang</dc:creator><dc:creator>Xin-Gao Gong</dc:creator><dc:creator>Ji-Hui Yang</dc:creator><category>arXiv</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: A single design choice determines whether machine learning models of materials make physically impossible predictions</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18714</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18714</guid><description>The study derives a group-theoretical parity-gap criterion for determining when a model&apos;s feature parity labels allow exact symmetry-forced zero predictions. The authors report that parity-labelled architectures stayed at the floating-point floor on centrosymmetric crystals while rotation-only models predicted forbidden piezoelectric responses on 90-96% of cases. Exposing this design choice makes exact physical constraints testable before costly training and benchmarking. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.18714</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.18714v1</dc:source><dc:creator>Can Polat</dc:creator><dc:creator>Mustafa Kurban</dc:creator><dc:creator>Erchin Serpedin</dc:creator><dc:creator>Hasan Kurban</dc:creator><category>arXiv</category><category>ai-materials</category><category>materials-representation</category></item><item><title>arXiv: ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20869</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-20869</guid><description>ReCurveflow learns transition-state geometries from curved reference paths derived from full NEB bands, with off-path correction during inference rollout. The authors report best results on most split-metric combinations against seven baselines and trajectories whose energy profiles closely track the reference NEB paths. Curved-path supervision and corrective fields address a mismatch between idealized straight interpolation and reaction trajectories. Preprint—not peer reviewed.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.20869</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.20869v1</dc:source><dc:creator>Seungheun Baek</dc:creator><dc:creator>Mogan Gim</dc:creator><dc:creator>Jaewoo Kang</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>reaction-modeling</category></item><item><title>arXiv: ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-21504</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-21504</guid><description>ChemDIRT evaluates chemical reasoning in large language models by varying instructions and molecular representations across eight task categories, measuring both accuracy and consistency. The authors report substantial prompt sensitivity, representation dependence, and uneven performance across task families among benchmarked open- and closed-source models. The framework helps distinguish robust chemical reasoning from performance that depends on a favorable problem formulation. Preprint—not peer reviewed.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.21504</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.21504v1</dc:source><dc:creator>Eric Inae</dc:creator><dc:creator>Tim Gunn</dc:creator><dc:creator>Chris Bond</dc:creator><dc:creator>Meng Jiang</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>datasets-benchmarks</category><category>foundation-models</category></item><item><title>arXiv: Diagnosing and narrowing the simulation-to-real gap in powder X-ray diffraction with a wet-dry agentic loop</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-22400</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-22400</guid><description>Xtalyst combines real-spectrum fine-tuning, peak-aligned reranking, and recalibration in an agent-orchestrated PXRD analysis system for phase identification, refinement, and property prediction. The authors report that correcting a small peak-position drift more than doubled median retrieval correlation and that a wet-dry recommend-rescan-reanalyze loop flipped a blinded silicon standard to a gated PASS. The work treats the simulation-to-real gap as a structural calibration problem inside the experimental loop rather than a generic denoising exercise. Preprint—not peer reviewed.</description><pubDate>Sun, 23 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.22400</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.22400v1</dc:source><dc:creator>Shaoguang Wang</dc:creator><dc:creator>Weiyu Guo</dc:creator><dc:creator>Ben Fei</dc:creator><dc:creator>Xiaohong Shao</dc:creator><dc:creator>Zhihui Wang</dc:creator><dc:creator>Wanli Ouyang</dc:creator><category>arXiv</category><category>ai-materials</category><category>scientific-agents</category></item><item><title>bioRxiv: A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-19-745729</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-19-745729</guid><description>scDrugPerturb-Bench links matched control and drug-treated single-cell profiles to literature-curated directional key-gene evidence, then measures mechanism fidelity with a composite score. The authors report that expression-similarity metrics aligned weakly with this score across 12 models, 3 baselines, and 10 data splits, while mechanism-aware selection improved early drug retrieval. The benchmark directs evaluation toward whether a model preserves drug-response signatures rather than only reconstructing expression. Preprint—not peer reviewed.</description><pubDate>Wed, 26 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.19.745729</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.19.745729v2</dc:source><dc:creator>Li, L.</dc:creator><dc:creator>Duan, S.</dc:creator><dc:creator>Zha, X.</dc:creator><dc:creator>Ye, F.</dc:creator><dc:creator>Zhang, Y.</dc:creator><dc:creator>Zhang, X.</dc:creator><dc:creator>Cao, Y.</dc:creator><dc:creator>Liu, C.</dc:creator><dc:creator>Fang, B.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>datasets-benchmarks</category></item><item><title>arXiv: Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18982</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-18982</guid><description>Monroe is a molecular foundation model pretrained on more than 81 million PM6 molecules, with stereochemistry-aware graph representation, multi-task learning, and TabPFN-based downstream prediction. The authors report that it matches or exceeds existing molecular foundation models on Polaris benchmarks and significantly improves on activity-cliff benchmarks. The combination of pretraining and adaptable downstream prediction may be especially useful where structure–activity changes are sharp and data remain limited. Preprint—not peer reviewed.</description><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.18982</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.18982v1</dc:source><dc:creator>Blazej Banaszewski</dc:creator><dc:creator>Andrew W. Fitzgibbon</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-22642</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-22642</guid><description>Mol-JEPA learns molecular representations with modality masking across structures, cellular phenotypes, binding affinities, ADMET profiles, quantum-chemistry simulations, and other drug-discovery data. The authors report strong representation performance across multiple benchmarks. Incorporating biochemical context through latent-space prediction could make molecular models less dependent on a single data modality. Preprint—not peer reviewed.</description><pubDate>Sun, 23 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.22642</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.22642v1</dc:source><dc:creator>Florian Rottach</dc:creator><dc:creator>Sebastian Schieferdecker</dc:creator><dc:creator>William Rudman</dc:creator><dc:creator>Randall Balestriero</dc:creator><dc:creator>Carsten Eickhoff</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Structure-Agnostic Prediction of the Electronic Density of States with a Chemical Language Model</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-24513</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-24513</guid><description>DOSSIER is a chemical language model that predicts electronic density of states directly from elemental composition, with an encoder pretrained by distillation from an interatomic potential. The authors report a mean absolute error of 3.76 states eV−1 on Mat2Spec, close to the best structure-aware model at 3.64, and favorable placement of known oxygen-reduction electrocatalysts in a composition screen. Predicting spectra before crystal structures are known could broaden electronic-structure screening across unsynthesized compositions. Preprint—not peer reviewed.</description><pubDate>Tue, 25 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.24513</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.24513v1</dc:source><dc:creator>Ivan D. Rubtsov</dc:creator><dc:creator>Ivan V. Dudakov</dc:creator><dc:creator>Vadim V. Korolev</dc:creator><category>arXiv</category><category>ai-materials</category><category>electronic-structure</category></item><item><title>arXiv: FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19808</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#arxiv-2608-19808</guid><description>FAR-DPO steers cyclic-peptide generators with feasibility-gated preference pairs and difficulty-aware group-robust optimization. The authors report fixed-budget success-rate gains from 46.89% to 57.79% on PepGLAD and from 47.96% to 49.57% on PepFlow. Putting feasibility directly into optimization could improve cyclic-peptide yield without relying primarily on post hoc filtering. Preprint—not peer reviewed.</description><pubDate>Thu, 20 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.19808</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.19808v1</dc:source><dc:creator>Guofeng Zhang</dc:creator><dc:creator>Rong Han</dc:creator><dc:creator>Xiaoyu Wang</dc:creator><dc:creator>Zhiyun Li</dc:creator><dc:creator>Zongbo Han</dc:creator><dc:creator>Xiaohong Liu</dc:creator><dc:creator>Guangyu Wang</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category></item><item><title>bioRxiv: De novo Design of Macrocyclic Molecular Glues</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-21-746227</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-21-746227</guid><description>EvoBind-multimer designs macrocyclic peptide molecular glues directly from protein sequences, bridging specified protein pairs without prior interface knowledge or existing ligands. The authors report NanoBRET evidence of design-induced proximity for VHL with KRAS and BRD4, along with ternary complexes that drove ligase-dependent degradation and signaling shutdown. Sequence-only glue design could expand induced-proximity programs beyond retrospective optimization of serendipitous binders. Preprint—not peer reviewed.</description><pubDate>Sat, 22 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.21.746227</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.21.746227v1</dc:source><dc:creator>Brunner, A.</dc:creator><dc:creator>Wierbilowicz, K.</dc:creator><dc:creator>Daumiller, D.</dc:creator><dc:creator>Bexell, D.</dc:creator><dc:creator>Karlsson, K.</dc:creator><dc:creator>Sangfelt, O.</dc:creator><dc:creator>Bryant, P.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category></item><item><title>bioRxiv: PhageLysData: an evidence-aware and AI-ready dataset of phage lytic enzymes and depolymerases</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-24-746620</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-26/#biorxiv-10-64898-2026-08-24-746620</guid><description>PhageLysData integrates dispersed sequence and annotation records through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The authors report 807,366 source observations consolidated into 759,105 unique sequences, including an evidence-supported Core of 11,867 entities. Its explicit separation of evidence-supported and prediction-only records gives machine-learning workflows a traceable starting point for phage-enzyme retrieval and comparison. Preprint—not peer reviewed.</description><pubDate>Tue, 25 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.24.746620</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.24.746620v1</dc:source><dc:creator>Medina-Ortiz, D.</dc:creator><dc:creator>Olivera-Nappa, A.</dc:creator><dc:creator>Lienqueo, M. E.</dc:creator><dc:creator>Opazo, R.</dc:creator><dc:creator>Romero, J.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>datasets-benchmarks</category></item><item><title>ChemRxiv: Autonomous mechanism discovery from minimal experiments</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007390</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007390</guid><description>MiMEDAL combines adaptive experimentation, symbolic regression, LLM reasoning, and first-principles falsification to refine physical mechanisms from sparse data. The authors report that the autonomous system stopped after 25 experiments, reached 96.1% accuracy across 103 unseen COFs, and guided synthesis of a COF with a 61% solid-state photoluminescence quantum yield. The approach matters because it joins data-efficient experimentation to physical falsification, offering a route from statistical prediction to transferable mechanism discovery. Preprint—not peer reviewed.</description><pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007390/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007390/v1?redirectToLatest=false</dc:source><dc:creator>Jiahui Du</dc:creator><dc:creator>Zikai Xie</dc:creator><dc:creator>Man Luo</dc:creator><dc:creator>Liang Zhang</dc:creator><dc:creator>Hengtao Lei</dc:creator><dc:creator>Fangming Gao</dc:creator><dc:creator>Chunxing Yan</dc:creator><dc:creator>Chengxi Zhao</dc:creator><dc:creator>Xini Chu</dc:creator><dc:creator>Ziyi Cheng</dc:creator><dc:creator>Ziyi Jin</dc:creator><dc:creator>Bing Huang</dc:creator><dc:creator>Xijun Wang</dc:creator><dc:creator>Linjiang Chen</dc:creator><dc:creator>Hexiang Deng</dc:creator><dc:creator>Jun Jiang</dc:creator><dc:creator>Yi Luo</dc:creator><category>ChemRxiv</category><category>ai-materials</category><category>autonomous-labs</category><category>scientific-agents</category></item><item><title>arXiv: Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14720</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14720</guid><description>MACROS uses multiple agents to automate iterative hypothesis testing for molecular structure elucidation from combinations of routine spectra. The authors report zero-shot identification of diverse real-world samples above 500 daltons with one-dimensional NMR, as well as faster and more accurate elucidation through chemist collaboration. This matters because scalable spectral reasoning could move automated structure elucidation beyond fixed reference-library matching. Preprint—not peer reviewed.</description><pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.14720</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.14720v1</dc:source><dc:creator>Bingsen Xue</dc:creator><dc:creator>Zhuojun Jiang</dc:creator><dc:creator>Jianhao Zhang</dc:creator><dc:creator>Mingcheng Gu</dc:creator><dc:creator>Yizhe Yuan</dc:creator><dc:creator>Yongtai Zhuo</dc:creator><dc:creator>Yifan Zhang</dc:creator><dc:creator>Li Wang</dc:creator><dc:creator>Ya Su</dc:creator><dc:creator>Yue Yuan</dc:creator><dc:creator>Jiang Liu</dc:creator><dc:creator>Xueqian Kong</dc:creator><dc:creator>Cheng Jin</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>autonomous-labs</category><category>foundation-models</category><category>molecular-representation</category><category>scientific-agents</category></item><item><title>ChemRxiv: High-throughput Molecular Dynamics Simulation on an AIpowered Platform</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007400</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007400</guid><description>The platform uses a message-passing neural network to predict a complete polarizable-force-field parameter set from molecular structure and couples that model to automated system building and trajectory analysis. The authors report density errors below 0.02 g/cm3, ionic conductivity near 1.5 mS/cm in agreement with measurement, and solubility predictions within 10%. Reducing force-field parameterization from weeks of expert work to a single forward pass could make high-throughput molecular dynamics practical across solvents, salts, and additives. Preprint—not peer reviewed.</description><pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007400/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007400/v1?redirectToLatest=false</dc:source><dc:creator>Dengpan Dong</dc:creator><dc:creator>Yani Guan</dc:creator><dc:creator>Shuang Luo</dc:creator><dc:creator>Jingxuan Ding</dc:creator><dc:creator>Dan C. Hannah</dc:creator><dc:creator>Yumin Zhang</dc:creator><dc:creator>Qichao Hu</dc:creator><dc:creator>Kang Xu</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>surrogate-modeling</category></item><item><title>arXiv: Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-13341</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-13341</guid><description>UltraIR pretrains a large infrared foundation model on simulated spectra and adapts it to multiple chemical-sensing tasks. The authors report stronger performance than conventional and task-specific learning across analytical settings, including limited-label and cross-instrument tests. This matters because a shared spectral representation could reduce the data burden of deploying infrared analysis across laboratories and sample types. Preprint—not peer reviewed.</description><pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.13341</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.13341v2</dc:source><dc:creator>Yusen Tan</dc:creator><dc:creator>Yixuan Chen</dc:creator><dc:creator>Zheng Fang</dc:creator><dc:creator>Pan Liu</dc:creator><dc:creator>Yifan Li</dc:creator><dc:creator>Qinyu Guo</dc:creator><dc:creator>Zhedong Lin</dc:creator><dc:creator>Yuqiang Li</dc:creator><dc:creator>Xiangxiang Zeng</dc:creator><dc:creator>Tong Wang</dc:creator><dc:creator>Jun Xia</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14076</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14076</guid><description>TransTS learns atom-level reaction transformations together with aligned reactant, transition-state, and product geometries to generate transition-state starting structures. The authors report more frequent convergence to validated saddle points and intended elementary reactions on challenging out-of-distribution benchmarks. This matters because chemically informed initial guesses can lower the quantum-chemical cost of mechanistic modeling. Preprint—not peer reviewed.</description><pubDate>Fri, 14 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.14076</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.14076v1</dc:source><dc:creator>Kaipeng Zeng</dc:creator><dc:creator>Wenxi Zhai</dc:creator><dc:creator>Shengrui Xu</dc:creator><dc:creator>Jie Zhao</dc:creator><dc:creator>Bowen Li</dc:creator><dc:creator>Shiyue Wang</dc:creator><dc:creator>Junchi Yan</dc:creator><dc:creator>Tong Zhu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category><category>reaction-modeling</category></item><item><title>ChemRxiv: PROBE: An Executable Physics-Validation Benchmark for LLM-Generated Process Model Code</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007535</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007535</guid><description>PROBE evaluates LLM-generated process-model code with executable checks for runnability, units, reference structure, physical invariants, and numerical agreement. The authors report no failures in 150 generations from the three Claude models across ten tasks, with an upper 95% failure-rate bound of 2% under the pooled interpretation. Replacing an LLM judge with executable physics tests makes process-model evaluation reproducible and exposes whether generated code obeys the specified scientific contract. Preprint—not peer reviewed.</description><pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007535/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007535/v1?redirectToLatest=false</dc:source><dc:creator>Nikhil Jayanth</dc:creator><dc:creator>Alexander Grunwald</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>datasets-benchmarks</category></item><item><title>arXiv: Coupled-cluster molecular properties across the main group that extrapolate beyond training size</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-18346</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-18346</guid><description>MEHnet-MG predicts an effective one-electron Hamiltonian from an inexpensive density-functional calculation and derives multiple molecular properties from it. The authors report three-point-eight- to 230-fold lower errors than several density-functional baselines while adding about 25 milliseconds per molecule. This matters because an architecture with the right size-scaling can extend coupled-cluster-quality trends beyond the sizes used for training. Preprint—not peer reviewed.</description><pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.18346</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.18346v1</dc:source><dc:creator>Wenhao He</dc:creator><dc:creator>Xu Chen</dc:creator><dc:creator>Noah Song</dc:creator><dc:creator>Haowei Xu</dc:creator><dc:creator>Tim S. Hindges</dc:creator><dc:creator>Bohan Li</dc:creator><dc:creator>Zihan Lin</dc:creator><dc:creator>Yu Yao</dc:creator><dc:creator>Avetik R. Harutyunyan</dc:creator><dc:creator>Fang Liu</dc:creator><dc:creator>Yao Wang</dc:creator><dc:creator>Hao Tang</dc:creator><dc:creator>Ju Li</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>electronic-structure</category><category>molecular-representation</category><category>surrogate-modeling</category></item><item><title>ChemRxiv: Neural Networks Accelerate Ab Initio Multiple Spawning Simulations: A Case Study of Using Machine Learning Potentials for Excited State Dynamics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007443</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007443</guid><description>The hybrid holey-ML approach switches from machine-learning interatomic potentials to ab initio quantum chemistry when the predicted gap between electronic states becomes small. The authors report an order-of-magnitude cost reduction while reproducing the excited-state population decay obtained from fully ab initio simulations. This adaptive handoff makes nonadiabatic dynamics more practical without trusting the learned potential near the conical intersections where it fails. Preprint—not peer reviewed.</description><pubDate>Fri, 14 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007443/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007443/v1?redirectToLatest=false</dc:source><dc:creator>Pablo A. Unzueta</dc:creator><dc:creator>Yuanheng Wang</dc:creator><dc:creator>Todd J. Martinez</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>electronic-structure</category><category>neural-potentials</category></item><item><title>arXiv: A nuclear-quantum-corrected machine-learning potential reveals quantum-enhanced hydrogen segregation at general grain boundaries in alpha-iron</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-16652</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-16652</guid><description>NQC-PACE relabels configurations from an iron-hydrogen machine-learning potential with finite-temperature quantum mean forces. The authors report stronger hydrogen segregation at general grain boundaries and trapping behavior closer to experimental trends in Monte Carlo and molecular-dynamics simulations. This matters because quantum effects for light solutes can be incorporated into large-scale defect simulations without new density-functional calculations. Preprint—not peer reviewed.</description><pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.16652</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.16652v1</dc:source><dc:creator>Kazuma Ito</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>multiscale-modeling</category><category>neural-potentials</category></item><item><title>ChemRxiv: SafeChem: A Benchmark Dataset for Multi-Label Chemical Hazard Prediction and LLM Safety Hallucination Evaluation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007436</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007436</guid><description>SafeChem combines a regulatory-grounded dataset of 32,211 substances and 30 hazard labels with separate tests of structure-based prediction and LLM safety reliability. The authors report macro-AUPRC values of 0.45–0.55 for high-prevalence labels and omission-hallucination rates above 0.21 for all eight LLMs in a 500-substance stress test. The benchmark supplies a common test bed for molecular models and general-purpose LLMs operating in chemical-safety workflows. Preprint—not peer reviewed.</description><pubDate>Fri, 14 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007436/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007436/v1?redirectToLatest=false</dc:source><dc:creator>Ran Elgedawy</dc:creator><dc:creator>Sanjay Das</dc:creator><dc:creator>Ethan Seefried</dc:creator><dc:creator>Ryan Burchfield</dc:creator><dc:creator>Gavin Wiggins</dc:creator><dc:creator>Kimberly Jeskie</dc:creator><dc:creator>Chrissi Schnell</dc:creator><dc:creator>Susan Fiscor</dc:creator><dc:creator>Subhamay Pramanik</dc:creator><dc:creator>Billey Thomas</dc:creator><dc:creator>Sudarshan Srinivasan</dc:creator><dc:creator>Tirthankar Ghosal</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>datasets-benchmarks</category></item><item><title>ChemRxiv: RiemannMol I: Molecular Generation with Learned Latent Metrics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007553</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007553</guid><description>RiemannMol adds an invertible, isometry-regularized metric head to a frozen molecular latent space so that distance tracks a chosen chemical property. The authors report more efficient property-guided sampling and optimization than decode-then-filter, with behavior validated against exact LogP ground truth. Learning a chemically purposeful latent metric gives molecular generators a clearer geometry for controllable sampling and optimization. Preprint—not peer reviewed.</description><pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007553/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007553/v1?redirectToLatest=false</dc:source><dc:creator>Xichen Zhang</dc:creator><dc:creator>Yizhou Ma</dc:creator><dc:creator>Xin Chen</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: Stochastic Control Policies for Robust Molecular Transition Path Sampling</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-13800</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-13800</guid><description>FS-TPS and LaS-TPS introduce stochastic control policies for molecular transition-path sampling during explicit molecular-dynamics rollouts. The authors report better transition success and path quality than deterministic-policy baselines across three biomolecular systems, with lower sensitivity to initialization. This matters because stochasticity can improve rare-event sampling without abandoning the physical trajectory dynamics. Preprint—not peer reviewed.</description><pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.13800</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.13800v1</dc:source><dc:creator>Jingqian Liu</dc:creator><dc:creator>Yu-Hsiang Wang</dc:creator><dc:creator>Yanru Qu</dc:creator><dc:creator>Ge Liu</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>biomolecular-structure</category></item><item><title>arXiv: ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-12219</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-12219</guid><description>ScreenShot is a hierarchical transformer pretrained on drug-screening datasets that predicts combination-therapy responses from few-shot functional observations without fine-tuning or molecular profiling. The authors report that it outperformed all baselines on four held-out datasets and matched uniform-screening hit detection with a reduced budget through active learning. This approach could prioritize combination experiments when molecular profiling or cohort-specific model training is impractical. Preprint—not peer reviewed.</description><pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.12219</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.12219v1</dc:source><dc:creator>Christopher Tosh Antoine de Mathelin</dc:creator><dc:creator>Wesley Tansey</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>foundation-models</category></item><item><title>ChemRxiv: Controllable molecular graph generation from natural-language chemical constraints</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007383</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007383</guid><description>MolWeaver uses a frozen language-model backbone, explicit constraint grounding, graph-level control, and scaffold-aware decoding to turn natural-language prompts into valid molecular graphs. The authors report evaluations on ZINC22-derived prompts across ring-size, Lipinski-guided, and phosphine-ligand generation tasks with specified chemical constraints. The framework makes qualitative chemical intent directly usable for molecular generation, reducing reliance on structured numerical or categorical inputs. Preprint—not peer reviewed.</description><pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007383/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007383/v1?redirectToLatest=false</dc:source><dc:creator>Li-Cheng Xu</dc:creator><dc:creator>Yu-Duo Qian</dc:creator><dc:creator>Fenglei Cao</dc:creator><dc:creator>Yuan Qi</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>foundation-models</category><category>generative-design</category><category>molecular-representation</category></item><item><title>arXiv: Active learning molecular beam epitaxy of complex quantum materials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-17742</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-17742</guid><description>The authors combine a random-forest surrogate with thermodynamic constraints and expected improvement in a sequential model-based active-learning loop for molecular beam epitaxy. They report that, after a small initial training set, four active-learning iterations halved the absolute predictive error to about 10% for Fe3Sn growth. The framework targets abrupt phase boundaries and narrow growth windows that make continuous optimization models a poor fit for closed-loop thin-film synthesis. Preprint—not peer reviewed.</description><pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.17742</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.17742v1</dc:source><dc:creator>Raghutheja Bollampally, Soumya Sankar, Yuqi Qin</dc:creator><dc:creator>Berthold Jack</dc:creator><category>arXiv</category><category>ai-materials</category><category>autonomous-labs</category><category>surrogate-modeling</category></item><item><title>ChemRxiv: Agentic AI in Process Analytical Technology: LLM Assistants for Chemometric Workflows</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007499</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#chemrxiv-10-26434-chemrxiv-15007499</guid><description>The PAT agents combine LLMs, chemometric software, and a multi-agent orchestrator in a workflow that translates user requests into reviewable analyses. The authors report 93–96% pass rates across nine Raman tasks while grounding numerical claims in the outputs of the analytical tools. This capability could make advanced process analytics more accessible while retaining structured logs and the expert oversight needed for industrial use. Preprint—not peer reviewed.</description><pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.26434/chemrxiv.15007499/v1</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://chemrxiv.org/doi/full/10.26434/chemrxiv.15007499/v1?redirectToLatest=false</dc:source><dc:creator>Jan G. Rittig</dc:creator><dc:creator>Luise F. Kaven</dc:creator><dc:creator>Philippe Schwaller</dc:creator><category>ChemRxiv</category><category>ai-chemistry</category><category>scientific-agents</category></item><item><title>arXiv: Electrostatic Phenomenology Benchmarks for Machine-Learned Interatomic Potentials in Electrochemistry: Beyond the Energy-Force Metric</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14153</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-14153</guid><description>EPhEct is a focused test suite that probes machine-learned interatomic potentials for electrochemically relevant phenomena beyond aggregate energy and force errors. The authors report that its cases test image-charge attraction, screening through optical-phonon splitting, interfacial-water dipoles, and Fermi-level pinning during ion discharge. Such diagnostics can expose physically consequential failures that an otherwise favorable error summary would leave hidden. Preprint—not peer reviewed.</description><pubDate>Fri, 14 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.14153</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.14153v1</dc:source><dc:creator>Barbara Sumic</dc:creator><dc:creator>Ria Vasdev</dc:creator><dc:creator>Sudheesh Kumar Ethirajan</dc:creator><dc:creator>Jing Yang</dc:creator><dc:creator>Clotilde S. Cucinotta</dc:creator><dc:creator>Richard G. Hennig</dc:creator><dc:creator>Karsten Reuter</dc:creator><dc:creator>Stefan Ringe</dc:creator><dc:creator>Mira Todorova</dc:creator><dc:creator>Christoph Freysoldt</dc:creator><dc:creator>Jorg Neugebauer</dc:creator><category>arXiv</category><category>ai-materials</category><category>datasets-benchmarks</category><category>neural-potentials</category></item><item><title>arXiv: Crystal-structure design by agentic AI in a language of motifs</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-15900</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-15900</guid><description>MatEvolve represents crystals as human-readable motif profiles that an agent edits to propose candidates, which are then tested by first-principles calculation. The authors report that its rare-earth-lean magnet designs reached new structural prototypes more than three times as often as generative models under an equal validation budget. Linking proposals to editable structural motifs offers a route to design that is both generative and inspectable at the level of recurring geometry. Preprint—not peer reviewed.</description><pubDate>Sun, 16 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.15900</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.15900v1</dc:source><dc:creator>Dinh-Khiet Le</dc:creator><dc:creator>Minh-Quyet Ha</dc:creator><dc:creator>Hong-Phuc Vu-Dinh</dc:creator><dc:creator>Takashi Miyake</dc:creator><dc:creator>Hiori Kino</dc:creator><dc:creator>Hieu-Chi Dam</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category><category>scientific-agents</category></item><item><title>arXiv: Unlocking Multi-Component Bulk-Materials Molecular Dynamics with a Small-Footprint Machine Learning Interatomic Potential</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-16329</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-16329</guid><description>The proposed machine-learned interatomic potential reduces feature-vector dimensionality with physical and chemical knowledge and eliminates intermediate tensors by kernel fusion. The authors report molecular dynamics of a six-component bulk system using 144 NVIDIA A100 GPUs. A substantially smaller memory footprint could make chemically heterogeneous bulk simulations accessible without the supercomputer scale previously associated with unary systems. Preprint—not peer reviewed.</description><pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.16329</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.16329v1</dc:source><dc:creator>Yucheng Ouyang</dc:creator><dc:creator>Xin Chen</dc:creator><dc:creator>Ying Liu</dc:creator><dc:creator>Lifang Wang</dc:creator><dc:creator>Xingyu Gao</dc:creator><dc:creator>Xiawei Du</dc:creator><dc:creator>Jianierken Habudelihan</dc:creator><dc:creator>Haifeng Song</dc:creator><dc:creator>Huimin Cui</dc:creator><dc:creator>Xiaobing Feng</dc:creator><dc:creator>Jingling Xue</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-17567</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-19/#arxiv-2608-17567</guid><description>The study benchmarks four molecular language models across six virtual libraries, then fine-tunes their encoders on structures drawn from each target library. The authors report that domain adaptation consistently improves sample efficiency and produces several top-performing representations across the benchmark tasks. This result makes the target library itself a practical part of representation design for adaptive screening and self-driving experimentation. Preprint—not peer reviewed.</description><pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.17567</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.17567v1</dc:source><dc:creator>Henrik Wille</dc:creator><dc:creator>Luis-Finley Schutz</dc:creator><dc:creator>Felix Strieth-Kalthoff</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-05336</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-05336</guid><description>The authors construct fixed-length molecular fingerprints from spectral graph theory using three-dimensional interaction-weighted graphs. The authors report strong benchmark performance across organic, inorganic, biological, reticular, and reaction-chemistry datasets while distinguishing structures with identical two-dimensional connectivity. This matters because an interpretable representation of three-dimensional similarity could scale beyond pairwise descriptors without depending on learned embedding coverage. Preprint—not peer reviewed.</description><pubDate>Wed, 05 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.05336</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.05336v1</dc:source><dc:creator>Jacob W. Toney</dc:creator><dc:creator>Ayleen Y. Farnood</dc:creator><dc:creator>Samir Darouich</dc:creator><dc:creator>Heather J. Kulik</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>materials-representation</category><category>molecular-representation</category></item><item><title>arXiv: A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-06694</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-06694</guid><description>CGMas coordinates agents for polymer topology construction, equilibration, coarse-graining, potential derivation, and validation from a natural-language specification. The authors report completion of all 27 polymer tasks, density agreement within five percent in 22 cases, and a reduction in simulation time to one minute. This matters because automated coarse-graining could make polymer models easier to build for mappings that otherwise require bespoke work. Preprint—not peer reviewed.</description><pubDate>Thu, 06 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.06694</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.06694v1</dc:source><dc:creator>Joohee Choi</dc:creator><dc:creator>Junhyeong Lee</dc:creator><dc:creator>Seunghwa Ryu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>multiscale-modeling</category><category>scientific-agents</category></item><item><title>arXiv: RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-06259</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-06259</guid><description>RxnCLF pretrains a reaction foundation model on condensed graphs that unify reactant and product information into an explicit transformation representation. The authors report improved yield-prediction performance over graph and sequence baselines across several reaction benchmarks. This matters because reaction representations that retain both centers and surrounding context may transfer to a wider range of reaction-informatics tasks. Preprint—not peer reviewed.</description><pubDate>Thu, 06 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.06259</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.06259v1</dc:source><dc:creator>Yiting Zheng</dc:creator><dc:creator>Cheng Fang</dc:creator><dc:creator>Anthony Donofrio</dc:creator><dc:creator>Haote Li</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category><category>reaction-modeling</category></item><item><title>bioRxiv: Boltz-Perturb: Improving Diversity and Accuracy in Protein-Ligand Co-Folding through Training-Free Conditioning Perturbation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-05-742877</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-05-742877</guid><description>Boltz-Perturb perturbs model-conditioning signals during inference through Token Bias Perturbation and Token Conditioning Perturbation, increasing exploration of alternative protein-ligand binding poses. Across diverse protein-ligand systems, the authors report that token conditioning improved top-20 oracle success rates by 2.6- to 7.8-fold, while Boltz-Perturb required over 75% less compute than the Boltz-2 high diffusion temperature variant. The method makes latent binding-mode diversity accessible without retraining, so a sampling deficiency can be addressed directly during co-folding inference. Preprint—not peer reviewed.</description><pubDate>Wed, 05 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.05.742877</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.05.742877v1</dc:source><dc:creator>Jung, H.</dc:creator><dc:creator>Lee, B.</dc:creator><dc:creator>Cheng, A. C.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>biomolecular-structure</category><category>foundation-models</category></item><item><title>arXiv: Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-07637</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#arxiv-2608-07637</guid><description>Agent-MD reserves language-model reasoning for campaign setup and event-triggered review while a persistent rule-based agent manages routine molecular-simulation operations. The authors report 120 segmented simulation cycles across 15 system-humidity states without live reasoning during production, with replay identifying workflow problems at review boundaries. This separation could make long-running simulations more auditable without turning every routine control action into a model decision. Preprint—not peer reviewed.</description><pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.07637</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.07637v1</dc:source><dc:creator>Yijie Wang</dc:creator><dc:creator>Zhen-Yu Yin</dc:creator><dc:creator>Zhenheng Tang</dc:creator><dc:creator>Xiaowen Chu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>multiscale-modeling</category><category>scientific-agents</category></item><item><title>bioRxiv: Move BeTween modAlities (MBTA) employs flow matching to predict single cell data modalities</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-05-743110</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-05-743110</guid><description>MBTA maintains modality-specific latent spaces and connects them via flow matching, preserving the structure of each modality while enabling cross-modal translation. Across multimodal single-cell benchmarks, the authors report that MBTA outperformed existing methods most strongly where structural mismatch was pronounced and identified transcriptomic lineage relationships corroborated by genomic variation in breast cancer profiles. By linking multiple molecular readouts without erasing their individual structure, the framework makes those structural differences usable in layered descriptions of cellular identity. Preprint—not peer reviewed.</description><pubDate>Sun, 09 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.05.743110</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.05.743110v1</dc:source><dc:creator>Xu, B.</dc:creator><dc:creator>Zhang, Y.</dc:creator><dc:creator>Michor, F.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>molecular-representation</category><category>surrogate-modeling</category></item><item><title>bioRxiv: Coevolution-informed Bayesian optimization for sample-efficient protein design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-06-743295</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-12/#biorxiv-10-64898-2026-08-06-743295</guid><description>ALSEBO couples Bayesian optimization to a generative protein-sequence landscape using direct-coupling-analysis coevolutionary features. The authors report reaching the avGFP optimum in about forty evaluations and transferring the approach to divergent GFPs and a non-GFP enzyme. This matters because a better low-data fitness representation can reduce the experimental cost of protein design. Preprint—not peer reviewed.</description><pubDate>Fri, 07 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.06.743295</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.06.743295v1</dc:source><dc:creator>Prasanna, D.</dc:creator><dc:creator>Shukla, D.</dc:creator><dc:creator>Potoyan, D. A.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category><category>surrogate-modeling</category></item><item><title>bioRxiv: Structure-free, site-resolved contrastive learning extends small-molecule discovery beyond the reach of structure-based modeling</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#biorxiv-10-64898-2026-07-28-741295</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#biorxiv-10-64898-2026-07-28-741295</guid><description>Ptarmigan-1 contrastively co-embeds protein residues and candidate small molecules from sequence and two-dimensional chemistry, avoiding explicit pose construction. The authors report that it matches or exceeds docking and co-folding models at covalent, cryptic, and disordered sites, and screens 3.4 billion compounds across the human proteome in under a day. A residue-resolved, pose-free index could extend virtual screening to targets where structural pockets are unavailable or poorly defined. Preprint—not peer reviewed.</description><pubDate>Thu, 20 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.28.741295</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.28.741295v2</dc:source><dc:creator>Fondrie, W. E.</dc:creator><dc:creator>Canzani, D.</dc:creator><dc:creator>Tatka, L.</dc:creator><dc:creator>Paez, J. S.</dc:creator><dc:creator>Prymolenna, A.</dc:creator><dc:creator>Gutierrez, A.</dc:creator><dc:creator>Robbins, J.</dc:creator><dc:creator>McEllin, B.</dc:creator><dc:creator>Hubbard, E.</dc:creator><dc:creator>Siebenthall, K.</dc:creator><dc:creator>Pino, L. K.</dc:creator><dc:creator>Federation, A. J.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>arXiv: Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-29158</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-29158</guid><description>Implicit force fields replace explicit neural-network stacks with self-consistent fixed-point equations whose intermediate representations can be reused between molecular-dynamics steps. The authors report two- to five-fold reductions in compute and memory across invariant, Cartesian-equivariant, and spherical-tensor graph networks. The speedup preserves atomistic resolution and the original integration timestep, opening longer trajectories and larger systems without spatial or temporal coarse graining. Preprint—not peer reviewed.</description><pubDate>Fri, 31 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.29158</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.29158v1</dc:source><dc:creator>Johannes Mae\ss</dc:creator><dc:creator>Leon Werner</dc:creator><dc:creator>J. Thorben Frank</dc:creator><dc:creator>Winfried Ripken</dc:creator><dc:creator>Martin Michajlow</dc:creator><dc:creator>Joshua Futterer</dc:creator><dc:creator>Klaus-Robert Muller</dc:creator><dc:creator>Stefan Chmiela</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-01378</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-01378</guid><description>The framework audits surrogate-driven search by reserving true evaluations for every certified conclusion, then derives rank-preservation and query-complexity conditions for when a model may act as an oracle. Across 432 surrogate fits over six task-regime conditions, the authors report that the audit statistic tracked deployed search performance at Spearman rank correlation 0.80-0.99, while audited screening reduced certified oracle cost by a measured factor of 25. The result turns trust in a surrogate into an auditable design condition, separating the model&apos;s role in proposing candidates from the true evaluations required to certify them. Preprint—not peer reviewed.</description><pubDate>Sun, 02 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.01378</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.01378v1</dc:source><dc:creator>Shuangxiu (Max) Ma</dc:creator><dc:creator>Wenhe (Zachary) Zhao</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>datasets-benchmarks</category><category>surrogate-modeling</category></item><item><title>bioRxiv: Seed-Guided De Novo Design Expands the Structural Diversity of Antitoxin Protein Binders</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#biorxiv-10-64898-2026-08-02-742339</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#biorxiv-10-64898-2026-08-02-742339</guid><description>The method guides diffusion-based protein-backbone generation with PDB-derived seed fragments selected for geometric complementarity to the target surface. The authors report that screening 1,402 designs produced multiple functional binders, including nanomolar-to-low-micromolar variants and one with neutralization comparable to the native antitoxin peptide. Surface-complementing seeds expand the structural and contact patterns available to generative binder design at difficult multi-site interfaces. Preprint—not peer reviewed.</description><pubDate>Tue, 04 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.08.02.742339</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.08.02.742339v1</dc:source><dc:creator>Britton, D.</dc:creator><dc:creator>Ghose, D. A.</dc:creator><dc:creator>Halpin, J. C.</dc:creator><dc:creator>Birnbaum, F.</dc:creator><dc:creator>Gundu, K.</dc:creator><dc:creator>Raval, S.</dc:creator><dc:creator>Papanastasiou, M.</dc:creator><dc:creator>Carr, S. A.</dc:creator><dc:creator>Keating, A. E.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: Fast and Accurate Foundation Models for Equivariant Machine-Learned Interatomic Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-28461</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-28461</guid><description>The work develops a family of equivariant foundation potentials in the NequIP and Allegro architectures, together with infrastructure for training on ultra-large datasets. The authors report leading inference speed, strong scaling, and high accuracy across materials discovery, thermal conductivity, and near-equilibrium mechanical and thermodynamic benchmarks. The analysis shifts attention from architecture alone toward the diversity and consistency of the energy surfaces used to train universal potentials. Preprint—not peer reviewed.</description><pubDate>Thu, 30 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.28461</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.28461v1</dc:source><dc:creator>Seán R. Kavanagh</dc:creator><dc:creator>Chuin Wei Tan</dc:creator><dc:creator>Menghang Wang</dc:creator><dc:creator>Marc L. Descoteaux</dc:creator><dc:creator>Gabriel de Miranda Nascimento</dc:creator><dc:creator>Ulrik Unneberg</dc:creator><dc:creator>Laura Zichi</dc:creator><dc:creator>Francesco Libbi</dc:creator><dc:creator>Norma Rivano</dc:creator><dc:creator>Austin Glover</dc:creator><dc:creator>Vivek Bharadwaj</dc:creator><dc:creator>Anders Johansson</dc:creator><dc:creator>William C. Witt</dc:creator><dc:creator>Albert Musaelian</dc:creator><dc:creator>Boris Kozinsky</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>foundation-models</category><category>neural-potentials</category></item><item><title>arXiv: Conditional grain-graph diffusion for property-guided inverse design of polycrystalline microstructures</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-00707</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-00707</guid><description>The method couples a grain graph neural-network surrogate for stress prediction with conditional reverse diffusion over alpha-phase fraction, elastic modulus, and yield-stress targets. Across four target regimes, the authors report that finite-element checks of the five best candidates reached a maximum absolute relative error of 1.0%, while diffusion used 32 evaluations per input graph rather than approximately 40,000 for the search baselines. Generating microstructures directly at prescribed properties reduces the search burden while retaining topology and crystallographic checks for physical consistency. Preprint—not peer reviewed.</description><pubDate>Sat, 01 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.00707</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.00707v1</dc:source><dc:creator>Yuheng Zhou</dc:creator><dc:creator>Xiao Shang</dc:creator><dc:creator>Huicong Chen</dc:creator><dc:creator>Yu Zou</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>materials-representation</category><category>surrogate-modeling</category></item><item><title>arXiv: SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-27431</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2607-27431</guid><description>SE(3)-MeanFlow extends MeanFlow to protein-frame Lie-group geometry and derives closed-form average-velocity targets for simulation-free training. The authors report matched or better backbone generation than flow-matching baselines using several times fewer sampling steps, with the largest gains in the few-step regime. Few-step generation directly targets the network-evaluation bottleneck in high-throughput protein design while making the associated diversity tradeoff explicit. Preprint—not peer reviewed.</description><pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.27431</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.27431v4</dc:source><dc:creator>Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang,</dc:creator><dc:creator>Guang Lin</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category></item><item><title>arXiv: A Unified Graph Neural Network Framework for Non-Equilibrium Carrier and Lattice Dynamics Driven by Electric Fields</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-03287</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-03287</guid><description>EFR-GNN predicts energies, forces, Born effective charge tensors, atom-resolved charges and magnetic moments, and supports long-time field-driven molecular dynamics that track localized electronic states. In hole-doped MgO, GaAs, and superionic alpha-AgI, the authors report field-driven polaron, phonon, and ionic dynamics that agree with a nearest-neighbor model, experiment, and temperature-dependent mobility. A single field-aware potential can therefore connect atomistic motion, electronic response, and localized-carrier dynamics over the long trajectories needed for finite-temperature materials simulations. Preprint—not peer reviewed.</description><pubDate>Tue, 04 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.03287</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.03287v1</dc:source><dc:creator>Jia-Wen Li, Sheng Meng, Xinghua Shi, Jin Zhang,</dc:creator><dc:creator>Wei-Hai Fang</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>electronic-structure</category><category>neural-potentials</category></item><item><title>arXiv: ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-03260</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-03260</guid><description>ED-DiT learns transferable molecular representations by reconstructing corrupted and masked electron-density fields while using an electron-number consistency constraint to preserve total electronic mass. The authors report that electron-density prediction RMSE fell from 2.2474 to 1.3753 and exceeded the available baseline. Pretraining on a physical field gives one encoder access to molecular properties, spin state, retrieval, and density generation rather than a single downstream label. Preprint—not peer reviewed.</description><pubDate>Tue, 04 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.03260</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.03260v1</dc:source><dc:creator>Liang Shuang</dc:creator><dc:creator>Haocheng Wang</dc:creator><dc:creator>Jiayi Song</dc:creator><dc:creator>Shuquan Ye</dc:creator><dc:creator>Ben Fei</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>electronic-structure</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-00641</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-08-05/#arxiv-2608-00641</guid><description>DASH decouples surrogate selection from acquisition control by scoring surrogates for predictive reliability, uncertainty calibration, and ranking consistency, then reallocating acquisition-function quotas before an LLM chooses from the resulting shortlist. Across four chemical optimization tasks, the authors report 12.51% better trajectory-level Acceleration Factor and 5.00% better endpoint Enhancement Factor than the strongest AutoBO baseline, with ablations supporting complementary contributions from all components. The separation gives automated optimization a way to adapt model reliability and search behavior independently as campaign feedback accumulates. Preprint—not peer reviewed.</description><pubDate>Thu, 06 Aug 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2608.00641</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2608.00641v2</dc:source><dc:creator>Changquan Zhao</dc:creator><dc:creator>Yuxiang Sun</dc:creator><dc:creator>Ruihao Zhu</dc:creator><dc:creator>Cheng Hua</dc:creator><dc:creator>Yulian He</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>scientific-agents</category><category>surrogate-modeling</category></item><item><title>arXiv: Uni-XAS: Alignment-Driven Bidirectional Multimodal Learning for X-ray Absorption Spectroscopy</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-20906</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-20906</guid><description>Uni-XAS treats spectra and atomic structures as a shared alignment-and-generation problem, with retrieval-anchored forward decoding and permutation-rectified inverse flow matching. The authors report strong cross-modal retrieval, absolute-spectrum prediction, and composition-conditioned three-dimensional structure generation on 328,839 paired structures and spectra. A common latent space makes forward and inverse spectroscopy mutually informative rather than two disconnected regressions. Preprint—not peer reviewed.</description><pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.20906</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.20906v1</dc:source><dc:creator>Suyang Zhong</dc:creator><dc:creator>Yuhao Zhao</dc:creator><dc:creator>Boying Huang</dc:creator><dc:creator>Fanjie Xu</dc:creator><dc:creator>Pengwei Xu</dc:creator><dc:creator>Haoyi Tao</dc:creator><dc:creator>Xi Fang</dc:creator><dc:creator>Jun Cheng</dc:creator><dc:creator>Fujie Tang</dc:creator><category>arXiv</category><category>ai-materials</category><category>datasets-benchmarks</category><category>generative-design</category><category>materials-representation</category></item><item><title>arXiv: MatDiffract: A Material-Informed Automated Analysis Platform for X-ray Powder Diffraction</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-20880</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-20880</guid><description>MatDiffract combines perturbation-augmented simulated patterns, multiscale vector retrieval, Rietveld refinement, and quantitative phase fitting in one automated workflow. The authors report 91.3% top-1 and 97.2% top-10 identification accuracy after automated refinement on 875 experimental single-phase patterns. Returning refined structures and quantitative compositions within seconds addresses a practical characterization bottleneck in high-throughput and autonomous materials work. Preprint—not peer reviewed.</description><pubDate>Sun, 26 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.20880</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.20880v2</dc:source><dc:creator>Hongqing Wang</dc:creator><dc:creator>Mingwei Chen</dc:creator><dc:creator>Hongjie Luo</dc:creator><dc:creator>Wen Yin</dc:creator><dc:creator>Fazhi Qi</dc:creator><dc:creator>Xuqing Chai</dc:creator><dc:creator>Miao Liu</dc:creator><dc:creator>Fengyao Hou</dc:creator><category>arXiv</category><category>ai-materials</category><category>autonomous-labs</category><category>materials-representation</category></item><item><title>arXiv: Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-25156</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-25156</guid><description>Vilya-2 extends an all-atom molecular representation from individual molecules to a diffusion transformer that models their interactions with protein targets. The authors report that calibrated structural ensembles recover 59.1% of peptide interfaces below 2 angstrom backbone RMSD, while the model also reaches state-of-the-art small-molecule docking and transfers to complex types unlike those in training. A common all-atom representation can support structure prediction across chemically distinct interface classes and then be fine-tuned for hit-to-lead enrichment. Preprint—not peer reviewed.</description><pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.25156</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.25156v1</dc:source><dc:creator>Vilya Research</dc:creator><dc:creator>:</dc:creator><dc:creator>Pascal Sturmfels</dc:creator><dc:creator>Naozumi Hiranuma</dc:creator><dc:creator>Milad Salem</dc:creator><dc:creator>Benjamin D. Sellers</dc:creator><dc:creator>Stephen Rettie</dc:creator><dc:creator>CJ San Felipe</dc:creator><dc:creator>Chase A. P. Wood</dc:creator><dc:creator>Jeffrey K. Holden</dc:creator><dc:creator>Adam P. Moyer</dc:creator><dc:creator>Patrick J. Salveson</dc:creator><dc:creator>Ivan Anishchanka</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>biomolecular-structure</category><category>foundation-models</category></item><item><title>arXiv: Rem3Di: Learning smooth, chiral 3D molecular descriptors from atomistic foundation models</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-19977</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-19977</guid><description>Rem3Di pools per-atom latent features from atomistic foundation models into smooth, order-invariant whole-molecule descriptors and adds pseudoscalar features that reverse sign under reflection to encode chirality. Across public drug-property benchmarks, the authors report that Rem3Di matched or exceeded published baselines without classical two-dimensional fingerprints and separated transition-metal complexes without predefined bonding rules. The representation carries information learned from quantum-mechanical simulation into property prediction and virtual screening while preserving three-dimensional structure and molecular handedness. Preprint—not peer reviewed.</description><pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.19977</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.19977v1</dc:source><dc:creator>Steffen Wedig, Felix Burton, Rokas Elijo\vsius, Christoph Schran,</dc:creator><dc:creator>Lars L. Schaaf</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>bioRxiv: Spaceland: Histology-Guided Reconstruction of High-Resolution Whole-Organ 3D Molecular Atlases from Sparse Spatial Transcriptomics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-739686</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-739686</guid><description>Spaceland learns continuous gene-expression fields in morphology-informed histological space, using optical-flow interpolation of foundation-model histology features to build a dense three-dimensional scaffold and decode sparse transcriptomic measurements. In mouse benchmarks, the authors report that Spaceland outperformed spatial-transcriptomic and two-dimensional histology baselines while resolving sub-spot organization; in planarian regeneration, four sections per stage supported time-resolved whole-organism analysis. The continuous-field formulation provides a scalable route from sparse sections to high-resolution whole-organ molecular reconstruction across tissues and regeneration stages. Preprint—not peer reviewed.</description><pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.23.739686</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.23.739686v1</dc:source><dc:creator>Xu, F.</dc:creator><dc:creator>Zhuang, Z.</dc:creator><dc:creator>Zhu, Y.</dc:creator><dc:creator>Ying, B.</dc:creator><dc:creator>Hou, N.</dc:creator><dc:creator>Lin, W.</dc:creator><dc:creator>Wang, L.</dc:creator><dc:creator>Yang, C.</dc:creator><dc:creator>song, j.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Transformer Atomic Cluster Expansion: TRACE</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-25652</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-25652</guid><description>TRACE combines atomic cluster-expansion density correlations with local multihead cross-attention in an energy-conserving equivariant architecture for interatomic potentials. With the same architecture, the authors report a methyl-migration activation free energy of 27.92 plus or minus 0.03 kcal/mol, close to the experimental 29.2 plus or minus 1.1 kcal/mol, together with experimental agreement for perovskite phase behavior and liquid-water structure. An equivariant potential that spans crystallization, liquid structure, and chemical reactivity reduces the need for separate models tied to individual phases or processes. Preprint—not peer reviewed.</description><pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.25652</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.25652v1</dc:source><dc:creator>Paramvir Ahlawat</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>bioRxiv: Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-740427</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-740427</guid><description>The workflow diversifies a parent protoglobin active site and uses active-learning-assisted directed evolution to predict promiscuous variant libraries across carbene and nitrene transfer reactions. Across 26 reactions, the authors report improved activity and selectivity for every parent-catalyzed reaction in at least one predicted-library member, plus activity on five of ten reactions absent from the parent enzyme. The strategy produces enzyme libraries that retain parent functions while adding transformations outside the parent&apos;s catalytic repertoire. Preprint—not peer reviewed.</description><pubDate>Fri, 24 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.23.740427</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.23.740427v1</dc:source><dc:creator>Lal, R.</dc:creator><dc:creator>Yang, J.</dc:creator><dc:creator>Zhang, Z.</dc:creator><dc:creator>Arnold, F. H.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>biomolecular-design</category><category>protein-engineering</category></item><item><title>arXiv: MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-23607</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-23607</guid><description>MS-GPT queries a conditional molecule-language model with a calibrated band of spectrum-induced fingerprints, then ranks pooled candidates by generation-frequency consensus. The authors report new state-of-the-art exact-match accuracy on NPLIB1 and MassSpecGym at both top-1 and top-10 ranks. Treating the noisy fingerprint as a posterior rather than a thresholded answer preserves uncertainty where de novo spectral identification needs it most. Preprint—not peer reviewed.</description><pubDate>Sun, 26 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.23607</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.23607v1</dc:source><dc:creator>Xin Zhao</dc:creator><dc:creator>Yumin Liu</dc:creator><dc:creator>Zhuo Li</dc:creator><dc:creator>Weichu Zheng</dc:creator><dc:creator>Feng Zhu</dc:creator><dc:creator>Xiaokang Yang</dc:creator><dc:creator>Yaohui Jin</dc:creator><dc:creator>Yanyan Xu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>bioRxiv: Spatium: A Protein Language Foundation Model for Spatial Proteomics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-740264</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#biorxiv-10-64898-2026-07-23-740264</guid><description>Spatium is a protein language foundation model trained on more than 51 million cells across spatial-proteomics platforms to learn protein co-expression hierarchies that remain stable across panel composition and measurement scale. The authors report that Spatium recovered cell identities with known marker patterns, separated spatial microenvironments by coherent enrichment signatures, and reconstructed missing proteins while preserving biological expression patterns. A panel-robust representation lets heterogeneous spatial-proteomics datasets support cell-identity, microenvironment, and missing-marker tasks with only lightweight adaptation. Preprint—not peer reviewed.</description><pubDate>Sun, 26 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.23.740264</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.23.740264v1</dc:source><dc:creator>Wang, T.</dc:creator><dc:creator>Wu, S.</dc:creator><dc:creator>Huang, L.</dc:creator><dc:creator>Liu, J.</dc:creator><dc:creator>Huang, K.</dc:creator><dc:creator>Zhou, X.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Aligning Heterogeneous DFT Datasets: A Graph Neural Network Approach to Cross-Functional Formation Energies</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-24327</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-29/#arxiv-2607-24327</guid><description>A structure-aware graph neural network learns cross-functional energy residuals from 380,190 paired PBE-r2SCAN structures, converting legacy PBE formation energies onto the r2SCAN scale. The authors report a mean absolute error of 14.3 meV/atom when converting PBE energies to r2SCAN-level accuracy, compared with 18.2 meV/atom for CHGNet. Cross-functional alignment makes large legacy DFT collections usable alongside higher-fidelity calculations for materials foundation models and thermodynamic prediction. Preprint—not peer reviewed.</description><pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.24327</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.24327v1</dc:source><dc:creator>Yidong Huang</dc:creator><dc:creator>Tenglong Lu</dc:creator><dc:creator>Hanwen Kang</dc:creator><dc:creator>Junfeng Huang</dc:creator><dc:creator>Sheng Meng</dc:creator><dc:creator>Miao Liu</dc:creator><category>arXiv</category><category>ai-materials</category><category>datasets-benchmarks</category><category>electronic-structure</category><category>materials-representation</category></item><item><title>arXiv: ATLAS: A Foundation Neural Sampler for Amorphous Materials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-19198</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-19198</guid><description>ATLAS trains an equivariant diffusion process to generate Boltzmann-distributed amorphous structures directly from a target energy function. The authors report less than 0.2% free-energy error in a low-temperature glass regime with more than 500-fold fewer energy evaluations than parallel tempering. A sampler that transfers across size, temperature, and composition could make amorphous-material thermodynamics and inverse design substantially more tractable. Preprint—not peer reviewed.</description><pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.19198</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.19198v1</dc:source><dc:creator>Mouyang Cheng</dc:creator><dc:creator>Denis Blessing</dc:creator><dc:creator>Botao Yu</dc:creator><dc:creator>Gerhard Neumann</dc:creator><dc:creator>Mingda Li</dc:creator><dc:creator>Carles Domingo-Enrich</dc:creator><dc:creator>Yuanqi Du</dc:creator><category>arXiv</category><category>ai-materials</category><category>foundation-models</category><category>generative-design</category></item><item><title>bioRxiv: BioPathfinder: Evidence-guided multi-agent platform enables hypothesis discovery for CAR-T engineering</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#biorxiv-10-64898-2026-07-15-738646</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#biorxiv-10-64898-2026-07-15-738646</guid><description>BioPathfinder builds a provenance-aware resource linking publications, patient single-cell RNA sequencing, and clinical metadata, then uses specialized language-model agents to generate, critique, and prioritize falsifiable hypotheses for computational and experimental validation. Applied to CAR-T patient evidence, the authors report that the workflow prioritized KLRC1/NKG2A and that in vitro and in vivo chronic-stimulation models linked NKG2A to exhaustion-associated CD8 CAR-T cells while blockade improved antitumor activity and persistence readouts. The platform connects fragmented clinical molecular evidence to testable engineering hypotheses and carries them through virtual perturbation, expert selection, and experimental validation. Preprint—not peer reviewed.</description><pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.15.738646</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.15.738646v2</dc:source><dc:creator>Wang, S.</dc:creator><dc:creator>Li, Y.-R.</dc:creator><dc:creator>Wang, Q.</dc:creator><dc:creator>Yang, Y.</dc:creator><dc:creator>Shen, X.</dc:creator><dc:creator>Li, H.</dc:creator><dc:creator>Nan, H.</dc:creator><dc:creator>Chen, Z.</dc:creator><dc:creator>Zhu, Y.</dc:creator><dc:creator>Zhang, B.</dc:creator><dc:creator>Ding, H.</dc:creator><dc:creator>Soto, J.</dc:creator><dc:creator>Park, S.</dc:creator><dc:creator>Zheng, Y.</dc:creator><dc:creator>Huang, X.</dc:creator><dc:creator>Li, D.</dc:creator><dc:creator>Li, S.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>scientific-agents</category></item><item><title>arXiv: Accelerated descriptor-free path sampling for protein-ligand binding kinetics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15101</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15101</guid><description>The method combines a shared descriptor-free equivariant committor with a basin-restricted bias that leaves the reactive region unbiased. The authors report rates consistent with references and experiments across roughly 17 orders of magnitude, together with unbinding mechanisms reconstructed without additional sampling. Minimal system-specific setup makes the approach a plausible foundation for comparing structure-kinetics relationships across ligand series. Preprint—not peer reviewed.</description><pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.15101</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.15101v1</dc:source><dc:creator>Simon M. Lichtinger</dc:creator><dc:creator>Roberto Covino</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>biomolecular-structure</category></item><item><title>arXiv: SevenNet-Polar for MultiTask Prediction of Energy, Forces, Stress, and Born Effective Charges: Development and Application to ZrO$_2$, Li$_3$PO$_4$, and Perovskites</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-14827</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-14827</guid><description>SevenNet-Polar extends an equivariant graph network to predict Born effective charges alongside energy, forces, and stress. The authors report multitask errors of 1.0 meV per atom for energy, 12 meV per angstrom for forces, 0.05 GPa for stress, and 0.0029 e for charge. Fast charge-aware potentials make large-scale molecular dynamics under electric fields accessible without separating polarization from the underlying mechanics. Preprint—not peer reviewed.</description><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.14827</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.14827v3</dc:source><dc:creator>Anh Khoa Augustin Lu</dc:creator><dc:creator>Shungo Arai</dc:creator><dc:creator>Yutack Park</dc:creator><dc:creator>Seungwu Han</dc:creator><dc:creator>Tsuyoshi Miyazaki</dc:creator><dc:creator>Satoshi Watanabe</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: STEP: Spin Tensor Equivariant Potential for Data-Efficient Learning of Magnetic Potential Energy Surfaces</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-17129</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-17129</guid><description>STEP treats vector magnetic moments as continuous geometric degrees of freedom and couples a central spin representation to its local spin-lattice environment through an equivariant tensor product. The authors report competitive or improved accuracy on FeAl, CrN, and Fe benchmarks, high-fidelity phonon and magnon spectra for CrI3, and Curie temperatures for CrI3 and bcc Fe in good agreement with experiment. Encoding spin-lattice symmetry directly into the potential makes data-efficient simulation of coupled magnetic excitations and finite-temperature behavior possible within one framework. Preprint—not peer reviewed.</description><pubDate>Sun, 19 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.17129</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.17129v1</dc:source><dc:creator>Yuanqing Gao</dc:creator><dc:creator>Wen-Hao Luo</dc:creator><dc:creator>Lei Zhang</dc:creator><dc:creator>Kun Cao</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15686</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15686</guid><description>S1-Omni maps language instructions, crystal and molecular encodings, protein sequences, spectra, and images into a shared representation, aligns that space with scientific laws and expert knowledge, and decodes it for domain-specific tasks. After training across 200 scientific tasks and evaluation on more than 60 benchmarks, the authors report that S1-Omni outperformed general-purpose comparison models on most tests and matched or exceeded specialized models on several. The common representation connects scientific understanding, prediction, and native generation across molecules, proteins, crystals, spectra, and images in one model. Preprint—not peer reviewed.</description><pubDate>Fri, 17 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.15686</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.15686v1</dc:source><dc:creator>Jiahao Zhao</dc:creator><dc:creator>Junyi Liu</dc:creator><dc:creator>Lifeng Xu</dc:creator><dc:creator>Nan Xu</dc:creator><dc:creator>Qingli Wang</dc:creator><dc:creator>Qingxiao Li</dc:creator><dc:creator>Tianle Chen</dc:creator><dc:creator>Xiaoyu Wu</dc:creator><dc:creator>Yawen Zheng</dc:creator><dc:creator>Zikai Wang</dc:creator><dc:creator>Guanming Liu</dc:creator><dc:creator>Hequn Zhou</dc:creator><dc:creator>Jingyi Wang</dc:creator><dc:creator>Jingyuan Shu</dc:creator><dc:creator>Keqi Wang</dc:creator><dc:creator>Li He</dc:creator><dc:creator>Songyang Diao</dc:creator><dc:creator>Wenhui Xu</dc:creator><dc:creator>Xinyu Ren</dc:creator><dc:creator>Yaqin Fan</dc:creator><dc:creator>Yujin Zhou</dc:creator><dc:creator>Zhanao Yao</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>ai-materials</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>bioRxiv: UniFlow: Unifying protein conformational ensemble generation and machine-learned force fields with a scalable normalizing Flow</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#biorxiv-10-64898-2026-07-17-739266</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#biorxiv-10-64898-2026-07-17-739266</guid><description>UniFlow uses an internal-coordinate normalizing flow to unify independent protein-ensemble sampling, exact likelihoods, and differentiable coarse-grained energies and forces. The authors report close agreement with reference molecular dynamics, transfer beyond the training proteins, and substantially faster sampling than diffusion-based ensemble models. Using one learned density for both equilibrium samples and stable dynamics closes a longstanding divide between generative modeling and molecular simulation. Preprint—not peer reviewed.</description><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.17.739266</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.17.739266v1</dc:source><dc:creator>Liu, Y.</dc:creator><dc:creator>Chen, M.</dc:creator><dc:creator>Lin, G.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-structure</category><category>neural-potentials</category></item><item><title>arXiv: Full-data accuracy with fewer labels for training and fine-tuning machine-learning force fields</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-14486</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-14486</guid><description>The active-learning workflow uses last-layer-projection regression as a cheap per-configuration uncertainty estimate for machine-learning force fields. The authors report that LLPR-selected subsets recovered full-data accuracy across molecular, condensed-phase, and electrolyte systems using only a small fraction of the electronic-structure labels. A forward-pass uncertainty measure avoids the separate fine-tuning runs that make model committees impractical for foundation force fields. Preprint—not peer reviewed.</description><pubDate>Wed, 15 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.14486</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.14486v1</dc:source><dc:creator>Sheng Bi</dc:creator><dc:creator>Yi-Ze Wang</dc:creator><dc:creator>Jun Cheng</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>neural-potentials</category></item><item><title>arXiv: Girsanov Reweighting for Uncertainty Propagation in Rare-Event Kinetics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-13757</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-13757</guid><description>The framework combines Adaptive Multilevel Splitting with Girsanov reweighting to propagate interatomic-potential parameter uncertainty into averaged committor probabilities without resampling reactive trajectories for each parameter realization. Across a Muller-Brown potential, a solvated dimer, and a butane conformational transition, the authors report recovery of reference rare-event probabilities and, under mild basin-accuracy assumptions, uncertainty bounds on reaction rates. Path-space reweighting provides uncertainty-aware committors and rate bounds from an existing rare-event trajectory ensemble, avoiding a new simulation for every potential realization. Preprint—not peer reviewed.</description><pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.13757</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.13757v2</dc:source><dc:creator>Leonard Moracchini</dc:creator><dc:creator>Thomas Pigeon</dc:creator><dc:creator>Morgane Menz</dc:creator><dc:creator>Thibault Faney</dc:creator><dc:creator>Thomas D. Swinburne</dc:creator><dc:creator>Mihai-Cosmin Marinica</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>neural-potentials</category><category>reaction-modeling</category></item><item><title>arXiv: Neural operators solve inverse problems for constitutive model discovery</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15049</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-22/#arxiv-2607-15049</guid><description>PANO and CANO map full-field displacement measurements and net reaction forces directly to hyperelastic strain-energy density functions, using Laplacian eigenfunctions and physically admissible output constraints. The authors report near-instantaneous constitutive inference in one forward pass and evaluate the operators on unseen, noisy, incomplete, differently discretized data and geometries of different sizes. The operator formulation makes constitutive discovery fast, discretization-independent, and constrained to physically admissible material responses. Preprint—not peer reviewed.</description><pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.15049</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.15049v1</dc:source><dc:creator>Moritz Flaschel</dc:creator><dc:creator>Burigede Liu</dc:creator><dc:creator>Ellen Kuhl</dc:creator><category>arXiv</category><category>ai-materials</category><category>surrogate-modeling</category></item><item><title>arXiv: Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-07708</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-07708</guid><description>SciReasoner discretizes molecular coordinates, topologies, and periodic connectivity into a shared vocabulary whose structural tokens remain addressable evidence during model reasoning. Across 86 benchmarks, the authors report state-of-the-art results on 67 tasks, and double-blind experts judged its reasoning traces preferred or comparable to a frontier language model in 98 percent of cases. A common but inspectable structural language could make predictions across proteins, molecules, and crystals easier to connect to the physical evidence that supports them. Preprint—not peer reviewed.</description><pubDate>Wed, 08 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.07708</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.07708v1</dc:source><dc:creator>Chen Tang</dc:creator><dc:creator>Yizhou Wang</dc:creator><dc:creator>Jianyu Wu</dc:creator><dc:creator>Lintao Wang</dc:creator><dc:creator>Shixiang Tang</dc:creator><dc:creator>Pengze Li</dc:creator><dc:creator>Encheng Su</dc:creator><dc:creator>Jun Yao</dc:creator><dc:creator>Jiabei Xiao</dc:creator><dc:creator>Yuqi Shi</dc:creator><dc:creator>Jielan Li</dc:creator><dc:creator>Hongxia Hao</dc:creator><dc:creator>Zhangyang Gao</dc:creator><dc:creator>Fang Wu</dc:creator><dc:creator>Ben Fei</dc:creator><dc:creator>Xiangyu Yue</dc:creator><dc:creator>Pan Tan</dc:creator><dc:creator>Bozitao Zhong</dc:creator><dc:creator>Jinouwen Zhang</dc:creator><dc:creator>Aoran Wang</dc:creator><dc:creator>Yan Lu</dc:creator><dc:creator>Jiaheng Liu</dc:creator><dc:creator>Xinzhu Ma</dc:creator><dc:creator>Liang Hong</dc:creator><dc:creator>Mingyue Zheng</dc:creator><dc:creator>Phil Torr</dc:creator><dc:creator>Bowen Zhou</dc:creator><dc:creator>Wanli Ouyang</dc:creator><dc:creator>Lei Bai</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>ai-materials</category><category>biomolecular-structure</category><category>foundation-models</category><category>materials-representation</category><category>molecular-representation</category></item><item><title>arXiv: Reaction-network reasoning with frontier models for experimentally confirmed catalyst-selectivity hypotheses</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-08003</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-08003</guid><description>The framework constrains frontier language models to reason over explicit reaction networks and couples their hypotheses to human review while preserving the network topology. For carbon-dioxide electroreduction, the authors report identification of selectivity-controlling pathways and levers that guided prospective synthesis of a copper-iron oxide catalyst with threefold higher acetate selectivity than matched copper-rich baselines. Reasoning over pathway competition can make an AI proposal experimentally useful by tying a materials choice to a mechanism that can be perturbed and tested. Preprint—not peer reviewed.</description><pubDate>Wed, 08 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.08003</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.08003v1</dc:source><dc:creator>Sutanay Choudhury</dc:creator><dc:creator>Anwesha Banerjee</dc:creator><dc:creator>Udishnu Sanyal</dc:creator><dc:creator>Jorin Dawidowicz</dc:creator><dc:creator>Chiezugolum Ijeoma Odilinye</dc:creator><dc:creator>Jesun Firoz</dc:creator><dc:creator>Liney Arnadottir</dc:creator><dc:creator>Simone Raugei</dc:creator><dc:creator>Johannes Lercher</dc:creator><dc:creator>Arnab Dutta</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>reaction-modeling</category><category>scientific-agents</category></item><item><title>arXiv: Transferable Implicit Solvent Machine Learning Potential for Drugs and Proteins Approaching Ab Initio Accuracy</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-10887</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-10887</guid><description>The Transferable Water Implicit Network represents aqueous environments with an equivariant graph-neural-network potential trained only on ab initio calculations and experimental labels. Across drug-like molecules, peptides, and proteins, the authors report better crystallographic and nuclear-magnetic-resonance results than earlier learned implicit-solvent or coarse-grained models, with timestep evaluation two orders of magnitude faster than explicit-solvent density-functional-theory potentials. A transferable implicit solvent at this accuracy and speed could extend first-principles-quality biomolecular simulation toward the longer timescales required in practice. Preprint—not peer reviewed.</description><pubDate>Sun, 12 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.10887</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.10887v1</dc:source><dc:creator>Jan Eckwert</dc:creator><dc:creator>Julija Zavadlav</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>bioRxiv: The projection basis determines the information ceiling for perturbation prediction</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#biorxiv-10-64898-2026-07-07-737004</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#biorxiv-10-64898-2026-07-07-737004</guid><description>The analysis derives an information-theoretic ceiling in which squared prediction-truth correlation is bounded by the variance explained by an orthonormal projection basis, so discarded signal cannot be recovered by model complexity. On chemical perturbations, the authors report that a gene-network eigenbasis captured only 10-12% of response variance while PCA captured 90-99%; graph wavelets recovered about 88%, and the ordering reversed for CRISPRa perturbations. Projection choice therefore determines which chemical or genetic response signal remains available to any downstream model. Preprint—not peer reviewed.</description><pubDate>Thu, 09 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.07.737004</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.07.737004v1</dc:source><dc:creator>BIANCO, S.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>datasets-benchmarks</category><category>molecular-representation</category></item><item><title>arXiv: Variable-Length Generative Protein Design via Generalized Poisson Flow</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09039</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09039</guid><description>Generalized Poisson Flow learns the rate function of an inhomogeneous counting process so protein length can be generated jointly with structure or sequence rather than fixed before sampling. The authors report exact recovery of the length distribution in unconditional design and first-place performance on 10 of 16 structure-based motif-scaffolding tasks, with more unique successes than fixed-length baselines. Allowing length to emerge from the design objective removes an artificial oracle from protein generation and enlarges the space of viable solutions. Preprint—not peer reviewed.</description><pubDate>Thu, 09 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.09039</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.09039v1</dc:source><dc:creator>Chaoran Cheng</dc:creator><dc:creator>Zhanghan Ni</dc:creator><dc:creator>Yanru Qu</dc:creator><dc:creator>Yuxin Chen</dc:creator><dc:creator>Ruihan Guo</dc:creator><dc:creator>Jiajun Fan</dc:creator><dc:creator>Ge Liu</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: Active rejection enables reliable generalization of universal machine-learning interatomic potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09456</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09456</guid><description>Adaptive Multi-Teacher Routing calibrates several pretrained interatomic potentials with a small set of high-fidelity labels, then accepts or rejects each proposed pseudo-label according to structure, teacher identity, and model disagreement. The authors report consistent gains over unrouted controls on held-out structures and stable finite-temperature trajectories in systems where baseline simulations collapse. Explicit rejection turns uncertainty from a descriptive score into a data-construction decision, which is the more useful role when one bad configuration can destabilize a long simulation. Preprint—not peer reviewed.</description><pubDate>Fri, 10 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.09456</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.09456v1</dc:source><dc:creator>Mingxiang Luo</dc:creator><dc:creator>Xinnan Mao</dc:creator><dc:creator>Lu Wang</dc:creator><dc:creator>Lei Bai</dc:creator><dc:creator>Feng Ding</dc:creator><dc:creator>Yuqiang Li</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>datasets-benchmarks</category><category>neural-potentials</category></item><item><title>arXiv: Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-12349</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-12349</guid><description>conDitar-dev combines a pretrained multiscale pocket representation, a pocket-conditioned diffusion model, and generation-time developability optimization to produce ligands with strong affinity and favorable ADMET properties. After synthesis and testing, the authors report two generated PD-L1 ligands with SPR-derived binding values of 3.49 and 3.75 micromolar and selective CSF1R inhibitors active at concentrations as low as 200 nanomolar. The modular design brings binding and developability into one generative workflow and carries selected candidates through synthesis and biological testing. Preprint—not peer reviewed.</description><pubDate>Tue, 14 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.12349</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.12349v1</dc:source><dc:creator>Ruoxi Gao</dc:creator><dc:creator>Jiangweizhi Peng</dc:creator><dc:creator>Ziqi Chen</dc:creator><dc:creator>Frazier N. Baker</dc:creator><dc:creator>David C. Kombo</dc:creator><dc:creator>John L. Kane</dc:creator><dc:creator>Andrew A. Scholte</dc:creator><dc:creator>Yi Li</dc:creator><dc:creator>Matthew J. LaMarche</dc:creator><dc:creator>Luigi I. Iconaru</dc:creator><dc:creator>Hans-Peter Biemann</dc:creator><dc:creator>Mingyi Hong</dc:creator><dc:creator>Xia Ning</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>bioRxiv: Hybrid quantum-classical de novo design of MHC-binding peptides</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#biorxiv-10-64898-2026-07-09-736951</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#biorxiv-10-64898-2026-07-09-736951</guid><description>The hybrid pipeline couples a generative adversarial network to latent vectors sampled from a photonic quantum processor to design MHC class I-binding peptides. Across 131 HLA alleles, the authors report that quantum-derived priors increased predicted strong-binder yield, with the largest gains for understudied alleles, and peptide-MHC stability ELISAs confirmed potent stabilizers among designs for three selected alleles. Hardware-derived nonclassical priors provide a structured way to widen sequence exploration when biological training data are sparse while preserving allele-specific anchor constraints. Preprint—not peer reviewed.</description><pubDate>Fri, 10 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.09.736951</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.09.736951v1</dc:source><dc:creator>Engdal, E. S.</dc:creator><dc:creator>Funk, J.</dc:creator><dc:creator>Bacarreza, O.</dc:creator><dc:creator>Machado, L.</dc:creator><dc:creator>Johansen, K. H.</dc:creator><dc:creator>Kemming, J.</dc:creator><dc:creator>Farnsworth, T.</dc:creator><dc:creator>Brasas, V.</dc:creator><dc:creator>Lefevre-Morand, R. Y. L.</dc:creator><dc:creator>Slysz, M.</dc:creator><dc:creator>Noerregaard, O. L.</dc:creator><dc:creator>Sandberg, O. A. D. A.</dc:creator><dc:creator>Makarovskiy, A.</dc:creator><dc:creator>Lodahl, P.</dc:creator><dc:creator>Acevedo-Rocha, C. G.</dc:creator><dc:creator>Kurowski, K.</dc:creator><dc:creator>Hadrup, S. R.</dc:creator><dc:creator>Clements, W. R.</dc:creator><dc:creator>Jenkins, T.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: The Precursor Genome: A Pairwise Reaction Dataset for Solid-State Synthesis</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09903</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-09903</guid><description>The Precursor Genome records 1,035 pairwise solid-state reactions executed by a self-driving laboratory, with thermal histories, masses, instrument settings, raw diffraction, refined structures, and reviewer annotations linked by provenance. The authors report 1,351 diffraction scans and 1,950 automated refinement cases, each evaluated by human experts on a three-tier quality scale. This combination of autonomous experiments and auditable raw-to-assignment records supplies the kind of training target that predictive models of solid-state reactivity have largely lacked. Preprint—not peer reviewed.</description><pubDate>Fri, 10 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.09903</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.09903v1</dc:source><dc:creator>Lauren N. Walters</dc:creator><dc:creator>Matthew J. McDermott</dc:creator><dc:creator>Bernardus Rendy</dc:creator><dc:creator>Yuxing Fei</dc:creator><dc:creator>Kristin A. Persson</dc:creator><dc:creator>Gerbrand Ceder</dc:creator><category>arXiv</category><category>ai-materials</category><category>autonomous-labs</category><category>datasets-benchmarks</category><category>reaction-modeling</category></item><item><title>arXiv: A vision foundation model for single-cell biology via spatial gene cartography</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-14163</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-15/#arxiv-2607-14163</guid><description>scVision uses optimal transport to place genes at fixed positions on a shared pan-tissue map, renders each transcriptome as a continuous image, and applies a masked-image-trained vision transformer as a frozen encoder. In zero-shot tests on six independent held-out studies, the authors report that scVision was the most accurate cell-type annotator, recovered gene programs without supervision, and lost sharply in accuracy when the gene layout was permuted. The fixed spatial map preserves gene relationships and expression magnitude in a representation that can reuse mature computer-vision methods across single-cell studies. Preprint—not peer reviewed.</description><pubDate>Tue, 14 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.14163</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.14163v1</dc:source><dc:creator>Ridvan Yesiloglu</dc:creator><dc:creator>Sakib Mostafa</dc:creator><dc:creator>James Zou</dc:creator><dc:creator>Ash Alizadeh</dc:creator><dc:creator>Jiajun Wu</dc:creator><dc:creator>Lei Xing</dc:creator><dc:creator>Ehsan Adeli</dc:creator><dc:creator>Md Tauhidul Islam</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>foundation-models</category><category>molecular-representation</category></item><item><title>arXiv: Agentic generation of verifiable rules for deterministic, self-expanding reaction classification</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01061</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01061</guid><description>A multi-agent language-model pipeline classifies patent reactions, writes deterministic rules, and tests every proposed rule against a corpus of 665,901 transformations. The authors report that the resulting system expands a 68-class taxonomy to 14,073 classes without human curation. A verified rule-writing loop turns reaction classification from a fixed taxonomy into a symbolic system that can extend itself when new chemistry appears. Preprint—not peer reviewed.</description><pubDate>Mon, 13 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.01061</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.01061v3</dc:source><dc:creator>Daniel Armstrong</dc:creator><dc:creator>Maarten Dobbelaere</dc:creator><dc:creator>Valentas Olikauskas</dc:creator><dc:creator>Helena Avila</dc:creator><dc:creator>Octavian Susanu</dc:creator><dc:creator>Jerome Waser</dc:creator><dc:creator>Philippe Schwaller</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>reaction-modeling</category><category>scientific-agents</category></item><item><title>arXiv: AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-03513</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-03513</guid><description>AquaGen samples all-atom molecular configurations with explicit solvent and periodic boundaries from a learned Boltzmann distribution, preserving compatibility with force-field evaluation and molecular dynamics refinement. For absolute hydration free energies, the authors report estimates with accuracy comparable to standard graphics-processor molecular dynamics at four- to tenfold lower cost. High-resolution ensemble generation retains the physical observables and post-processing hooks that are often lost when generative models remove solvent or coarse-grain the system. Preprint—not peer reviewed.</description><pubDate>Fri, 03 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.03513</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.03513v1</dc:source><dc:creator>Emmanuel Bengio</dc:creator><dc:creator>Sanjeev Raja</dc:creator><dc:creator>Yui Tik Pang</dc:creator><dc:creator>Kerstin Klaeser</dc:creator><dc:creator>Cristian Gabellini</dc:creator><dc:creator>Nikhil Shenoy</dc:creator><dc:creator>Francesco Di Giovanni</dc:creator><dc:creator>Prudencio Tossou</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>generative-design</category></item><item><title>arXiv: Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01408</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01408</guid><description>SpinGTP generalizes the Gaunt tensor product from scalar functions to spin-weighted spherical harmonics, restoring antisymmetric interactions while retaining its asymptotic efficiency. Across Tetris, 3BPA, SPICE-MACE-OFF, and OC20 benchmarks, the authors report accuracy comparable to full CGTP and better performance on chiral materials and non-centrosymmetric geometries. The construction provides a complete scalable equivariant basis for parity-sensitive interactions in large three-dimensional atomistic simulations. Preprint—not peer reviewed.</description><pubDate>Fri, 03 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.01408</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.01408v2</dc:source><dc:creator>Chenxing Liang</dc:creator><dc:creator>Yuchao Lin</dc:creator><dc:creator>Andrii Kryvenko</dc:creator><dc:creator>Wendi Yu</dc:creator><dc:creator>Chuan Li</dc:creator><dc:creator>Jianwen Xie</dc:creator><dc:creator>Xiaofeng Qian</dc:creator><dc:creator>Shuiwang Ji</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>molecular-representation</category></item><item><title>arXiv: EquiFiLM: Charge-Conditioned Equivariant Force Fields via Feature-wise Linear Modulation</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-05559</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-05559</guid><description>EquiFiLM adds continuous external-state conditioning to an equivariant foundation force field by modulating only scalar channels in each interaction layer. The authors report stable molecular dynamics across all tested charge states and prediction of the charge-dependent first-shell response measured by ultrafast electron diffraction. The lightweight conditioning axis offers a practical way to adapt ground-state foundation potentials to driven electronic processes without rebuilding them from scratch. Preprint—not peer reviewed.</description><pubDate>Mon, 06 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.05559</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.05559v1</dc:source><dc:creator>Samuel Sahel-Schackis</dc:creator><dc:creator>Ken-ichi Nomura</dc:creator><dc:creator>Aiichiro Nakano</dc:creator><dc:creator>Matthias F. Kling</dc:creator><dc:creator>Thomas Linker</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: MARLIN: De Novo Molecular Structure Elucidation from Tandem Mass Spectra without a Ground-Truth Formula</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-04774</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-04774</guid><description>MARLIN predicts a molecular fingerprint directly from tandem mass-spectrum peaks and uses a block-diffusion language model with an exact mass-shell constraint to generate structures without assuming a molecular formula. On NPLIB1, the authors report the strongest results among methods denied the ground-truth formula across exact match, structural distance, and fingerprint similarity, while recovering the correct formula about as often as a dedicated predictor. Removing the formula oracle places de novo structure elucidation closer to the untargeted setting where genuinely novel metabolites are first encountered. Preprint—not peer reviewed.</description><pubDate>Mon, 06 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.04774</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.04774v1</dc:source><dc:creator>Xujun Che</dc:creator><dc:creator>Xiuxia Du</dc:creator><dc:creator>Depeng Xu</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>molecular-representation</category></item><item><title>bioRxiv: Breaking the Synthesis Barrier for AI-Designed DNA Libraries</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#biorxiv-10-64898-2026-07-07-736931</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#biorxiv-10-64898-2026-07-07-736931</guid><description>Policy Gradients for Library Design optimizes a synthesis-aware parameterization of stochastic DNA libraries against a chosen sequence objective. The authors report that the method supports multi-round lab-in-the-loop design and was used to synthesize a large influenza-antibody sequence library for about 700 dollars. This formulation lets generative sequence models exploit the scale of multiplexed experiments without requiring every proposed sequence to be synthesized separately. Preprint—not peer reviewed.</description><pubDate>Tue, 07 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.07.736931</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.07.736931v1</dc:source><dc:creator>Sussex, S.</dc:creator><dc:creator>Borevkovic, E.</dc:creator><dc:creator>Lohmann, F.</dc:creator><dc:creator>Chen, N.</dc:creator><dc:creator>Lüthi, E.</dc:creator><dc:creator>Reddy, S. T.</dc:creator><dc:creator>Krause, A.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: A Physics-Regulated Neural Framework for Learning 3D Grain Growth Dynamics</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-04680</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-04680</guid><description>3D-PRIMME learns a physics-regulated local evolution rule for three-dimensional grain growth from only two consecutive time steps. After training on a 100^3 grid with 512 grains, the authors report applying the operator without retraining to 1024^3 grids with 550,000 grains while preserving coarsening kinetics and grain topology. The local rule allows grain-growth surrogates to move several orders of magnitude in system size without sacrificing the physical statistics they were trained to reproduce. Preprint—not peer reviewed.</description><pubDate>Mon, 06 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.04680</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.04680v1</dc:source><dc:creator>Zhihui Tian</dc:creator><dc:creator>Kang Yang</dc:creator><dc:creator>Michael Tonks</dc:creator><dc:creator>Amanda R. Krause</dc:creator><dc:creator>Joel B. Harley</dc:creator><category>arXiv</category><category>ai-materials</category><category>multiscale-modeling</category><category>surrogate-modeling</category></item><item><title>arXiv: Complex crystal structure prediction using ML-enhanced multi-minima iterative genetic algorithm</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01004</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-01004</guid><description>The multi-minima iterative genetic algorithm couples an artificial-neural-network interatomic potential to a metadynamics-inspired penalty that steers search away from previously explored basins. The authors report recovery of the synthesized ground-state structure of La4Co4Pb and an exact match to the independently measured crystal structure of La5CoPb2 using composition alone. Navigating both the global minimum and nearby metastable states gives data-driven structure prediction a route beyond recombining entries from known crystal databases. Preprint—not peer reviewed.</description><pubDate>Wed, 01 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.01004</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.01004v1</dc:source><dc:creator>Ling Tang</dc:creator><dc:creator>Weiyi Xia</dc:creator><dc:creator>Tyler J. Slade</dc:creator><dc:creator>Paul C. Canfield</dc:creator><dc:creator>Cai-Zhuang Wang</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>generative-design</category><category>neural-potentials</category></item><item><title>bioRxiv: Improving Generalizability in Whole-Cell Antibiotic Discovery Through Active Learning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#biorxiv-10-64898-2026-07-04-736489</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#biorxiv-10-64898-2026-07-04-736489</guid><description>The study compares three active-learning policies for whole-cell bacterial bioactivity, selects a balance between predicted novelty and hit rate in retrospective tuberculosis data, and then closes the loop in a Borrelia antibiotic screen. The authors report a five-fold hit-rate increase from 0.2% to 1.0% in the closed-loop screen, 53-fold enrichment with 11.0% validation in prospective selections, and intended narrow-spectrum activity for all validated hits. The acquisition strategy couples exploration and exploitation to experimental feedback, allowing a whole-cell predictor to generalize into chemically diverse, out-of-distribution screens. Preprint—not peer reviewed.</description><pubDate>Sun, 05 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.07.04.736489</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.07.04.736489v1</dc:source><dc:creator>Serrano, L. R.</dc:creator><dc:creator>Zhou, A.</dc:creator><dc:creator>Wei, Z.</dc:creator><dc:creator>Stocks, K.-L. K.</dc:creator><dc:creator>Ektefaie, Y.</dc:creator><dc:creator>Gwynne, P. J.</dc:creator><dc:creator>Chen, E.</dc:creator><dc:creator>Krieger, I.</dc:creator><dc:creator>Sacchettini, J.</dc:creator><dc:creator>Aldridge, B.</dc:creator><dc:creator>Hu, L. T.</dc:creator><dc:creator>Farhat, M. R.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>molecular-representation</category></item><item><title>arXiv: Dyna-Mat: End-to-end benchmarking of foundation machine learning interatomic potentials in finite-temperature ensembles</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-03433</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-08/#arxiv-2607-03433</guid><description>Dyna-Mat evaluates 15 foundation interatomic potentials against finite-temperature first-principles trajectories, comparing static energy and force errors with structural and dynamical observables from model-driven simulations. The authors report that low force error usually tracks better ensemble behavior but can still coincide with qualitative structural failure, while pressure remains poorly described across most models. The benchmark makes trajectory-level physical behavior, rather than a favorable static error alone, the standard for judging a deployable potential. Preprint—not peer reviewed.</description><pubDate>Fri, 03 Jul 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2607.03433</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2607.03433v1</dc:source><dc:creator>Mikołaj J. Gawkowski</dc:creator><dc:creator>Nongnuch Artrith</dc:creator><dc:creator>Silvia Bonfanti</dc:creator><dc:creator>Abhijeet Sadashiv Gangan</dc:creator><dc:creator>Hendrik H. Heenen</dc:creator><dc:creator>Joseph Kioseoglou</dc:creator><dc:creator>Ivor Lončarić</dc:creator><dc:creator>Hemanadhan Myneni</dc:creator><dc:creator>Janosh Riebesell</dc:creator><dc:creator>Mariana Rossi</dc:creator><dc:creator>Matthias Rupp</dc:creator><dc:creator>Jonathan Schmidt</dc:creator><dc:creator>Shubham Sharma</dc:creator><dc:creator>Benjamin X. Shi</dc:creator><dc:creator>Antoni Wadowski</dc:creator><dc:creator>Lukas Hörmann</dc:creator><dc:creator>Venkat Kapil</dc:creator><category>arXiv</category><category>ai-materials</category><category>atomistic-modeling</category><category>datasets-benchmarks</category><category>neural-potentials</category></item><item><title>arXiv: Contrastive Regularization of Machine Learning Potentials</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-31660</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-31660</guid><description>Contrastive Regularized MSE adds a distribution-aware term to ordinary energy-and-force fitting and uses persistent Langevin samples from the potential itself to expose configurations that should be raised in energy. On ethanol and aspirin, the authors report that the correction restores energy, distance, and free-energy distributions to near-quantitative agreement with density functional theory while preserving force accuracy. The result makes a useful point for molecular simulation: a potential meant to generate trajectories must be trained against the distribution it produces. Preprint—not peer reviewed.</description><pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.31660</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.31660v1</dc:source><dc:creator>Dimitrios Tzivrailis</dc:creator><dc:creator>Georgios Sotiropoulos</dc:creator><dc:creator>Alberto Rosso</dc:creator><dc:creator>Eiji Kawasaki</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>atomistic-modeling</category><category>neural-potentials</category></item><item><title>arXiv: Autoregressive Boltzmann Generators</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-27361</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-27361</guid><description>Autoregressive Boltzmann Generators replace flow-based Boltzmann generators with an autoregressive framework that avoids flow topology constraints and permits sequential interventions. The authors report that their 132-million-parameter transferable model reduces zero-shot energy error by more than 60% on 8-residue systems. The framework offers a likelihood-bearing route to more scalable equilibrium sampling for molecular systems. Preprint—not peer reviewed.</description><pubDate>Thu, 25 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.27361</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.27361v1</dc:source><dc:creator>Danyal Rehman</dc:creator><dc:creator>Charlie B. Tan</dc:creator><dc:creator>Yoshua Bengio</dc:creator><dc:creator>Avishek Joey Bose</dc:creator><dc:creator>Alexander Tong</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>generative-design</category></item><item><title>arXiv: Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-29776</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-29776</guid><description>NMRAgent combines specialized spectral-analysis tools with chemical knowledge graphs to plan structure elucidation, test peak-to-atom consistency, and refine candidate structures. The authors report a 46.5% improvement in top-1 accuracy and a 0.502 improvement in Tanimoto similarity on a scaffold-split benchmark. The resulting workflow makes molecular-structure proposals more inspectable and correctable when the test scaffolds are novel. Preprint—not peer reviewed.</description><pubDate>Mon, 29 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.29776</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.29776v1</dc:source><dc:creator>Zheng Fang</dc:creator><dc:creator>Chen Yang</dc:creator><dc:creator>Yusen Tan</dc:creator><dc:creator>Yunpeng Zhao</dc:creator><dc:creator>Fanjie Xu</dc:creator><dc:creator>Hongxin Xiang</dc:creator><dc:creator>Hanyu Sun</dc:creator><dc:creator>Hanyu Gao</dc:creator><dc:creator>Xiaojian Wang</dc:creator><dc:creator>Wenjie Du</dc:creator><dc:creator>Yuqiang Li</dc:creator><dc:creator>Jun Xia</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>molecular-representation</category><category>scientific-agents</category></item><item><title>bioRxiv: BoltzProt-1: Towards Efficient De Novo Binder Design with Good Developability</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#biorxiv-10-64898-2026-06-23-733997</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#biorxiv-10-64898-2026-06-23-733997</guid><description>BoltzProt-1 combines a refined generative binder model with BoltzPPI, a protein-interaction predictor used to rank designed nanobodies. Across ten novel targets, the authors report that confirmed-binder hit rates rise from 3.3 percent to 8.0 percent, while 58 percent of confirmed designs pass every measured developability criterion. Ranking designs by interaction quality rather than structure-prediction confidence connects de novo generation to two practical experimental bottlenecks: finding real binders and keeping the successful ones developable. Preprint—not peer reviewed.</description><pubDate>Sat, 27 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.06.23.733997</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.06.23.733997v1</dc:source><dc:creator>Ucar, T.</dc:creator><dc:creator>Bates, J.</dc:creator><dc:creator>Fu, Y.</dc:creator><dc:creator>Shi, W.</dc:creator><dc:creator>Stark, H.</dc:creator><dc:creator>Nava, D.</dc:creator><dc:creator>Cavalleri, L.</dc:creator><dc:creator>Wohlwend, J.</dc:creator><dc:creator>Corso, G.</dc:creator><dc:creator>Passaro, S.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>biomolecular-design</category><category>generative-design</category><category>protein-engineering</category></item><item><title>arXiv: Shape-Constrained Bayesian Active Learning of Self-Limiting Saturation Curves</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-26577</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-26577</guid><description>This active-learning platform uses Bayesian monotonic I-spline regression so each posterior saturation curve rises from zero and never decreases. The authors report that it reaches noise-floor accuracy within a 20-measurement budget in every regime, in as few as seven measurements. The same shape-constrained surrogate can make sparse-data experiment selection physically consistent across self-limiting chemical and materials responses. Preprint—not peer reviewed.</description><pubDate>Wed, 24 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.26577</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.26577v1</dc:source><dc:creator>Pouyan Navabi</dc:creator><dc:creator>Christos G. Takoudis</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>surrogate-modeling</category></item><item><title>arXiv: Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-28578</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-28578</guid><description>The workflow inserts a Gaussian-process acquisition gate between crystal generation and a property oracle in an RL-steered materials-design loop. The authors report that the gate comes within about 9% of exhaustive oracle spending at roughly one-fifth of the calls, while a density-functional-theory check confirms bulk-modulus predictions within 2.5% on average. This lets generative materials searches spend expensive calculations where a surrogate expects them to be most useful. Preprint—not peer reviewed.</description><pubDate>Fri, 26 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.28578</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.28578v1</dc:source><dc:creator>Sk Md Ahnaf Akif Alvi</dc:creator><dc:creator>Jan Janssen</dc:creator><dc:creator>Danny Perez</dc:creator><dc:creator>Douglas Allaire</dc:creator><dc:creator>Raymundo Arroyave</dc:creator><category>arXiv</category><category>ai-materials</category><category>generative-design</category><category>surrogate-modeling</category></item><item><title>bioRxiv: Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#biorxiv-10-64898-2026-06-29-735309</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#biorxiv-10-64898-2026-06-29-735309</guid><description>The authors introduce the Novelty-Tiered Affinity Benchmark, which partitions test data into ligand novelty tiers to control target-mirroring leakage. They report that a ChEMBL 36 meta-analysis identifies more than 6,000 such assay pairs and that ligand-only models fall to r = 0.14 in the most challenging tier. The benchmark gives affinity models a more direct test of whether they generalize beyond localized leakage and memorized training data. Preprint—not peer reviewed.</description><pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>10.64898/2026.06.29.735309</dc:identifier><dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights><dc:source>https://www.biorxiv.org/content/10.64898/2026.06.29.735309v1</dc:source><dc:creator>Mattsson, B.</dc:creator><dc:creator>Walters, W.</dc:creator><category>bioRxiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>biomolecular-structure</category><category>datasets-benchmarks</category></item><item><title>arXiv: Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-29459</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-29459</guid><description>LLM4MOF assigns one language-model agent to propose chemically interpretable design hypotheses and another to translate those hypotheses into constrained metal-organic framework candidates for simulation and revision. The authors report that ten autonomous iterations concentrate searches on strong candidates across six tasks within 400 property evaluations and outperform random search and a genetic algorithm during de novo simulated design. Because each candidate remains tied to explicit choices about nodes, linkers, pores, and functional chemistry, the loop can expose why a design succeeds instead of returning only a score. Preprint—not peer reviewed.</description><pubDate>Sun, 28 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.29459</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.29459v1</dc:source><dc:creator>Kyungmin Nam</dc:creator><dc:creator>Seunghee Han</dc:creator><dc:creator>Jihan Kim</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>ai-materials</category><category>generative-design</category><category>scientific-agents</category></item><item><title>arXiv: Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-30687</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-30687</guid><description>The paper defines an action-operator framework for molecular diffusion models and uses a noisy operator bridge to read out free-energy differences from endpoint ensembles. The authors report that endpoint coordinates and binary labels alone are sufficient to partially recover the operator shape and a centered free-energy scale without force or action supervision. This provides a route for turning diffusion models from coordinate samplers into thermodynamic estimators with an explicit physical interpretation. Preprint—not peer reviewed.</description><pubDate>Sun, 28 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.30687</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.30687v1</dc:source><dc:creator>Wenjie Xi</dc:creator><category>arXiv</category><category>ai-biochemistry</category><category>ai-chemistry</category><category>atomistic-modeling</category><category>generative-design</category></item><item><title>arXiv: ReactionAtlas: Ab origine exploration of chemical reaction networks with machine learning</title><link>https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-30778</link><guid isPermaLink="true">https://brettsavoie.com/projects/useful-chemistry/issues/2026-07-01/#arxiv-2606-30778</guid><description>ReactionAtlas builds reaction networks from a handful of seed molecules without hand-crafted rules, using a generative model to propose reactions and a DFT-trained machine-learned force field to filter valid transition states. The authors report roughly 47,000 reactions among roughly 12,000 compounds from eight pre-biotic seeds, with 85% of machine-learned transition states within 0.5 Å RMSD of PBE0 references. This combination makes deep network exploration more tractable when the relevant products and transition states are not known in advance. Preprint—not peer reviewed.</description><pubDate>Mon, 29 Jun 2026 12:00:00 GMT</pubDate><dc:identifier>arxiv:2606.30778</dc:identifier><dc:rights>https://creativecommons.org/publicdomain/zero/1.0/</dc:rights><dc:source>https://arxiv.org/abs/2606.30778v1</dc:source><dc:creator>Stefan Gugler</dc:creator><dc:creator>Max Eissler</dc:creator><dc:creator>Khaled Kahouli</dc:creator><dc:creator>Klaus-Robert Müller</dc:creator><category>arXiv</category><category>ai-chemistry</category><category>generative-design</category><category>neural-potentials</category><category>reaction-modeling</category></item></channel></rss>