Mouyang Cheng, Denis Blessing, Botao Yu, Gerhard Neumann, Mingda Li, Carles Domingo-Enrich, Yuanqi Du
ATLAS trains an equivariant diffusion process to generate Boltzmann-distributed amorphous structures directly from a target energy function. The authors report less than 0.2% free-energy error in a low-temperature glass regime with more than 500-fold fewer energy evaluations than parallel tempering. A sampler that transfers across size, temperature, and composition could make amorphous-material thermodynamics and inverse design substantially more tractable.
Dimitrios Tzivrailis, Georgios Sotiropoulos, Alberto Rosso, Eiji Kawasaki
Contrastive Regularized MSE adds a distribution-aware term to ordinary energy-and-force fitting and uses persistent Langevin samples from the potential itself to expose configurations that should be raised in energy. On ethanol and aspirin, the authors report that the correction restores energy, distance, and free-energy distributions to near-quantitative agreement with density functional theory while preserving force accuracy. The result makes a useful point for molecular simulation: a potential meant to generate trajectories must be trained against the distribution it produces.
Jiahui Du, Zikai Xie, Man Luo, Liang Zhang, Hengtao Lei, Fangming Gao, Chunxing Yan, Chengxi Zhao, Xini Chu, Ziyi Cheng, Ziyi Jin, Bing Huang, Xijun Wang, Linjiang Chen, Hexiang Deng, Jun Jiang, Yi Luo
MiMEDAL combines adaptive experimentation, symbolic regression, LLM reasoning, and first-principles falsification to refine physical mechanisms from sparse data. The authors report that the autonomous system stopped after 25 experiments, reached 96.1% accuracy across 103 unseen COFs, and guided synthesis of a COF with a 61% solid-state photoluminescence quantum yield. The approach matters because it joins data-efficient experimentation to physical falsification, offering a route from statistical prediction to transferable mechanism discovery.
Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai
SciReasoner discretizes molecular coordinates, topologies, and periodic connectivity into a shared vocabulary whose structural tokens remain addressable evidence during model reasoning. Across 86 benchmarks, the authors report state-of-the-art results on 67 tasks, and double-blind experts judged its reasoning traces preferred or comparable to a frontier language model in 98 percent of cases. A common but inspectable structural language could make predictions across proteins, molecules, and crystals easier to connect to the physical evidence that supports them.
BioPathfinder builds a provenance-aware resource linking publications, patient single-cell RNA sequencing, and clinical metadata, then uses specialized language-model agents to generate, critique, and prioritize falsifiable hypotheses for computational and experimental validation. Applied to CAR-T patient evidence, the authors report that the workflow prioritized KLRC1/NKG2A and that in vitro and in vivo chronic-stimulation models linked NKG2A to exhaustion-associated CD8 CAR-T cells while blockade improved antitumor activity and persistence readouts. The platform connects fragmented clinical molecular evidence to testable engineering hypotheses and carries them through virtual perturbation, expert selection, and experimental validation.
The framework constrains frontier language models to reason over explicit reaction networks and couples their hypotheses to human review while preserving the network topology. For carbon-dioxide electroreduction, the authors report identification of selectivity-controlling pathways and levers that guided prospective synthesis of a copper-iron oxide catalyst with threefold higher acetate selectivity than matched copper-rich baselines. Reasoning over pathway competition can make an AI proposal experimentally useful by tying a materials choice to a mechanism that can be perturbed and tested.
Suyang Zhong, Yuhao Zhao, Boying Huang, Fanjie Xu, Pengwei Xu, Haoyi Tao, Xi Fang, Jun Cheng, Fujie Tang
Uni-XAS treats spectra and atomic structures as a shared alignment-and-generation problem, with retrieval-anchored forward decoding and permutation-rectified inverse flow matching. The authors report strong cross-modal retrieval, absolute-spectrum prediction, and composition-conditioned three-dimensional structure generation on 328,839 paired structures and spectra. A common latent space makes forward and inverse spectroscopy mutually informative rather than two disconnected regressions.
Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas, Helena Avila, Octavian Susanu, Jerome Waser, Philippe Schwaller
A multi-agent language-model pipeline classifies patent reactions, writes deterministic rules, and tests every proposed rule against a corpus of 665,901 transformations. The authors report that the resulting system expands a 68-class taxonomy to 14,073 classes without human curation. A verified rule-writing loop turns reaction classification from a fixed taxonomy into a symbolic system that can extend itself when new chemistry appears.
Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin
MACROS uses multiple agents to automate iterative hypothesis testing for molecular structure elucidation from combinations of routine spectra. The authors report zero-shot identification of diverse real-world samples above 500 daltons with one-dimensional NMR, as well as faster and more accurate elucidation through chemist collaboration. This matters because scalable spectral reasoning could move automated structure elucidation beyond fixed reference-library matching.
Emmanuel Bengio, Sanjeev Raja, Yui Tik Pang, Kerstin Klaeser, Cristian Gabellini, Nikhil Shenoy, Francesco Di Giovanni, Prudencio Tossou
AquaGen samples all-atom molecular configurations with explicit solvent and periodic boundaries from a learned Boltzmann distribution, preserving compatibility with force-field evaluation and molecular dynamics refinement. For absolute hydration free energies, the authors report estimates with accuracy comparable to standard graphics-processor molecular dynamics at four- to tenfold lower cost. High-resolution ensemble generation retains the physical observables and post-processing hooks that are often lost when generative models remove solvent or coarse-grain the system.
Dengpan Dong, Yani Guan, Shuang Luo, Jingxuan Ding, Dan C. Hannah, Yumin Zhang, Qichao Hu, Kang Xu
The platform uses a message-passing neural network to predict a complete polarizable-force-field parameter set from molecular structure and couples that model to automated system building and trajectory analysis. The authors report density errors below 0.02 g/cm3, ionic conductivity near 1.5 mS/cm in agreement with measurement, and solubility predictions within 10%. Reducing force-field parameterization from weeks of expert work to a single forward pass could make high-throughput molecular dynamics practical across solvents, salts, and additives.
The method combines a shared descriptor-free equivariant committor with a basin-restricted bias that leaves the reactive region unbiased. The authors report rates consistent with references and experiments across roughly 17 orders of magnitude, together with unbinding mechanisms reconstructed without additional sampling. Minimal system-specific setup makes the approach a plausible foundation for comparing structure-kinetics relationships across ligand series.
The Transferable Water Implicit Network represents aqueous environments with an equivariant graph-neural-network potential trained only on ab initio calculations and experimental labels. Across drug-like molecules, peptides, and proteins, the authors report better crystallographic and nuclear-magnetic-resonance results than earlier learned implicit-solvent or coarse-grained models, with timestep evaluation two orders of magnitude faster than explicit-solvent density-functional-theory potentials. A transferable implicit solvent at this accuracy and speed could extend first-principles-quality biomolecular simulation toward the longer timescales required in practice.
Fondrie, W. E., Canzani, D., Tatka, L., Paez, J. S., Prymolenna, A., Gutierrez, A., Robbins, J., McEllin, B., Hubbard, E., Siebenthall, K., Pino, L. K., Federation, A. J.
Ptarmigan-1 contrastively co-embeds protein residues and candidate small molecules from sequence and two-dimensional chemistry, avoiding explicit pose construction. The authors report that it matches or exceeds docking and co-folding models at covalent, cryptic, and disordered sites, and screens 3.4 billion compounds across the human proteome in under a day. A residue-resolved, pose-free index could extend virtual screening to targets where structural pockets are unavailable or poorly defined.
MatDiffract combines perturbation-augmented simulated patterns, multiscale vector retrieval, Rietveld refinement, and quantitative phase fitting in one automated workflow. The authors report 91.3% top-1 and 97.2% top-10 identification accuracy after automated refinement on 875 experimental single-phase patterns. Returning refined structures and quantitative compositions within seconds addresses a practical characterization bottleneck in high-throughput and autonomous materials work.
Danyal Rehman, Charlie B. Tan, Yoshua Bengio, Avishek Joey Bose, Alexander Tong
Autoregressive Boltzmann Generators replace flow-based Boltzmann generators with an autoregressive framework that avoids flow topology constraints and permits sequential interventions. The authors report that their 132-million-parameter transferable model reduces zero-shot energy error by more than 60% on 8-residue systems. The framework offers a likelihood-bearing route to more scalable equilibrium sampling for molecular systems.
Vilya Research, :, Pascal Sturmfels, Naozumi Hiranuma, Milad Salem, Benjamin D. Sellers, Stephen Rettie, CJ San Felipe, Chase A. P. Wood, Jeffrey K. Holden, Adam P. Moyer, Patrick J. Salveson, Ivan Anishchanka
Vilya-2 extends an all-atom molecular representation from individual molecules to a diffusion transformer that models their interactions with protein targets. The authors report that calibrated structural ensembles recover 59.1% of peptide interfaces below 2 angstrom backbone RMSD, while the model also reaches state-of-the-art small-molecule docking and transfers to complex types unlike those in training. A common all-atom representation can support structure prediction across chemically distinct interface classes and then be fine-tuned for hit-to-lead enrichment.
SpinGTP generalizes the Gaunt tensor product from scalar functions to spin-weighted spherical harmonics, restoring antisymmetric interactions while retaining its asymptotic efficiency. Across Tetris, 3BPA, SPICE-MACE-OFF, and OC20 benchmarks, the authors report accuracy comparable to full CGTP and better performance on chiral materials and non-centrosymmetric geometries. The construction provides a complete scalable equivariant basis for parity-sensitive interactions in large three-dimensional atomistic simulations.
NMRAgent combines specialized spectral-analysis tools with chemical knowledge graphs to plan structure elucidation, test peak-to-atom consistency, and refine candidate structures. The authors report a 46.5% improvement in top-1 accuracy and a 0.502 improvement in Tanimoto similarity on a scaffold-split benchmark. The resulting workflow makes molecular-structure proposals more inspectable and correctable when the test scaffolds are novel.
UltraIR pretrains a large infrared foundation model on simulated spectra and adapts it to multiple chemical-sensing tasks. The authors report stronger performance than conventional and task-specific learning across analytical settings, including limited-label and cross-instrument tests. This matters because a shared spectral representation could reduce the data burden of deploying infrared analysis across laboratories and sample types.
Anh Khoa Augustin Lu, Shungo Arai, Yutack Park, Seungwu Han, Tsuyoshi Miyazaki, Satoshi Watanabe
SevenNet-Polar extends an equivariant graph network to predict Born effective charges alongside energy, forces, and stress. The authors report multitask errors of 1.0 meV per atom for energy, 12 meV per angstrom for forces, 0.05 GPa for stress, and 0.0029 e for charge. Fast charge-aware potentials make large-scale molecular dynamics under electric fields accessible without separating polarization from the underlying mechanics.
The analysis derives an information-theoretic ceiling in which squared prediction-truth correlation is bounded by the variance explained by an orthonormal projection basis, so discarded signal cannot be recovered by model complexity. On chemical perturbations, the authors report that a gene-network eigenbasis captured only 10-12% of response variance while PCA captured 90-99%; graph wavelets recovered about 88%, and the ordering reversed for CRISPRa perturbations. Projection choice therefore determines which chemical or genetic response signal remains available to any downstream model.
STEP treats vector magnetic moments as continuous geometric degrees of freedom and couples a central spin representation to its local spin-lattice environment through an equivariant tensor product. The authors report competitive or improved accuracy on FeAl, CrN, and Fe benchmarks, high-fidelity phonon and magnon spectra for CrI3, and Curie temperatures for CrI3 and bcc Fe in good agreement with experiment. Encoding spin-lattice symmetry directly into the potential makes data-efficient simulation of coupled magnetic excitations and finite-temperature behavior possible within one framework.
Chaoran Cheng, Zhanghan Ni, Yanru Qu, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu
Generalized Poisson Flow learns the rate function of an inhomogeneous counting process so protein length can be generated jointly with structure or sequence rather than fixed before sampling. The authors report exact recovery of the length distribution in unconditional design and first-place performance on 10 of 16 structure-based motif-scaffolding tasks, with more unique successes than fixed-length baselines. Allowing length to emerge from the design objective removes an artificial oracle from protein generation and enlarges the space of viable solutions.
Mingxiang Luo, Xinnan Mao, Lu Wang, Lei Bai, Feng Ding, Yuqiang Li
Adaptive Multi-Teacher Routing calibrates several pretrained interatomic potentials with a small set of high-fidelity labels, then accepts or rejects each proposed pseudo-label according to structure, teacher identity, and model disagreement. The authors report consistent gains over unrouted controls on held-out structures and stable finite-temperature trajectories in systems where baseline simulations collapse. Explicit rejection turns uncertainty from a descriptive score into a data-construction decision, which is the more useful role when one bad configuration can destabilize a long simulation.
Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen, Frazier N. Baker, David C. Kombo, John L. Kane, Andrew A. Scholte, Yi Li, Matthew J. LaMarche, Luigi I. Iconaru, Hans-Peter Biemann, Mingyi Hong, Xia Ning
conDitar-dev combines a pretrained multiscale pocket representation, a pocket-conditioned diffusion model, and generation-time developability optimization to produce ligands with strong affinity and favorable ADMET properties. After synthesis and testing, the authors report two generated PD-L1 ligands with SPR-derived binding values of 3.49 and 3.75 micromolar and selective CSF1R inhibitors active at concentrations as low as 200 nanomolar. The modular design brings binding and developability into one generative workflow and carries selected candidates through synthesis and biological testing.
Zhuotao Jin, Mohamed S. Abdallah, Boris Kozinsky, Kyle Bystrom
RS-CIDER fits a non-local exchange functional to ground-state energies and single-particle levels, replacing the short-range Hartree-Fock term of HSE06 during self-consistent evaluation. The authors report agreement with HSE06 for reaction energies and band gaps, with a measured per-self-consistent-field-step wall time more than an order of magnitude lower. The work puts a more costly hybrid-functional reference within reach for larger periodic calculations without dropping self-consistency.
Lin Huang, Frank Peng, JiaJun Cheng, Zion Wang, Hao Yin, Hao Li, Ji Zhang, Jack Jia, Junping Zhao, Arthur Jiang,, Jia Zhang
UBio-MolFM is a foundation model trained on 160 million quantum-chemical labels, designed with a receptive field spanning noncovalent distances at near-linear cost. The authors report force errors near 20 meV/Å past a thousand atoms and a 108,964-atom KcsA channel simulation that formed the anhydrous knock-on geometry in four of five replicas. Extending quantum-trained modeling to such system sizes can make electronic structure accessible where fixed-charge approximations obscure the mechanism.
Steffen Wedig, Felix Burton, Rokas Elijo\vsius, Christoph Schran,, Lars L. Schaaf
Rem3Di pools per-atom latent features from atomistic foundation models into smooth, order-invariant whole-molecule descriptors and adds pseudoscalar features that reverse sign under reflection to encode chirality. Across public drug-property benchmarks, the authors report that Rem3Di matched or exceeded published baselines without classical two-dimensional fingerprints and separated transition-metal complexes without predefined bonding rules. The representation carries information learned from quantum-mechanical simulation into property prediction and virtual screening while preserving three-dimensional structure and molecular handedness.
Samuel Sahel-Schackis, Ken-ichi Nomura, Aiichiro Nakano, Matthias F. Kling, Thomas Linker
EquiFiLM adds continuous external-state conditioning to an equivariant foundation force field by modulating only scalar channels in each interaction layer. The authors report stable molecular dynamics across all tested charge states and prediction of the charge-dependent first-shell response measured by ultrafast electron diffraction. The lightweight conditioning axis offers a practical way to adapt ground-state foundation potentials to driven electronic processes without rebuilding them from scratch.
Ucar, T., Bates, J., Fu, Y., Shi, W., Stark, H., Nava, D., Cavalleri, L., Wohlwend, J., Corso, G., Passaro, S.
BoltzProt-1 combines a refined generative binder model with BoltzPPI, a protein-interaction predictor used to rank designed nanobodies. Across ten novel targets, the authors report that confirmed-binder hit rates rise from 3.3 percent to 8.0 percent, while 58 percent of confirmed designs pass every measured developability criterion. Ranking designs by interaction quality rather than structure-prediction confidence connects de novo generation to two practical experimental bottlenecks: finding real binders and keeping the successful ones developable.
TransTS learns atom-level reaction transformations together with aligned reactant, transition-state, and product geometries to generate transition-state starting structures. The authors report more frequent convergence to validated saddle points and intended elementary reactions on challenging out-of-distribution benchmarks. This matters because chemically informed initial guesses can lower the quantum-chemical cost of mechanistic modeling.
MARLIN predicts a molecular fingerprint directly from tandem mass-spectrum peaks and uses a block-diffusion language model with an exact mass-shell constraint to generate structures without assuming a molecular formula. On NPLIB1, the authors report the strongest results among methods denied the ground-truth formula across exact match, structural distance, and fingerprint similarity, while recovering the correct formula about as often as a dedicated predictor. Removing the formula oracle places de novo structure elucidation closer to the untargeted setting where genuinely novel metabolites are first encountered.
PROBE evaluates LLM-generated process-model code with executable checks for runnability, units, reference structure, physical invariants, and numerical agreement. The authors report no failures in 150 generations from the three Claude models across ten tasks, with an upper 95% failure-rate bound of 2% under the pooled interpretation. Replacing an LLM judge with executable physics tests makes process-model evaluation reproducible and exposes whether generated code obeys the specified scientific contract.
S1-Omni maps language instructions, crystal and molecular encodings, protein sequences, spectra, and images into a shared representation, aligns that space with scientific laws and expert knowledge, and decodes it for domain-specific tasks. After training across 200 scientific tasks and evaluation on more than 60 benchmarks, the authors report that S1-Omni outperformed general-purpose comparison models on most tests and matched or exceeded specialized models on several. The common representation connects scientific understanding, prediction, and native generation across molecules, proteins, crystals, spectra, and images in one model.
Sussex, S., Borevkovic, E., Lohmann, F., Chen, N., Lüthi, E., Reddy, S. T., Krause, A.
Policy Gradients for Library Design optimizes a synthesis-aware parameterization of stochastic DNA libraries against a chosen sequence objective. The authors report that the method supports multi-round lab-in-the-loop design and was used to synthesize a large influenza-antibody sequence library for about 700 dollars. This formulation lets generative sequence models exploit the scale of multiplexed experiments without requiring every proposed sequence to be synthesized separately.
Wenhao He, Xu Chen, Noah Song, Haowei Xu, Tim S. Hindges, Bohan Li, Zihan Lin, Yu Yao, Avetik R. Harutyunyan, Fang Liu, Yao Wang, Hao Tang, Ju Li
MEHnet-MG predicts an effective one-electron Hamiltonian from an inexpensive density-functional calculation and derives multiple molecular properties from it. The authors report three-point-eight- to 230-fold lower errors than several density-functional baselines while adding about 25 milliseconds per molecule. This matters because an architecture with the right size-scaling can extend coupled-cluster-quality trends beyond the sizes used for training.
UniFlow uses an internal-coordinate normalizing flow to unify independent protein-ensemble sampling, exact likelihoods, and differentiable coarse-grained energies and forces. The authors report close agreement with reference molecular dynamics, transfer beyond the training proteins, and substantially faster sampling than diffusion-based ensemble models. Using one learned density for both equilibrium samples and stable dynamics closes a longstanding divide between generative modeling and molecular simulation.
Engdal, E. S., Funk, J., Bacarreza, O., Machado, L., Johansen, K. H., Kemming, J., Farnsworth, T., Brasas, V., Lefevre-Morand, R. Y. L., Slysz, M., Noerregaard, O. L., Sandberg, O. A. D. A., Makarovskiy, A., Lodahl, P., Acevedo-Rocha, C. G., Kurowski, K., Hadrup, S. R., Clements, W. R., Jenkins, T.
The hybrid pipeline couples a generative adversarial network to latent vectors sampled from a photonic quantum processor to design MHC class I-binding peptides. Across 131 HLA alleles, the authors report that quantum-derived priors increased predicted strong-binder yield, with the largest gains for understudied alleles, and peptide-MHC stability ELISAs confirmed potent stabilizers among designs for three selected alleles. Hardware-derived nonclassical priors provide a structured way to widen sequence exploration when biological training data are sparse while preserving allele-specific anchor constraints.
Lauren N. Walters, Matthew J. McDermott, Bernardus Rendy, Yuxing Fei, Kristin A. Persson, Gerbrand Ceder
The Precursor Genome records 1,035 pairwise solid-state reactions executed by a self-driving laboratory, with thermal histories, masses, instrument settings, raw diffraction, refined structures, and reviewer annotations linked by provenance. The authors report 1,351 diffraction scans and 1,950 automated refinement cases, each evaluated by human experts on a three-tier quality scale. This combination of autonomous experiments and auditable raw-to-assignment records supplies the kind of training target that predictive models of solid-state reactivity have largely lacked.
Ridvan Yesiloglu, Sakib Mostafa, James Zou, Ash Alizadeh, Jiajun Wu, Lei Xing, Ehsan Adeli, Md Tauhidul Islam
scVision uses optimal transport to place genes at fixed positions on a shared pan-tissue map, renders each transcriptome as a continuous image, and applies a masked-image-trained vision transformer as a frozen encoder. In zero-shot tests on six independent held-out studies, the authors report that scVision was the most accurate cell-type annotator, recovered gene programs without supervision, and lost sharply in accuracy when the gene layout was permuted. The fixed spatial map preserves gene relationships and expression magnitude in a representation that can reuse mature computer-vision methods across single-cell studies.
Theo Jaffrelot Inizan, Prathami Divakar Kamath, Alin Marin Elena, Kristin A. Persson
The authors release uMOF, combining a density functional theory dataset, a literature-mined benchmark, and universal metal-organic framework potentials. They report that uMOF models outperform tested baselines for dynamics-sensitive adsorption properties and reduce error by more than 80% to within experimental uncertainty. The package links broad physically diverse training data to experimental benchmarks for transferable metal-organic framework simulation.
Sourin Dey, Dipannoy Das Gupta, Lai Wei, Sadman Sadeed Omee, Jianjun Hu
uFlowCSP learns an average probability-flow velocity to generate crystal structures in a small number of evaluations. On MP-20, the authors report that one step matches CrystalFlow accuracy with far fewer evaluations, while additional steps improve accuracy further. The result shifts the emphasis in crystal generation from peak accuracy alone to useful accuracy per network evaluation.
Johannes Mae\ss, Leon Werner, J. Thorben Frank, Winfried Ripken, Martin Michajlow, Joshua Futterer, Klaus-Robert Muller, Stefan Chmiela
Implicit force fields replace explicit neural-network stacks with self-consistent fixed-point equations whose intermediate representations can be reused between molecular-dynamics steps. The authors report two- to five-fold reductions in compute and memory across invariant, Cartesian-equivariant, and spherical-tensor graph networks. The speedup preserves atomistic resolution and the original integration timestep, opening longer trajectories and larger systems without spatial or temporal coarse graining.
Joshua W. Sin, David Ming Segura, Bojana Rankovic, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt Puntener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
The method learns a reaction representation from textual condition descriptions with a fine-tuned language model and Gaussian-process surrogates in multi-objective Bayesian optimization. In prospective studies, the authors report that high-throughput experiments produced conditions translating directly to gram scale with high isolated yields and enantiomeric excess. The approach treats the description of a reaction system as a learned experimental representation rather than a fixed descriptor library.
AIDEN separates element-dependent one-center charge density from environment-induced redistribution. The authors report state-of-the-art accuracy on periodic crystal benchmarks, competitive molecular performance, and zero-shot transfer across structurally distinct out-of-distribution cases. It advances charge-density surrogates by separating transferable local structure from the spatial grid used to evaluate the field.
Xu, F., Zhuang, Z., Zhu, Y., Ying, B., Hou, N., Lin, W., Wang, L., Yang, C., song, j.
Spaceland learns continuous gene-expression fields in morphology-informed histological space, using optical-flow interpolation of foundation-model histology features to build a dense three-dimensional scaffold and decode sparse transcriptomic measurements. In mouse benchmarks, the authors report that Spaceland outperformed spatial-transcriptomic and two-dimensional histology baselines while resolving sub-spot organization; in planarian regeneration, four sections per stage supported time-resolved whole-organism analysis. The continuous-field formulation provides a scalable route from sparse sections to high-resolution whole-organ molecular reconstruction across tissues and regeneration stages.
This active-learning platform uses Bayesian monotonic I-spline regression so each posterior saturation curve rises from zero and never decreases. The authors report that it reaches noise-floor accuracy within a 20-measurement budget in every regime, in as few as seven measurements. The same shape-constrained surrogate can make sparse-data experiment selection physically consistent across self-limiting chemical and materials responses.
TRACE combines atomic cluster-expansion density correlations with local multihead cross-attention in an energy-conserving equivariant architecture for interatomic potentials. With the same architecture, the authors report a methyl-migration activation free energy of 27.92 plus or minus 0.03 kcal/mol, close to the experimental 29.2 plus or minus 1.1 kcal/mol, together with experimental agreement for perovskite phase behavior and liquid-water structure. An equivariant potential that spans crystallization, liquid structure, and chemical reactivity reduces the need for separate models tied to individual phases or processes.
Sk Md Ahnaf Akif Alvi, Jan Janssen, Danny Perez, Douglas Allaire, Raymundo Arroyave
The workflow inserts a Gaussian-process acquisition gate between crystal generation and a property oracle in an RL-steered materials-design loop. The authors report that the gate comes within about 9% of exhaustive oracle spending at roughly one-fifth of the calls, while a density-functional-theory check confirms bulk-modulus predictions within 2.5% on average. This lets generative materials searches spend expensive calculations where a surrogate expects them to be most useful.