AI is helping chemists choose experiments, design molecules and interpret measurements. The results range from catalysts and enzymes tested in a laboratory to software identifying compounds from spectra. A useful chemical result and a demonstrated advantage of the algorithm are related questions, but they require different comparisons.

These ten studies appeared as journal papers or a preprint in 2026, through 30 September. Nine report physical experimental validation; the tenth evaluates molecular identification from spectra. Publication this year does not mean that all the underlying experiments or first public disclosures happened in 2026.

1. Flex-Cat: optimising which product a reaction makes

North Carolina State University and collaborators’ Flex-Cat combines parallel batch reactors with Bayesian optimisation. It tunes rhodium-catalysed propylene hydroformylation to favour straight-chain or branched aldehydes.

Its 680 experiments comprised 320 initial measurements and 360 model-guided experiments across 16 ligands. Maximum turnover frequency rose from 2,900 to 7,500 h⁻¹, about 2.6-fold: moles of aldehyde produced per mole of rhodium per hour. This compares the best guided result with the initial maximum, not an industrial benchmark.

Separate, equal-budget simulations compared the selection policy with random sampling and continued Latin hypercube sampling under the same operating constraints. The adaptive policy found promising regions earlier. That supports the search strategy computationally; it is not a parallel physical trial of all three methods. In 20 mL reactor retests, product balance largely survived, but activity fell in every listed comparison.

Aldehydes supply plastics, detergents and solvents. The practical prospect is helping chemists switch between desired products. Transferring those recipes to production still requires work on mixing and gas transfer.

2. BASF and FHI: searching catalyst formulations with a reactor in the loop

A BASF–FHI–TU Berlin study used adaptive experiment design to find catalysts that convert propane into propylene. An interpretable statistical model links performance to platinum, added elements and their interactions; this is model-guided chemistry, rather than a neural-network system.

After 48 experimental evaluations, the team modelled a composition space of 6.08 × 10¹³ formulations. A Pt–Sn–K–Ce–Zn catalyst gave a second-cycle integrated propylene yield of 32.8%, versus 29.5% for Pt–Sn–K: a 3.3-percentage-point increase. Tests ran at 600°C in industrial-like gas conditions, with two 20-hour cycles separated by four-hour regeneration.

Its late-cycle yield-retention ratio was 98.3%, versus 95.3%: yield during hours 15–20 divided by that during hours 10–15. This measures slower deactivation within the cycle, not a catalyst’s industrial lifetime. Models omitting promoter interactions failed to reproduce the full predicted performance space; that comparison tests the model, not equal-budget discovery campaigns.

Propylene supplies plastics and synthetic fibres. The opportunity is to balance output with durability while identifying useful ingredient combinations. Repeated regeneration and production-scale testing remain necessary; synthesis was only partly automated.

3. ChatHEA: language models connected to catalyst screening

ChatHEA links a specialist language-model assistant with catalyst screening. GPT-4o extracted literature information, experts curated it, and a fine-tuned Llama-3-8B supported alloy enumeration, planning and analysis.

The team synthesised 100 five-element alloys. In a 5 cm² hydrogen–air fuel cell, FeCoCuPtIr delivered peak power density of 0.789 W/cm², versus 0.724 W/cm² for commercial platinum-on-carbon, at 80°C and 150 kPa absolute pressure. The experimental methods specify cathode noble-metal loadings of 0.05 and 0.10 mg/cm² respectively. This is a device comparison under those different loadings, not an isolated alloy comparison at equal loading.

The result demonstrates the combined workflow. It does not measure the assistant’s contribution against a workflow without it. The authors explicitly say intelligent candidate prioritisation and active learning to reduce experimental workload have not yet been implemented.

The intended applications include vehicle and backup-power fuel cells. Useful device performance with lower precious-metal loading is a development lead; the alloy still contains platinum and iridium. Larger stacks, durability and manufacturing economics need separate evaluation.

4. Photocatalytic materials: screening frameworks, then making eight

A Nature Synthesis study used data from 55 papers to select porous organic materials that produce hydrogen peroxide with light. Its predictive model combines molecular features, reaction conditions and physics-informed descriptors. The March 2026 paper followed code and data released in May 2025.

The researchers screened 10,881 structures computationally and synthesised eight. COF-343 produced 12,978.7 μmol of peroxide per hour per gram of catalyst. The assay used 3 mg of catalyst in 20 mL water under ambient air, with a 300 W xenon lamp filtered to wavelengths above 420 nm, at 25 ± 1°C. This mass-normalised laboratory rate does not establish sunlight efficiency or production throughput.

A conventional deep neural network outperformed the tested graph, Transformer and ChemBERTa alternatives on the initial small dataset. That is a prediction comparison, not a controlled test of how many synthesis experiments each strategy needs. Figure 5’s predicted-versus-measured comparison uses single experimental values, limiting estimates of repeatability.

The prospect is another route to an industrial oxidant using water, oxygen and light. Feedback from new experiments could refine material selection; product concentration, energy input and larger-scale operation still determine whether the route is practical.

5. Lila: a promising catalyst result that is still a preprint

Lila Sciences’ September preprint describes an AI-guided, human-supervised search for acidic oxygen-evolution catalysts: the oxygen-making half of water electrolysis. It screened 2,942 catalysts across 53 material systems and 26 elements.

An indium–manganese–palladium oxide stayed below 0.5 V overpotential for over 1,000 hours, while a palladium-oxide reference crossed that threshold around 200 hours. Testing used thin films in 1 M sulphuric acid at 10 mA/cm², in repeated 12-hour measurements with interruptions; overpotential was corrected for electrical resistance. Palladium accounts for about 99.9% of the lead’s metal atoms, despite its lack of iridium and ruthenium.

A separate retrospective replay found that adaptive selection advanced the activity–stability frontier faster than fixed-policy Bayesian optimisation or LLM-only selection. Those methods selected from completed experimental data, rather than running independent live campaigns. This supports the policy under replay conditions, not measured savings in a working laboratory.

The intended benefit is another catalyst option for hydrogen production. Lila is testing forms closer to industrial use. The chemical lead and algorithm comparison remain preprint results; neither establishes a commercial electrolyser.

6. BioStrucTag: choosing enzymes for a particular stereoisomer

The BioStrucTag paper addresses a synthesis problem: using enzymes to make different spatial arrangements of the same alcohol product. It combines protein-sequence features and three-dimensional representations of a substrate in an enzyme’s binding pocket with rational protein design.

Engineered alcohol dehydrogenase variants collectively accessed all eight alcohol-product stereoisomers in the studied reaction family. The authors report enantiomeric and diastereomeric ratios reaching 99:1 for most tested substrates. These measure the balance of product forms, not product yield, and do not mean that every substrate produced every form at that ratio.

This establishes access to eight product forms through joint use of predictions and rational design; it does not quantify the model’s separate contribution. The August 2026 publication also draws on a dataset produced in 2021–2025.

For synthesis chemists, the prospect is an enzyme toolkit for choosing a required molecular shape and comparing the properties of its alternatives. Transferring it to another reaction or production process requires new experiments.

7. REAP: improving enzyme activity through successive experiments

REAP combines a protein model, ranking-based learning and robotic experiments. Measurements from each round guide the next enzyme variants to test.

For P450 BM3, five rounds produced a 44.93-fold activity gain over the engineered FL#62 reference in an assay converting deoxypodophyllotoxin towards an anticancer-drug building block. Activity was defined by target product formed in a fixed time under standardised conditions. A separate 57-fold result came from an additional rationally designed variant, not the same campaign endpoint.

The paper also compared its ranking objective with other training losses and EVOLVEpro on matched data splits. With 50 training variants, it was comparable to EVOLVEpro; with 400, it led across eight benchmark datasets. Those comparisons isolate prediction-method performance. The 44.93-fold chemical gain is measured against the starting enzyme, not against a parallel evolution campaign using another algorithm.

The team tested Sortase A as well and released research code. The prospect is better biocatalysts for synthesis: higher activity could help a route become practical if stability and selectivity survive process conditions. Each new reaction still needs a reliable assay and experimental validation.

8. PolyCAML: antifungal polymers from an active-learning loop

PolyCAML combines a graph transformer, automated polymer synthesis and biological testing. Its library contained 516,114 possible copolymers; measured feedback guided subsequent candidate selection.

Within 18 days, the workflow identified 11 leading candidates. Four selected for further evaluation had minimum fungicidal concentrations of 8 μg/mL or less and concentrations reducing L929 cell viability by 50% of at least 512 μg/mL. These endpoints measure fungal killing and effects on a laboratory cell line, respectively; they are not patient efficacy or a human safety margin.

The reported 18 days is the duration of this campaign, not time saved against a matched conventional campaign. The independent research highlight describes four design–synthesis–test–learn rounds; it discusses the same study rather than replicating it.

The polymers also form nanoparticles that carry fluconazole into fungal biofilms. The authors report treatment effects in animal models of bloodstream and eye infection. For infection researchers, the development lead is a material combining antifungal action with drug delivery. Clinical benefit, formulation and safety remain to be established.

9. Chemical language models: generating ligands, then measuring potency

A Nature Communications study trained recurrent chemical language models incrementally on existing molecules ordered by potency. It then generated designs for synthesis and testing; this was not repeated retraining after each new laboratory round.

For PPARγ, the researchers selected ten top-ranked designs, synthesised nine, and tested all nine. Five benzoic-acid derivatives had EC50 values of 0.6–3.1 nM, versus 37 nM for the retested reference: roughly 12–62 times lower. EC50 is the concentration producing half the maximum response. Those five are a subgroup of selected designs, not the hit rate across all generated molecules.

Computational comparisons favoured incremental training over training once on all templates or only the most potent examples. Before synthesis, the team excluded designs also generated by the all-template baseline. The laboratory tests validate the selected molecules, but do not compare the hit rates of both strategies through parallel synthesis campaigns.

The study also examined RORγ and released implementation code. The prospect is additional candidates for refining a known molecular scaffold in drug research. Greater potency alone does not establish selectivity, safety or clinical benefit.

10. SECS: interpreting spectra exposes the transfer problem

SECS combines contrastive learning and evolutionary search to propose structures from NMR, infrared and mass spectra. The authors released research code; a web implementation accepts proton NMR and a molecular formula. The June 2026 paper followed spectral-model resources released in 2024.

On 323 Chemotion compounds within specified elemental limits and containing at most 15 heavy atoms, the correct structure ranked first in 27.9% of cases and within the top 20 in 58.5%. Across the full 1,486-compound set, those figures were 10.3% and 24.8%. These are different evaluation populations: restricting the chemical space makes the task easier. The higher percentages should not be applied to arbitrary samples.

The paper also isolates training effects: augmentation and fine-tuning improved performance on a separate experimental dataset. Those results describe that dataset, not a replacement accuracy estimate for Chemotion.

The intended use is a second opinion when checking what a chemist made. A shortlist can reveal alternative structures or mistaken assignments, but still needs expert assessment. Wider spectral support is a stated development direction.

How AI is changing chemistry

These studies show several overlapping roles for AI in chemistry. Bayesian optimisation and active learning help select the next experiments. Predictive neural networks estimate properties using inputs such as molecular structure, protein sequence or reaction conditions. Generative models propose molecular candidates, while LLM assistants extract and organise literature to support research decisions. In analytical chemistry, contrastive learning and search algorithms help connect measured spectra with possible molecular structures. One system can combine several of these approaches.

The common thread is a closer connection between computation and chemical evidence. Some systems narrow a search before synthesis; others learn again after each laboratory round. That can help chemists tune a catalyst’s product mix, select an enzyme that makes the required molecular form, or find materials worth developing further. The strongest examples here include making and testing what the model proposes. A better catalyst, enzyme or ligand validates a chemical lead. Establishing that AI found it more efficiently also needs a comparable search strategy tested on the same task and budget. Here, several algorithm comparisons are simulations or retrospective tests; they should not be mistaken for measured reductions in laboratory cost or time.

Turning these results into a wider capability requires workflows and models that transfer between laboratories. A 2026 review of self-driving laboratories identifies that transfer, greater throughput and complete experimental records as priorities. An Argonne battery-electrolyte campaign illustrates a physical limit: sourcing chemicals and manually preparing reagents constrained throughput even after experiments were automated. AI can expand the search for useful chemistry, but scaling that contribution also requires better ways to make, measure and reproduce the proposed results.

← Back to blog