Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript.
Advertisement
Nature Computational Science (2026)
A preprint version of the article is available at ChemRxiv.
The Buchwald–Hartwig cross-coupling is a cornerstone of modern pharmaceutical synthesis, yet predictive modeling of its outcomes remains constrained by data quality and chemical space coverage. Electronic laboratory notebooks contain heterogeneous, noisy records, while open-source high-throughput experimentation (HTE) datasets are fragmented and narrow in scope, leading to poor model performance on unseen substrates and conditions. Here we introduce a framework that systematically standardizes and integrates multiple reaction datasets into a high-quality, unique-structure-per-entity dataset, coupled with active learning to strategically expand chemical space. By merging published Buchwald–Hartwig HTE data with new experimental results, we achieve a model with predictive power across novel substrates and conditions, delivering improved out-of-distribution predictions compared with previous approaches. Crucially, model-guided reagent recommendations were validated experimentally, confirming the framework’s utility to uncover unexplored reactivity. This work establishes a blueprint for robust machine learning in synthetic chemistry and enables preemptive in silico reagent screening to accelerate pharmaceutical discovery.
The Buchwald–Hartwig (BH) cross-coupling reaction1,2 has become an indispensable tool in modern synthetic chemistry3, particularly in the pharmaceutical industry, where it is used to synthesize structurally diverse (hetero)aryl amine compounds4. The ability to reliably forecast the outcomes of BH reactions is thus invaluable, with machine learning (ML) methods offering a pathway to streamline and accelerate lead discovery efforts5,6. A recent example introduced a tool7 to estimate substrate-adaptive conditions for palladium-catalyzed C–N couplings, trained on a diverse experimental dataset of reactant pairings and systematically refined through active learning. Beyond predicting reaction outcomes8,9,10,11, the integration of reagent dictionaries12 enables yield prediction models to suggest the most suitable catalysts, bases, solvents and other conditions to achieve desired transformations. Retrosynthesis planning13,14,15, which identifies viable synthetic routes to target molecules, can further benefit from accurate reaction-level predictors, as robust models of cross-coupling chemistry directly expand the reliability of multistep route design. However, the inherent complexity and sensitivity to conditions in these reactions pose substantial challenges for traditional predictive models16,17,18. Addressing these complexities requires not only vast numbers of data but also a degree of consistency and precision in data generation that is often difficult to achieve, in part due to the heterogeneity of the data sources.
Mining data from electronic laboratory notebooks (ELNs) is one possible avenue towards building predictive synthesis models, but the performance of these models is capped due to data quality. ELN and literature scraped datasets19 can vary widely in reaction set-up, purification processes and yield reporting. Additionally, such datasets have heterogeneous, and sometimes contradictory, representations of the same chemical entities, do not capture all reaction components in a tabular format and, in the case of literature, lack negative examples.
While better algorithms are still required to tackle this industry-wide challenge17, high-throughput experimentation (HTE) has emerged as a powerful method for generating reaction data with high consistency and quality through controlled and automated processes. Recently, a large-scale release of over 39,000 previously proprietary HTE reactions has further enriched the chemical data landscape20, enabling systematic exploration, which unveiled hidden statistical relationships between reaction components and uncovered dataset biases.
Similarly, parallel library synthesis has been instrumental in drug discovery, enabling the efficient production of structurally diverse analogs through robust synthetic methodologies21. Dombrowski et al.22 analyzed over 5,000 parallel libraries across 14 years, showcasing the evolution and proven success of these approaches. Notably, BH cross-coupling reactions were identified as one of the most challenging, with success rates among the lowest. Nonetheless, HTE platforms have demonstrated notable success in optimizing BH cross-coupling reactions, as reported in ref. 23, offering robust tools for synthesizing complex molecules on micro- and nanomolar scales and under automation-friendly conditions.
Collectively, these advancements established a foundation for reliable and standardized data generation, essential for navigating the inherent complexities of predictive BH reaction modeling. Moreover, HTE allows for the quick accumulation of a vast number of data points tailored to specific needs, encompassing a wide range of substrates and diverse reaction conditions. The ability to efficiently explore diverse reaction environments and generate comprehensive data makes HTE an ideal source of data for ML models aimed at estimating reaction outcomes, particularly when robust out-of-distribution (OOD) performance is required.
ML models excel at identifying complex patterns within large datasets24, and when trained on high-quality data, such as HTE data, they can more accurately predict reaction outcomes, even for previously unseen substrates. This is the reason why in recent years there has been an increase in studies combining ML and HTE7,25,26,27. The integration of HTE and ML therefore has the potential to transform reaction prediction by delivering models capable of OOD generalization—where predictions remain reliable despite differences between training and test data. This synergy is particularly impactful for applications in synthetic and medicinal chemistry, where the ability to forecast reaction outcomes in new, untested scenarios can substantially improve the time and resource efficiency required for chemical exploration and development.
The goal of this study is to develop a predictive model with wide OOD capabilities for BH reactions by integrating a newly generated HTE dataset with available open-source HTE data (Fig. 1). To this end, we employed a data-driven active learning28 approach that iteratively refined the model’s accuracy and generalization capacity across diverse chemical spaces. By standardizing and unifying disparate datasets into a single, high-quality format, we have created a comprehensive dataset that supports a model capable of reliable predictions across a meaningful area of the BH reaction landscape. This work not only advances the state of predictive chemistry for a critical reaction type but also demonstrates the potential of HTE and ML integration as a general framework for achieving robust, adaptable reaction prediction models and reagent-recommender systems.
A model with robust predictive performance is obtained through the combination of internally generated datasets and open-source HTE data. All open-source datasets used in this study are listed with an abbreviated journal name and publication year, and their corresponding reference number. By applying this model, previously unsuccessful reactant pairs were rescued, and challenging, OOD substrates were successfully coupled. JnJ, Johnson & Johnson.
In this project, we strategically leverage automated workflows using 96-well and 384-well plates to execute our HTE designs, with product quantification performed using ultraperformance liquid chromatography (UPLC)29-MS-CAD (mass spectrometry–charged aerosol detector)30 (Fig. 2). We employed an iterative approach for data generation, purposefully selecting experimental conditions across three iterations. The first iteration focused on establishing key variables and increasing diversity in reaction conditions by exploring combinations of catalysts, bases and solvents. For iterations 2 and 3, we employed model-guided experimental design, targeting broader coverage of the substrate chemical space. Reactions with high model uncertainty were prioritized, enabling continuous refinement of the predictive model’s accuracy on the broader space. This systematic approach substantially enhanced data diversity and quality, resulting in the largest and most comprehensive BH HTE dataset in literature to date, with 11,300 (11.3k) high-quality reactions. The production of these data followed a quantitative methodology consisting of two large phases. The first phase aimed to produce a starting baseline ML model, which was used in a second phase to perform experimental design in an active learning formulation. Additionally, we curated and standardized all other published BH HTE datasets into a unified format, creating an extensive dataset with ~27.5k reactions. To ensure comparability across heterogeneous quantification methodologies (for example CAD, isolated yield, MisER and so on), we harmonized all reactions under a binary classification framework using a 10% yield threshold. This cutoff reflects a standard go/no-go criterion in medicinal chemistry and permits consistent integration of multisource data.
Left: chemical-space coverage and overlap of data-production compounds with ELN entries. tSNE, t-distributed stochastic neighbor embedding. Middle: automated platforms utilized in the execution of HTE campaigns. Right: continued improvements in model quality were observed under the active learning approach, as indicated by the balanced accuracy. info., informer set.
Source data
The strategy for the first phase was centered around four factors: (1) informer set, (2) structure diversity criteria, (3) space density criteria and (4) undesired substructure filtering (Supplementary Section 2).
The second phase embraced an active learning methodology utilizing (1) BERT Enriched Embedding12, an extension of a language model for yield prediction16, which was trained on the data from the previous phase, (2) exploration of a virtual BH space of 64,000,000 (64M) reactions, (3) the Ranked Batch-Mode Active Learning31 function using model uncertainty and Tanimoto distance from previous selections and (4) diversity and density (Supplementary Section 3).
We address model bias in benchmarking by training six ML models with two types of features, the differential reaction fingerprint32 (DRFP) and a vector of physicochemical attributes, and by training one artificial intelligence Simplified Molecular Input Line Entry System (SMILES)-based model. We address data bias33 by testing on OOD data; OOD was created by clustering all available reactions with DRFP and k-means34 and only testing in the clusters where source is not represented. Finally, we addressed metric bias35 by considering a total of eight performance metrics and identifying four that can summarize the capabilities of interest (uncertainty estimation, good precision and serviceable recall in class imbalance OOD scenarios) of the different models.
The data produced in this study, labeled JnJ25 in Fig. 3, surpass the previous studies in our comparison not only in overall size but also in diversity of reagent combinations, with the second highest diversity of unique substrate pairs, behind a recently published work36 (Fig. 3a). The combination of JnJ25 data with all remaining BH HTE open-source datasets is denominated ‘All w JnJ25’.
a, Number of unique substrate pairs against number of unique reagent combinations. Each of the eight open-source BH HTE datasets used in the project is represented in size; these are colored according to CRDS (a data-centric score with strong correlation to OOD performance). ‘All w/o JnJ25’ denotes the combined open-source Buchwald–Hartwig HTE datasets excluding the JnJ25 dataset. b–d, Performance across four critical metrics for the best ML model for each given dataset when tested on OOD clusters built from all remaining open-source data. Each square marker represents the mean performance across n = 10 independent random cluster initializations, where each initialization uses a different random seed for k-means clustering (k = 12 clusters). Error bars indicate s.d. across the ten cluster initializations. For each data source, substrates present in the training set were excluded from test clusters, creating a strict OOD evaluation. Hexagon markers represent the combined All w JnJ25 dataset, evaluated across n = 10 iterations where in each iteration four cycles of train–test splits (three clusters held out, nine used for training) were performed until all 12 clusters appeared in the test set once, with performance averaged first across the four cycles within each iteration, then across all ten iterations. b–d, AUPRC reflects how effectively the model enriches for successful reactions when ranking candidates (all thresholds) (b); balanced accuracy reflects how effectively the model makes correct predictions for both successful and unsuccessful reactions, independent of class imbalance (c); F1 score reflects how effectively the model converts predictions into correct operational go/no-go decisions at a chosen threshold (d). AUPRC, area under the precision–recall curve.
Source data
Unexpectedly, we discovered that model performance on OOD predictions was determined not primarily by dataset size, but rather by a previously unrecognized relationship between data diversity and prediction accuracy, as shown in Table 1. This finding challenges the conventional wisdom that larger datasets necessarily lead to better generalization, instead showing that strategic diversity in reaction space coverage is the key driver of robust predictions. We introduce a new data-centric metric, Compound-Reaction Diversity Score (CRDS), which aims to estimate the ability of potential models trained on a given dataset to perform in OOD scenarios. This metric can be employed retrospectively to elucidate the key dataset attributes that underpin the effective training of ML models for robust OOD predictions, as well as in prospective planning of a data production campaign.
Its formula combines the number of unique reactant pairs, unique reagent combinations and the size of the dataset. Additionally, each component has adjustable thresholds to induce diminishing returns and varying weights. When applying this calculation, we found a 0.79 Pearson correlation between CRDS and receiver operating characteristic-area under the curve (ROC AUC) in the most challenging OOD test, indicating a high degree of correlation between substrate diversity quantity and OOD performance (Table 1). More details of the CRDS formula and calculation can be found in Supplementary Section 4.
Figure 3b–d presents the results of two benchmarking experiments. In these panels, the square symbols represent the highest-performing combinations of ML model, feature vector and data source, evaluated on OOD clusters in which the corresponding data source is not represented. Furthermore, any test instances containing a substrate present in the training set were excluded. This creates a very challenging OOD test, as the DRFP contained reagents, leading to cluster formation impacted by both reagent and substrate scope. The objective of this analysis was not to maximize predictive performance, but rather to compare the relative utility of each data source in enabling the training of models that are effective for reagent screening and experimental design in previously unexplored BH chemical spaces.
Hexagon symbols quantify the performance gain obtained by incorporating the JnJ25 dataset into all previously available open-source BH HTE datasets. This evaluation was conducted using the DRFP OOD cluster test, in which reactions—represented as DRFPs incorporating both reagents and substrates—were partitioned into 12 clusters via k-means clustering. The test procedure comprised ten iterations: in each iteration, there is a cycle where three clusters were selected without replacement as the test set, while the remaining nine were used for training; the cycle progresses until all clusters have appeared in the test set once. Performance was first averaged across the cycle with the held-out clusters within each iteration, and then across all iterations to determine the metrics value represented by the hexagon in the scatter plot. We also demonstrate how a binary reactivity classification model can be used as a reagent recommender and HTE plate designer (recommending both substrates and reagents) when combining it with a reagent dictionary and uncertainty estimation. The reagent-recommender workflow using a reagent dictionary, feature construction and scoring pipeline is provided as a self-contained notebook in the repository (Supplementary Section 5).
Before utilizing these types of models in the laboratory or executing experimental validation, it is important to assess whether the models are uncertainty calibrated, as uncertainty will be used to rank reagent and substrate pair combinations. We tackle this by binning the predictions of OOD reactions into five confidence bins. The confidence value is inversely derived from uncertainty and normalized between 0.5 and 1 (Supplementary Section 4.3, Supplementary equation (1)) as described in ref. 12. If the model is perfectly calibrated then performance will perfectly correlate with confidence, following the gray dashed line in Fig. 4.
a,b, Calibration curves showing the relationship between model confidence (x axis) and ROC AUC (a) (right y axis) or balanced accuracy (b) (left y axis) across five confidence bins for OOD test reactions (n = 20 independent random cluster initializations). The right y axis shows the percentage of test data predicted within the respective confidence interval. Each point represents the mean performance within a confidence bin, with error bars indicating s.d. across the 20 cluster initializations. The gray dashed line represents perfect calibration where confidence is equal to success rate. Confidence values are derived from model prediction probabilities, inversely related to uncertainty and normalized between 0.5 and 1. LR, logistic regression; conf., confidence.
Source data
Through the use of the confidence metrics, it is possible to rank reagent combinations for any given substrate pair by their likelihood of success. Furthermore, we found that a relatively simple random forest (RF)37 model, when trained on our properly curated dataset, achieved >90% ROC AUC on high-confidence OOD predictions, demonstrating that model architecture sophistication may be less critical than previously thought, at least for achieving reliable chemical prediction in a single reaction type with chemical space defined in ~27,500 samples.
To further investigate this observation, we performed hyperparameter optimization for the three top-performing model families (RF, histogram-based gradient boosting (XGB) and multilayer perceptron (MLP)) across two feature representations: full-reaction DRFP and a split representation separating substrate–product DRFP from reagent physicochemical descriptors (Supplementary Section 5). The results revealed that the choice of feature representation had a larger impact on OOD performance than the choice among top ML models. This suggests that, when the feature representation cleanly encodes the relevant chemical information, separating substrate structural identity from reagent properties, a simple model that does not distort it is sufficient. RF is particularly well suited to this task because tree-based splits on binary fingerprint bits directly correspond to ‘substructure present/absent’ decisions, and the ensemble’s natural subsampling prevents overfitting to training-distribution artifacts. We expect pretrained transformers to become more advantageous as additional reaction types and larger chemical spaces are incorporated. A detailed discussion of the model selection rationale is provided in Supplementary Section 5. It is also worth noting that richer descriptor families such as quantum-mechanics-informed features enriching deep learning embeddings represent promising directions for future work.
The first experiment was aimed at rescuing historically unsuccessful substrates (Fig. 5a). In this scenario, we tested the practical applicability of our model by using it as a reagent recommender12 to identify successful reaction conditions for 11 historically unsuccessful aryl halide–amine substrate combinations gathered from HTE literature reactions as well as data produced during this project. Previously, three substrate pairs had each failed across 24 reaction conditions, two pairs had each failed across two reaction conditions and the remaining six pairs had each failed in one reaction condition. We then selected catalysts, bases and solvents that were frequently covered in the dataset (minimum 400 samples), leading to a total of 1,092 reagent condition combinations. Leveraging the model, we systematically ranked all these combinations according to predicted confidence. For each substrate pair, only the top four recommendations were experimentally tested. Out of these 44 reactions, 27 succeeded (yield > 10%), efficiently providing viable synthetic routes for 10 of the 11 substrates (91%). This illustrates the model’s ability to efficiently identify effective reaction conditions, substantially reducing the experimental burden and conserving laboratory resources.
a, 11 previously unsuccessful aryl–amine pairs were tested across the four highest-ranked predicted reaction conditions, achieving viable synthetic routes for 10 of the 11 substrate pairs. b, Prospective validation of the model across a novel chemical reaction space; experimental results present fold enrichment of reactions selected from a vast chemical domain (~2.46 × 1011 reactions) predicted by the model with moderate confidence (<0.85) and high confidence (>0.85) versus a baseline set. The dashed horizontal line indicates the no-enrichment baseline. DMA, dimethylacetamide; Tx, transformation x.
Source data
The second experiment (Fig. 5b) was a prospective validation across a differentiated chemical space from the training data. We assessed the model’s predictive generalizability in a vast, unexplored chemical domain sampling from a virtual space consisting of 13 catalysts, 12 bases, 7 solvents, 12,252 aryl halides and 18,364 amines, representing approximately 2.46 × 1011 possible BH reactions. We specifically selected entirely novel aryl and amine combinations (absent from any previous training data) to rigorously test the model’s ability to identify successful reactions. From this space, we experimentally evaluated three groups of reactions.
A baseline group of 4,800 reactions, selected via diversity, space density, pharmacological relevance and model uncertainty, which yielded around 19% success rate.
Reactions predicted with moderate confidence (<0.85; for reference 0.5 indicates complete uncertainty) achieved a success rate of 21%, a fold enrichment of 1.11 over the baseline (Fig. 5b).
Reactions predicted with high confidence (>0.85) achieved a success rate of 33%, representing a fold enrichment of 1.64 over the baseline (Fig. 5b).
Critically, these results highlight the model’s exceptional capacity to distinguish reactions where modeling can enrich from those out of domain, even when dealing with entirely novel chemical combinations drawn from an extraordinarily expansive chemical space.
Collectively, these experiments validate that our unified, actively curated ML model substantially enhances BH reaction predictions by (1) effectively rescuing historically unsuccessful reaction pairs through minimal additional experimentation and (2) confidently identifying viable reactions in previously unexplored chemical space. This approach offers a powerful, transferable framework for accelerating the discovery and optimization of pharmaceutically relevant chemical transformations, and the growing high-quality curated dataset enables further modeling research.
We have developed a unified, actively curated ML framework capable of robust OOD prediction for BH cross-coupling reactions. By systematically standardizing heterogeneous datasets, strategically expanding reaction space through active learning and validating predictions experimentally, we demonstrate that data diversity, rather than dataset size or model complexity alone, is the key driver of generalizable performance. This insight not only advances predictive chemistry for a pivotal reaction class but also establishes principles broadly applicable to other transformations.
The models developed in this project operate within the reagent space covered during training (32 palladium precatalysts, 12 bases, 12 solvents), which, while comprehensive for most experimental objectives, represents a subset of commercially available options. Predictions for entirely novel reagent classes not represented in training should be interpreted accordingly. The binary classification framework (10% yield threshold) reflects practical go/no-go decisions in medicinal chemistry but does not capture quantitative yield information. Notably, heterogeneous yield quantification methods across aggregated datasets (CAD, isolated yield, MisER, liquid chromatography area percent or LCAP) were harmonized through this binary threshold, with only 11.9% of reactions falling in the 5–15% boundary region where measurement variability could influence classification. The CRDS shows strong correlation with OOD performance for BH chemistry (Pearson r = 0.79 versus r = 0.60 for dataset size alone), but the weighting exponents were empirically fitted to this reaction class and would require recalibration for other transformations where combinatorial complexity differs.
An opportunity for future work remains the harmonization of noisy, heterogeneous reaction data (for example, industry ELNs), inclusion of conditions in a global setting and scaling predictive frameworks across reaction types. Better algorithms are still required to tackle these industry-wide challenges; potentially in the future the use of knowledgeable chemistry artificial intelligence agents38,39 will make possible a harmonization of the free-text protocol and tabular data, while noise detection techniques will be able to remove the most ‘obvious’ noisy samples from the training dataset40,41,42,43. Looking ahead, the integration of additional HTE campaigns and richer reaction context representations can further enhance predictive power.
Standard glovebox techniques were utilized for the handling of air- and moisture-sensitive reagents. All reagents and starting materials were procured from commercial suppliers and employed as received without further purification. Solvents were sourced from sealed anhydrous containers to ensure minimal exposure to atmospheric moisture. Catalysts were obtained from commercial suppliers and stored within an MBraun glovebox under an inert atmosphere.
1-ml clear glass shell vials (8 mm × 30 mm, Analytical Sales Services) and 96-position reactor blocks were used for 96-well-plate reactions. For 384-well-plate reactions, Axygen 384-well clear V-bottom 120-µL polypropylene deep-well plates were employed.
Appropriate safety precautions were taken when handling air- and moisture-sensitive reagents, as well as strong acids and bases.
A CHRONECT Quantos was used to dispense solid materials. Tecan Fluent, Unchained Lab Junior, SPT Labtech dragonfly and Apricot PP5 were used as liquid handlers for different purposes. An IKA MATRIX Orbital Delta+ shaker was used for shaking and heating the 384-well plate in a Nano Nest (Analytical Sales & Services). UPLC-MS analysis was performed using a Waters ACQUITY system equipped with as SQD2 detector and a CAD.
The HTE was designed using Katalyst software, which facilitated the generation of automation execution files for both solid-dispensing and liquid-handler robots. All palladium Buchwald precatalysts (1 mmol, 10 mol%) and solid bases (12 mmol, 1.2 equiv.) were accurately dispensed into 1-ml vials arranged in a 96-well plate, using a CHRONECT Quantos system within a glovebox environment. Subsequently, aryl halide (10 mmol, 1.0 equiv., 50 ml of a 0.2-M stock solution) and amine (12 mmol, 1.2 equiv., 50 ml of a 0.24-M stock solution) were introduced via an Unchained Lab Junior liquid-handling system. Finally, the liquid bases (12 mmol, 1.2 equiv.) were directly added to the corresponding vials. The resulting reaction mixtures were sealed and stirred at 100 °C on an Unchained Lab Junior deck within the glovebox for a duration of 16 h.
The HTE was designed utilizing Katalyst software. Aryl halide (2 mmol, 1.0 equiv., 10 ml of a 0.2-M stock solution in 1,2-dichloroethane) and amine (2.4 mmol, 1.2 equiv., 10 ml of a 0.24-M stock solution in 1,2-dichloroethane) were dispensed into the corresponding wells of a 384-well plate using a Tecan Fluent liquid-handling system. Following this, the plates were subjected to a drying process to remove the solvent using a Genevac. Subsequently, the plates were transferred into a glovebox, where 10 µl of palladium Buchwald precatalysts (0.2 mmol, 10 mol%) and liquid bases (2.4 mmol, 1.2 equiv.), each dissolved in their respective reaction solvents, were added using an SPT dragonfly liquid dispenser. This ensured that the total volume of solvent for each reaction was 20 ml. After the addition of the reagents, the plates were sealed with Agilent PlateLoc and placed in a Nano Nest. The Nano Nest was maintained at a temperature of 100 °C and agitated at 800 r.p.m. on an orbital shaker for a duration of 16 h.
Following the reactions, analytical plates were prepared using an SPT Apricot PP5 system. For the 96-well-plate HTE, the reaction mixtures were diluted with 400 ml of dimethylsulfoxide (DMSO) and mixed. Subsequently, 50 ml of the resulting solution was transferred to a 96-well-plate analytical block containing 450 ml of CH3CN for further analysis. For the 384-well plates, the reaction mixtures were diluted with 40 ml of DMSO and mixed. From this solution, 2.5-ml aliquots were transferred into another 384-well plate containing 47.5 ml of CH3CN. The resulting analytical plates were then centrifuged and analyzed using UPLC-MS/CAD.
Calibration samples for CAD analysis were prepared using noscapine as the universal internal standard. The noscapine calibration samples were solutions in DMSO at the following concentrations: 1.024 mg ml−1, 0.512 mg ml−1, 0.256 mg ml−1, 0.128 mg ml−1, 0.064 mg ml−1, 0.032 mg ml−1, 0.016 mg ml−1 and 0.008 mg ml−1. Each of these calibration samples, along with the reaction samples, were analyzed using UPLC-MS/CAD in duplicate.
Solvent A, 0.1% trifluoroacetic acid in Milli-Q water;
Solvent B, 0.1% trifluoroacetic acid in Optima grade acetonitrile;
Gradient, 95% A/B to 0% A/B over 1.6 min, hold 0.4 min at 95% B;
Stop time, 2.3 min;
Flow rate, 2 ml min−1;
Column, CSH C18, 2.1 mm × 50 mm.
Data processing was carried out using Virscidian software. The calibration curve, represented by a quadratic equation, was drawn on the basis of the CAD peak areas of the noscapine calibration samples. The amount of product in the analytical samples was calculated using the CAD peak area corresponding to the products in the experimental samples, in conjunction with the established calibration curve. Product yields were determined on the basis of calibrated UPLC-MS/CAD peak areas using external standard curves and are reported as assay yields.
In the validation process, 6 substrate pairs and 64 reaction conditions were employed. The reactions within the 384-well-plate format were executed according to the layout depicted below, following the procedure outlined earlier. Subsequently, we compared these results with those obtained from the 96-well plate, also following the previously described procedure. To facilitate easier comparison, the results from the 96-well plate were reformatted to align with the 384-well layout. In conclusion, out of the 384 reactions, 228 exhibited differences of less than 5% when comparing the data from the 384-well plate and the 96-well plate. Additionally, 105 reactions showed differences ranging from 5% to 10%, while 51 reactions displayed differences between 10% and 13.2%. No reactions exhibited differences greater than 13.2%.
For the experimental data generation via HTE, no statistical method was used to predetermine sample size. Experimental data were generated using a model-guided selection strategy, where reactions were chosen on the basis of prediction uncertainty. No data were excluded from the analyses. The experiments were not randomized or blinded. Regarding statistical analysis and reproducibility, the model performance was evaluated using ten independent train–test splits with different random seeds (random_state = 0–9) for the data-source benchmark, creating distinct OOD cluster assignments via k-means clustering. Each train–test configuration enforced substrate-level OOD by excluding test reactions with aryl halides or amines present in training data. Performance metrics (balanced accuracy, ROC AUC, F1 score, AUPRC, precision, recall) were computed for each split, and mean ± s.d. were reported across all runs.
For the JnJ data impact analysis, models were trained with and without industrial HTE data across ten random seeds, with results aggregated to assess performance differences. For CRDS correlation analysis, Pearson and Spearman correlation coefficients were computed between diversity scores and OOD performance metrics across data sources.
Model training used fixed hyperparameters with consistent random seeds (random_state = 42 where applicable) to ensure reproducibility. All code, processed data and analysis notebooks are publicly available to enable full reproduction of results.
Regarding experimental reproducibility, for the HTE experimental validation, 384 reactions were performed in both 96-well and 384-well plate formats across 6 substrate pairs and 64 reaction conditions. Yield quantification via CAD showed that 228/384 reactions (59%) differed by <5%, 105/384 (27%) differed by 5–10% and 51/384 (13%) differed by 10–13.2% between formats, with no differences exceeding 13.2%, demonstrating reproducible quantification across experimental platforms.
The data for this project are available via GitHub at https://github.com/schwallergroup/bh-hte-ood and via Zenodo at https://doi.org/10.5281/zenodo.19636649 (ref. 44) a pickled version of the data with additional columns Source aggregated data with BH HTE datasets from which all figures are derived are available with this paper in the aforementioned GitHub and Zenodo. Specific tables with in silico results and datasets produced during experimental validation are also available in .zip format for Figs. 2–5. Source data are provided with this paper.
The code for this project is available via GitHub at https://github.com/schwallergroup/bh-hte-ood. A static version of the code repository can be found in Zenodo at https://doi.org/10.5281/zenodo.19635488 (ref. 45).
Guram, S. A. & Buchwald, S. L. Palladium-catalyzed aromatic aminations with in situ generated aminostannanes. J. Am. Chem. Soc. 116, 7901–7902 (1994).
Kutchukian, S. P. et al. Chemistry informer libraries: a chemoinformatics enabled approach to evaluate and advance synthetic methods. Chem. Sci. 7, 2604–2613 (2016).
Article  Google Scholar 
Ruiz-Castillo, P. & Buchwald, S. L. Applications of palladium-catalyzed C–N cross-coupling reactions. Chem. Rev. 116, 12564–12649 (2016).
Article  Google Scholar 
Buskes, M. J. & Blanco, M. J. Impact of cross-coupling reactions in drug discovery and development. Molecules 25, 3493 (2020).
Article  Google Scholar 
Struble, T. J. et al. Current and future roles of artificial intelligence in medicinal chemistry synthesis. J. Med. Chem. 63, 8667–8682 (2020).
Article  Google Scholar 
Ghiandoni, G. M., Evertsson, E., Riley, D. J., Tyrchan, C. & Rathi, P. C. Augmenting DMTA using predictive AI modelling at AstraZeneca. Drug Discov. Today 29, 103945 (2024).
Article  Google Scholar 
Rinehart, N. I. et al. A machine-learning tool to predict substrate-adaptive conditions for Pd-catalyzed C–N couplings. Science 381, 965–972 (2023).
Article  Google Scholar 
de Almeida, A. F., Moreira, R. & Rodrigues, T. Synthetic organic chemistry driven by artificial intelligence. Nat. Rev. Chem. 3, 589–604 (2019).
Article  Google Scholar 
Jiang, S. et al. When SMILES smiles, practicality judgment and yield prediction of chemical reaction via deep chemical language processing. IEEE Access 9, 85071–85083 (2021).
Article  Google Scholar 
Haywood, A. L. et al. Kernel methods for predicting yields of chemical reactions. J. Chem. Inf. Model. 62, 2077–2092 (2022).
Article  Google Scholar 
Schwaller, P. & Laino, T. Data-driven learning systems for chemical reaction prediction: an analysis of recent approaches. In ACS Symposium Series Vol. 1326, 61–79 (American Chemical Society, 2019).
Neves, P. et al. Global reactivity models are impactful in industrial synthesis applications. J. Cheminform. 15, 20 (2023).
Article  Google Scholar 
Meng, Z., Zhao, P., Yu, Y. & King, I. A unified view of deep learning for reaction and retrosynthesis prediction: current status and future challenges. In Proc. Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI-23) (ed. Elkind, E.) 6723–6731 (IJCAI Organization, 2023); https://doi.org/10.24963/ijcai.2023/753
Seidl, P. et al. Improving few- and zero-shot reaction template prediction using modern Hopfield networks. J. Chem. Inf. Model. 62, 2111–2120 (2022).
Article  Google Scholar 
Segler, M. H. S. & Waller, M. P. Neural-symbolic machine learning for retrosynthesis and reaction prediction. Chem. Eur. J. 23, 5966–5971 (2017).
Article  Google Scholar 
Schwaller, P., Vaucher, A. C., Laino, T. & Reymond, J.-L. Prediction of chemical reaction yields using deep learning. Mach. Learn. Sci. Technol. 2, 015016 (2021).
Article  Google Scholar 
Saebi, M. et al. On the use of real-world datasets for reaction yield prediction. Chem. Sci. 14, 4997–5005 (2023).
Hernández, M. A. & Stolfo, S. J. Real-world data is dirty: data cleansing and the merge/purge problem. Data Min. Knowl. Discov. 2, 9–37 (1998).
Article  Google Scholar 
Lowe, D. Chemical reactions from US patents (1976-Sep2016). figshare https://doi.org/10.6084/m9.figshare.5104873 (2017).
King-Smith, E. et al. Probing the chemical ‘reactome’ with high-throughput experimentation data. Nat. Chem. 16, 633–643 (2024).
Article  Google Scholar 
Zhong, H. et al. Towards global reaction feasibility and robustness prediction with high throughput data and bayesian deep learning. Nat. Commun. 16 (2025).
Dombrowski, A. W., Aguirre, A. L., Shrestha, A., Sarris, K. A. & Wang, Y. The chosen few: parallel library reaction methodologies for drug discovery. J. Org. Chem. 87, 1880–1897 (2022).
Article  Google Scholar 
Buitrago Santanilla, A. et al. Nanomole-scale high-throughput chemistry for the synthesis of complex molecules. Science 347, 49–53 (2015).
Article  Google Scholar 
Sarker, I. H. Machine learning: algorithms, real-world applications and research directions. SN Comput. Sci. 2, 160 (2021).
Ahneman, D. T., Estrada, J. G., Lin, S., Dreher, S. D. & Doyle, A. G. Predicting reaction performance in C–N cross-coupling using machine learning. Science 360, 186–190 (2018).
Article  Google Scholar 
Xu, J. et al. Roadmap to pharmaceutically relevant reactivity models leveraging high-throughput experimentation. Preprint at ChemRxiv https://doi.org/10.26434/chemrxiv-2022-x694w (2022).
Gandhi, S. S. et al. Data science-driven discovery of optimal conditions and a condition-selection model for the Chan–Lam coupling of primary sulfonamides. ACS Catal. 15, 2292–2304 (2025).
Ren, P. et al. A survey of deep active learning. ACM Comput. Surv. 54, 180 (2021).
Google Scholar 
Lewis, M. R. et al. Development and application of ultra-performance liquid chromatography-TOF MS for precision large scale urinary metabolic phenotyping. Anal. Chem. 88, 9004–9013 (2016).
Article  Google Scholar 
Vehovec, T. & Obreza, A. Review of operating principle and applications of the charged aerosol detector. J. Chromatogr. A 1217, 1549–1556 (2010).
Article  Google Scholar 
Cardoso, T. N. C., Silva, R. M., Canuto, S., Moro, M. M. & Gonçalves, M. A. Ranked batch-mode active learning. Inf. Sci. 379, 313–337 (2017).
Article  Google Scholar 
Probst, D., Schwaller, P. & Reymond, J. L. Reaction classification and yield prediction using the differential reaction fingerprint DRFP. Digit. Discov. 1, 91–97 (2022).
Article  Google Scholar 
Teney, D. et al. On the value of out-of-distribution testing: an example of Goodhart’s law. Adv. Neural Inf. Process. Syst. 33, 407–417 (2020).
Ahmed, M., Seraj, R. & Islam, S. M. S. The k-means algorithm: a comprehensive survey and performance evaluation. Electronics 9, 1295 (2020).
Article  Google Scholar 
Hort, M., Chen, Z., Zhang, J. M., Harman, M. & Sarro, F. Bias mitigation for machine learning classifiers: a comprehensive survey. ACM J. Responsib. Comput. 1, 11 (2024).
Article  Google Scholar 
Ha, S. K. et al. Developing pharmaceutically relevant Pd-catalyzed C–N coupling reactivity models leveraging high-throughput experimentation. J. Am. Chem. Soc. 147, 19602–19613 (2025).
Article  Google Scholar 
Breiman, L. Random forests. Mach. Learn. 45, 5–32 (2001).
Article  Google Scholar 
M. Bran, A. et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell. 6, 525–535 (2024).
Article  Google Scholar 
Ramos, M. C., Collison, C. J. & White, A. D. A review of large language models and autonomous agents in chemistry. Chem. Sci. 16, 2514–2572 (2025).
Gupta, S. & Gupta, A. Dealing with noise problem in machine learning data-sets: a systematic review. Procedia Comput. Sci. 161, 466–474 (2019).
Neves, P., Wegner, J. K. & Schwaller, P. Gradient guided hypotheses: a unified solution to enable machine learning models on scarce and noisy data regimes. Preprint at https://arxiv.org/abs/2405.19210 (2024).
Karimi, D., Dou, H., Warfield, S. K. & Gholipour, A. Deep learning with noisy labels: exploring techniques and remedies in medical image analysis. Med. Image Anal. 65, 101759 (2020).
Article  Google Scholar 
Neves, P. et al. Robust out-of-distribution prediction of Buchwald–Hartwig reactions. Preprint at ChemRxiv https://doi.org/10.26434/chemrxiv-2025-xcr4 (2025).
Neves, P. et al. Collection of curated Buchwald–Hartwig HTE reactions (2026-04-17). Zenodo https://doi.org/10.5281/zenodo.19636649 (2026).
Pfvneves, PauloNeves-Git & Schwaller, P. schwallergroup/bh-hte-ood: code and data repository for robust out-of-distribution prediction of Buchwald-Hartwig reactions. Zenodo https://doi.org/10.5281/zenodo.19635488 (2026).
Download references
We thank R. Nugmanov, K. Chernichenko, J. Ash and J. Verhoeven for discussions during the early stages of this project. P.S. acknowledges support from the NCCR Catalysis (grant number 180544/225147), a National Centre of Competence in Research funded by the Swiss National Science Foundation.
Open access funding provided by EPFL Lausanne
These authors contributed equally: Paulo Neves, Bo Hao, Santeri Aikonen.
These authors jointly supervised this work: Jörg K. Wegner, Philippe Schwaller, Iulia I. Strambeanu.
Drug Discovery Data Science, In Silico Discovery, Johnson & Johnson, Porto Salvo, Portugal
Paulo Neves
Laboratory of Artificial Chemical Intelligence (LIAC), Institut des Sciences et Ingénierie Chimiques, Ecole Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland
Paulo Neves & Philippe Schwaller
Chemistry Capabilities, Analytical and Purification, Global Discovery Chemistry, Johnson & Johnson, Spring House, PA, USA
Bo Hao, Justin B. Diccianni & Iulia I. Strambeanu
Drug Discovery Data Science, In Silico Discovery, Johnson & Johnson, Spring House, PA, USA
Santeri Aikonen
Drug Discovery Data Science, In Silico Discovery, Johnson & Johnson, Cambridge, MA, USA
Jörg K. Wegner
National Centre of Competence in Research (NCCR) Catalysis, Ecole Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland
Philippe Schwaller
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
PubMed Google Scholar
HTE workflow design, B.H.; HTE execution and data analysis, B.H., J.B.D.; modeling work, P.N., S.A.; active learning, P.N., S.A.; open-source data aggregation and CRDS, P.N.; writing thepaper draft, P.N., B.H., S.A., P.S., I.I.S.; revising the paper, P.N., B.H., S.A., J.B.D., P.S., J.K.W., I.I.S.; supervision and principal investigators, J.K.W., P.S., I.I.S.
Correspondence to Jörg K. Wegner, Philippe Schwaller or Iulia I. Strambeanu.
The authors declare no competing interests.
Nature Computational Science thanks the anonymous reviewer(s) for their contribution to the peer review of this work. Primary Handling Editor: Kaitlin McCardle, in collaboration with the Nature Computational Science team.
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Sections 1–5, Figs. 1–9 and Tables 1–8.
Statistical source data for bar charts.
Statistical source data: dataset metrics and CRDS.
Statistical source data: model calibration data.
Statistical source data: experimental validation.
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Reprints and permissions
Neves, P., Hao, B., Aikonen, S. et al. Robust out-of-distribution prediction of Buchwald–Hartwig reactions. Nat Comput Sci (2026). https://doi.org/10.1038/s43588-026-01017-6
Download citation
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s43588-026-01017-6
Anyone you share the following link with will be able to read this content:
Sorry, a shareable link is not currently available for this article.

Provided by the Springer Nature SharedIt content-sharing initiative
Advertisement
Nature Computational Science (Nat Comput Sci)
ISSN 2662-8457 (online)
© 2026 Springer Nature Limited
Sign up for the Nature Briefing: AI and Robotics newsletter — what matters in AI and robotics research, free to your inbox weekly.