Molecular Design: OpenADMET: Response to Pat Walters
Tuesday, 4 August 2026
OpenADMET: Response to Pat Walters
Pat Walters responded to my post on the OpenADMET initiative and I am responding to his post. As is normal for posts here my responses are italicised in red and enclosed in square brackets. Many thanks, Pat, for taking the time to respond to my comments.<br>A recent post on the Molecular Design blog by Peter Kenny critiques the OpenADMET initiative. While constructive debate is essential for open science, the post contains several fundamental misconceptions about the roles of machine learning, structural biology, and ADMET optimization in modern drug discovery that need clarification.<br>Peter,<br>While I appreciate your advocacy for Open Science and your pragmatic view on the complexity of human biology, your critique of the OpenADMET initiative relies on several fundamental mischaracterizations of how modern computational chemistry operates and what the initiative actually aims to achieve.<br>Here are the key areas where your arguments fall short:<br>1. Dismissing ML Model Utility Ignores Current Industry Reality<br>You argue that project teams historically delivered clinical candidates without predictive models, and suggest that we need "new assays that are more predictive… as opposed to new ML models." This severely underestimates how drug discovery is actually conducted today.<br>[This is a misrepresentation of what I argued (although I concede that the points could have been more clearly articulated). I'm certainly not dismissing ML modelling in drug discovery although I remain sceptical that it will prove to be the panacea that some seem to think that it will be.<br>First, I actually stated that "drug discovery project teams have delivered (and continue to deliver) clinical development candidates without ever having sufficient data for building ML models that can accurately predict all the quantities of interest to the project teams". What I was getting at is that many of the decisions that shape the future of a drug discovery project get made before there are enough project-specific data to enable ML models to be built. This will be less of an issue if models for predicting affinity such as Boltz-2 do indeed turn out to be universal. However, I consider it less likely that universal models will be found for prediction of ADME-related properties such as aqueous solubility, permeability and turnover by CYPs.<br>Second, I stated "All that said, decisions to take compounds into clinical development are based on measurements made in a range of assays and failure in clinical development reflects an inability of these assays to predict clinical outcomes. To more effectively address attrition we actually need new assays that are more predictive of outcomes in clinical development as opposed to new ML models that are more predictive of quantities that will need to be measured anyway." This does not constitute dismissal of ML model utility and I'm merely making the point that decisions as to whether to take a compound into clinical development are based on assay measurements as opposed to predicted assay results. I certainly believe that better predictions for ADME and off-target bioactivity will lead to faster discovery of clinical development candidates. However, clinical outcomes are uncertain even when you've got a full set of measured data (bear in mind that PK/PD modelling to set doses for Phase 2 trials uses measured human PK data from Phase 1).]
ML models actively drive modern discovery: The vast majority of drug discovery programs across pharma, biotech, and academia are now directly driven by ML models. [While I would not dispute that ML models are widely used in drug discovery, your claim that ML models are actually driving the majority of drug discovery programs is in need of evidence.] Far from being academic novelties, predictive models are integrated into daily multi-parameter optimization (MPO) cycles to guide design, prioritize synthesis, and shorten cycle times before a molecule ever touches a wet-lab assay.<br>Models complement, rather than replace, assays: Models do not eliminate the need for physical assays; they ensure that expensive synthesis and wet-lab capacity are spent on the most promising molecules. Relying purely on physical assays for every ideated variant is wildly inefficient. [Agreed and I never suggested otherwise. I was simply making the point that physical assays are required to generate the data for selection of candidates for clinical development.]<br>The throughput mismatch requires models: While ADMET modeling typically focuses on local lead-optimization spaces rather than vast, multi-billion-compound virtual libraries, the number of ideated compounds evaluated in any series is still orders of magnitude larger than the number that can realistically be synthesized and assayed. Predictive models are essential to filter that ideation space down to the highest-probability candidates....