In this thought leadership interview, Eray Watts, Distinguished Scientist, High-Throughput Chemistry at insitro, explores how high-throughput chemistry, DNA-encoded libraries and machine learning are coming together to reshape small-molecule drug discovery.
Watts explains that his work at insitro focuses on developing laboratory techniques and protocols capable of producing fit-for-purpose, high-quality datasets for machine learning. Rather than building broad models intended to identify drugs for any disease or target, insitro takes a more focused approach, developing high-resolution models designed to support individual medicinal chemistry campaigns.
A key component of this strategy is the use of DNA-encoded libraries (DELs). Insitro has developed technologies to create more focused and tailorable libraries, alongside screening approaches that generate greater dimensionality in the resulting data. Capturing information with direct measurements from DEL screening can provide richer datasets and enable more predictive models.
These models can then support the medicinal chemistry design–make–test–learn cycle, helping researchers make better decisions about which molecules should be synthesised and tested. By considering target engagement alongside properties such as ADMET, machine learning can help reduce the number of unnecessary compounds that need to be made and tested while improving the efficiency and impact of discovery campaigns.
Looking ahead, Watts suggests that the biggest opportunity may not necessarily be developing increasingly sophisticated AI architectures. Instead, the availability and quality of data could become the critical limiting factor. As machine learning models become increasingly capable and data-hungry, finding new ways to generate larger quantities of higher-quality data more efficiently and cost-effectively could have a major impact on small-molecule discovery.
For Watts, the future of AI-enabled drug discovery therefore depends not only on better models, but on the ability to generate the rich experimental data needed to unlock their full potential.







