Drug discovery suffers from many issues including "herding," a phenomenon where companies crowd the same well-known targets and eventually lose money to both failed clinical trials and increased competition for successful drugs. Here, we suggest a protocol to avoid herding and other issues by finding new druggable pockets in underexplored proteins. We introduce the Target X Method, which uses normal mode-driven weighted ensemble molecular dynamics simulations with xenon probes to detect both cryptic and pre-existing pockets from a single static protein structure. Target X achieves a 92% success rate across 26 diverse pocket types from our validation dataset, far exceeding the sub-50% rate of conventional tools. We also introduce the Groovy Dataset, which addresses data bias and leakage in existing training sets by using SiteHopper-based clustering to focus on pocket diversity rather than global protein metrics. Using over 100,000 Spruce-prepped structures as input, and with several steps of filtering and deduplication, we prepared Groovy to contain 1,847 non-redundant protein-ligand structures. A random forest regression model (Target X Model) trained on Groovy has a 67% success rate in ranking known binding pockets among the top three candidates. Overall, Target X successfully predicts challenging cryptic pockets such as NKG2D and generates holo-like conformations suitable for virtual screening, as demonstrated on BTK kinase. Future work focuses on completing the virtual screening pipeline to go from discovered pockets to active molecules ready for lead optimization.

David LeBard, Scientific Software Director and Head of Target Exploration, OpenEye, Cadence Molecular Sciences