Skip to content
Live newsroom 23 readers online
Friday, September 4, 2026 Live Sync: Just now
Breaking‘Obviously adds a bit of weight’: Butters reveals chat with Butters over potential Cats team up
Important AVOID AMZN Stage 4 (Conv: 3/5 | Size: 10%)

AI-enhanced adaptive virtual screening of large libraries for ligand discovery

Nature Biotechnology (2026) Cite this article Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of […]

By deepak · September 1, 2026 · 9 min read

Nature Biotechnology
(2026) Cite this article

Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.

A central challenge in early-stage drug discovery is the identification and optimization of initial hit and lead compounds that bind specifically and potently to biological macromolecules. Historically, experimental high-throughput screening (HTS) has served as the primary method for discovering such hits1. While these approaches have yielded many successful leads, they are inherently constrained by high costs, long timelines and the need for robust, scalable assays. These limitations are further compounded by the vastness of chemical space: the number of synthetically accessible, drug-like small molecules is estimated to exceed 1060 (ref. 2), yet typical HTS campaigns screen only hundreds of thousands of compounds. This minuscule sampling poses a major bottleneck, particularly for challenging targets such as protein–protein interaction interfaces or allosteric sites, where binding pockets are often shallow, dynamic or poorly defined. Virtual screenings have emerged as a compelling alternative, offering the ability to computationally evaluate extremely large numbers of molecules. Virtual screening surpassed millions of compounds over a decade ago3 and, in recent years, has expanded to hundreds of millions to billions of molecules in silico, giving rise to ultralarge virtual screenings (ULVSs). Compared to traditional HTS, ULVSs dramatically reduce cost and time, while expanding the scope of chemical space that can be explored. In addition to identifying potent candidate compounds, virtual screening can provide insights into binding mechanisms and structure–activity relationships, making it especially valuable for targets that are difficult to modulate using conventional experimental approaches.

In recent years, multiple studies have demonstrated the remarkable effectiveness of ULVSs in identifying potent hits across a wide range of protein targets, including enzymatic active sites, orthosteric sites of G-protein-coupled receptors and protein–protein interaction interfaces4,5,6,7,8. Driven by the success of billion-compound screens and advancements in synthesis and computational technologies, the number of available on-demand molecules with an established synthetic route has expanded rapidly, with current estimates reaching up to 3 trillion compounds across multiple vendors9.

While ULVSs enable the exploration of vast chemical libraries, notable challenges remain in making these searches efficient, affordable and intelligent. Navigating such immense chemical space to rapidly identify promising hits requires additional strategies that go beyond brute-force docking. The financial cost alone is a major bottleneck8, rendering it impractical as chemical databases continue to grow exponentially, with newer libraries exceeding 1 trillion compounds and growing. Moreover, effective virtual screening requires the ability to flexibly integrate a wide range of docking algorithms, each built on different theoretical foundations and requiring distinct input formats, to better match the characteristics of diverse target classes. Compounding these challenges is the need to fully leverage modern computational infrastructure; while many existing platforms are limited to CPUs, there is increasing demand to exploit both CPU clusters and GPU architectures to scale performance and reduce turnaround time. These limitations—high cost, limited flexibility and insufficient scalability—highlight the urgent need for a new solution. To end this, we developed AdaptiveFlow, an open-source platform designed to intelligently navigate ultralarge chemical spaces, seamlessly incorporate diverse docking protocols and operate efficiently across heterogeneous computing environments, including cloud-scale CPU and GPU resources.

A key innovation of the platform is the implementation of Adaptive Target-Guided Virtual Screens (ATG-VSs), which use an 18-dimensional grid of molecular properties to identify and hone in on the most promising regions of chemical space, favorable for a chosen target site, substantially reducing computational cost relative to exhaustive docking. An optional active learning layer can be added to further refine and accelerate this process by iteratively selecting the most promising compounds of each tranche. To demonstrate the utility of AdaptiveFlow, we applied it to two therapeutically relevant targets: ferroptosis suppressor protein 1 (FSP1), a key regulator of lipid peroxidation and ferroptotic cell death with emerging roles in cancer resistance10, and poly(ADP-ribose) polymerase 1 (PARP1), a well-established DNA repair enzyme and clinical target in oncology11. In both cases, the platform identified nanomolar inhibitors, which were experimentally validated through high-resolution co-crystal structures confirming target engagement.

Deploying ULVSs at the scale of billions of compounds requires not only massive computational resources but also software systems that can efficiently coordinate the preparation, management and docking of libraries across heterogeneous high-performance computing environments. To meet these requirements, we developed AdaptiveFlow, an open-source platform built to streamline, scale and intelligently guide ULVS workflows from end to end. AdaptiveFlow combines deep flexibility with performance, offering support for diverse ligand and receptor types, GPU acceleration and modular integration of classical and machine learning (ML)-based docking protocols.

AdaptiveFlow introduces major enhancements across three core components: the AdaptiveFlow Ligand Preparation (AFLP) module for efficient preprocessing of massive libraries; the AdaptiveFlow for Virtual Screening (AFVS) engine, which supports over 1,500 docking protocols; and the AdaptiveFlow Unity (AFU) module, which integrates these components into a unified, modular workflow adaptable to both CPU-based and GPU-based infrastructures (Fig. 1 and Supplementary Table 1). These modules have been engineered to support both traditional and emerging use cases while optimizing for large-scale parallel execution on cloud and on-premise infrastructures. AdaptiveFlow is natively compatible with cloud services such as Amazon Web Services (AWS) (Supplementary Fig. 1), where it demonstrates near-linear scaling up to 5.6 million virtual CPUs (vCPUs) (Supplementary Fig. 2), enabling efficient exploration of ultralarge libraries. The use of spot (preemptible) instances further reduces the costs.

AdaptiveFlow introduces several key advances that greatly expand the capabilities of ULVS, including (1) curation of the Enamine REAL Space library, comprising 69 billion molecules in a ready-to-dock format compatible in multiple chemical file formats; (2) integration of an extensive suite of approximately 1,500 docking protocols (defined as specific combinations of sampling algorithms and scoring functions), enabling the targeting of a broad range of receptor types, including RNA and DNA; (3) support for emerging hardware architectures such as ARM CPUs, along with GPU acceleration to enhance computational performance; (4) advanced ligand preparation (prep.) features, including stereoisomer enumeration and the calculation of 28 molecular (mol.) properties, to enable the creation of more comprehensive screening libraries; and (5) built-in support for ML and deep learning approaches, including compatibility with SELFIES and neural-network-based docking protocols. The integrated module, AFU, further streamlines workflows by combining ligand preparation and docking into a single environment. Importantly, AdaptiveFlow also introduces a novel screening paradigm known as ATG-VSs, which enables efficient, target-focused exploration of ultralarge compound libraries at a fraction of the computational cost typically associated with conventional ULVSs. DL, deep learning; Multidim., multidimensional.

The AFLP module provides a pipeline for preparing massive compound libraries for docking. In addition to generating three-dimensional (3D) conformers, it enumerates stereoisomers and tautomers, validates geometries and calculates 28 molecular properties (Fig. 2, Extended Data Figs. 1–3, Supplementary Tables 2 and 3 and Supplementary Figs. 3–10). These properties are used to disperse the molecules within a multidimensional grid, up to 18 dimensions, on the basis of physicochemical descriptors. This grid-based representation forms the foundation for ATG-VSs, which allow users to intelligently select and prioritize chemically diverse subspaces relevant to a given target, thereby optimizing computational resources focusing on the most promising regions of chemical space. Furthermore, an optional active learning component can be layered on top of ATG-VSs to iteratively refine the search using data-driven feedback, dynamically selecting molecules predicted to yield the highest information gain or binding likelihood (Supplementary Figs. 1–3 and 11 and Extended Data Fig. 4). Because each screening stage generates new labeled data, this cycle can be repeated for an arbitrary number of rounds, retraining the model on the newly docked and, where available, experimentally assayed compounds and re-selecting the next batch to screen, thereby forming a closed-loop active-learning cycle that progressively concentrates sampling on the most promising regions of chemical space.

The initial version of the REAL Space (enumerated release, 2022q1-2) comprised approximately 31 billion molecules, represented as SMILES strings in an unprocessed format (for example, without full stereoisomer enumeration or structural standardization). A selected subset of 18 of these properties were used to partition the library into an 18-dimensional property matrix, where each dimension was discretized into multiple intervals. In this figure, the high-dimensional matrix is illustrated as a 3D schematic for clarity, with three representative properties each split into 6–10 intervals, resulting in 8 × 6 × 6 = 288 tranches. The actual prepared REAL Space matrix has around 12 million occupied tranches. Each box represents a tranche that contains ligands sharing the corresponding property ranges. This structure enables flexible selection of compound subsets and serves as the foundation for ATG-VSs that are introduced in this work.

The AFVS module enables large-scale docking campaigns using over 1,500 distinct docking protocols. We define a docking protocol as a specific combination of a sampling algorithm (pose prediction) and a scoring function, allowing for a vast number of unique workflows assembled from more than 40 individual docking programs. AFVS supports a wide range of ligand and target types, including peptides, DNA and RNA, and includes recent advances in ML-based docking such as DiffDock and TANKBind12,13, which have shown promise in speed and accuracy, especially for blind or flexible docking scenarios (Supplementary Tables 4–9 and Supplementary Figs. 12 and 13).

To unify and simplify large-scale workflows, we developed AFU, a module that integrates the functionalities of AFLP and AFVS into a cohesive interface. AFU enables users to perform ligand preparation and docking within a single automated pipeline, dramatically reducing setup complexity and increasing accessibility for nonspecialists. This streamlined design supports rapid iteration and facilitates integration with other tools in computational drug discovery and artificial intelligence (AI)-driven design workflows (Supplementary Figs. 3–7 and 14 and Methods).

Virtual screening performance scales notably with library size, because larger screening sizes tend to yield both higher potencies and improved hit rates (that is, the number of confirmed hits divided by the number of tested compounds)4,6,7,8. Until recently, the largest publicly available ready-to-dock libraries, including the 2018 version of the REAL Database8 and the ZINC20 library14, each contained approximately 1.5 billion molecules. To enable virtual screens of an order-of-magnitude larger scale, we prepared the 2022 version of the Enamine REAL Space into a ready-to-dock format, comprising 68.7 billion commercially available on-demand small molecules.

The current version of the Enamine REAL Space (2022q1-2, accessed 15 November 2022) contains 31,507,987,117 enumerated molecules (becoming 69 billion after ligand preparation, for example, because of stereoisomer and tautomer enumeration). These compounds are synthetically accessible derivatives of 137,000 established building blocks, assembled using 167 internally validated one-pot synthetic protocols developed at Enamine15,16,17,18,19,20,21,22. This combinatorial framework enables high-throughput compound generation with a reported synthetic success rate of 82%, on the basis of 386,000 experimental reactions conducted in 2021. In total, 99% of the compounds in the library conform to Lipinski’s Rule of Five23. Structurally, it offers exceptional chemical diversity, comprising 1,098,811,629 unique Murcko scaffolds.

Source: Read the original article on www.nature.com

Important Legal & Financial Disclaimer

FutureKnowledge is an automated financial intelligence aggregator. The information provided on this website does not constitute investment advice, financial advice, trading advice, or any other sort of advice and you should not treat any of the website's content as such. We are not registered with the SEC, SEBI, or any regulatory agency. Automated AI-generated content may contain errors. Always conduct your own due diligence and consult your financial advisor before making any investment decisions.

© 2026 FutureKnowledge Intelligence. All rights reserved.