-
Technology
-

Virtual Biology Initiative Pools $1.8 Billion to Build Open AI Training Data for Drug Discovery

By
Distilled Post Editorial Team

A public-private partnership has pledged $1.8 billion (£1.34 billion) to create large-scale biological datasets for training artificial intelligence models. The scheme, called the Virtual Biology Initiative, will combine government, philanthropic and corporate funding to generate standardised data on how cells respond to a wide range of conditions. Supporters say the initiative could reduce drug development timelines from years to just months.

The Chan Zuckerberg Biohub provided the starting capital, pledging $500 million (£373 million) in April. The US Department of Energy has pledged more than $500 million (£373 million) over five years to fund laboratory measurements, computational modelling and data processing.A further $500 million or more in federal money, coordinated by the National Institutes of Health, will go towards standardising existing datasets and repositories. Meta Platforms, Google DeepMind and Isomorphic Labs have together committed $300 million (£224 million) to the commercial effort.

The stated ambition is to turn biology from a discipline built on observation and discovery into one that can predict outcomes. Researchers would be able to ask how a given cell will behave under a given condition and receive a reliable answer from a model, rather than running each experiment by hand. That requires far more data than exists today. Current biological datasets cover hundreds of millions of cells. Predictive models of the kind the initiative envisages need billions, and possibly trillions.

Organisers say generating data on this scale would typically take decades, but the initiative aims to complete the work within five years, with the first dataset expected to be released in about a year.

Two experimental approaches will supply much of the new material. Spatial transcriptomics maps molecular activity inside intact tissue, which preserves information about where each cell sits and what surrounds it. That context shapes cell behaviour and is lost when tissue is broken apart for analysis. The second approach is high-throughput screening, which records how cells respond to controlled changes in their environment. Running such experiments at scale produces the variety of conditions that predictive models need to learn from.

Existing data presents a different problem. Laboratories around the world have generated large volumes of experimental results over many years, but they were collected for individual studies, with little coordination on format or design. Much of the funding, including support from the NIH, will go towards standardising the material into unified formats that can be used directly by AI systems.

The access terms reflect the mix of funders. The long-term commitment is to release every dataset publicly, so that it serves as a shared resource for the research community. Commercial backers, however, will have exclusive early access for a limited period before each dataset is released. Data generated directly through US government funding will be exempt from this arrangement, with no embargo and immediate public release.

The arrangement gives the technology companies a head start on the data they have helped to pay for, while keeping the underlying material available to universities, smaller firms and public laboratories once the early-access window closes. How long that window will last has not been detailed in the initiative's announcement.

The Biohub says it plans to seek additional financial backing from pharmaceutical companies and philanthropic foundations. The total could therefore rise above $1.8 billion as the programme develops.

The initiative arrives as major AI laboratories compete to secure biological data of their own. Anthropic has established wet laboratory facilities, which allow it to generate experimental results directly rather than relying on published sources. The OpenAI Foundation has launched grant funding programmes aimed at biological research. The Virtual Biology Initiative differs from both in its scale of pooled funding and in its commitment to eventual open release, though it shares the same underlying premise: that the quality of biological AI models will depend heavily on the quantity and consistency of the data used to train them.

‍