The $110 million bet on cells, not just algorithms
When GSK expanded its partnership with London-based Relation Therapeutics in a deal worth up to $110 million, the pharma giant wasn’t just buying another AI platform. It was paying for something far scarcer: high-quality biological data.
The collaboration centers on generating large-scale datasets that measure how human cells respond to genetic changes and drug interventions. That data will train AI models to spot potential drug targets, including those within Relation’s MORGAN platform.
It’s a telling shift. For years, the buzz in drug discovery has been about smarter algorithms. Now, the bottleneck is moving to the fuel that powers them.
What Relation Therapeutics actually does
Relation calls its approach Lab-in-the-Loop. It’s a cycle: run lab experiments, feed the results into computational models, use those models to design the next experiment. The company handles tissue profiling, single-cell and spatial transcriptomics, sequencing, and target validation. Machine learning sits at the center, guiding target identification and experimental design.
Its perturbation experiments are particularly interesting. These measure how genetic tweaks alter cellular characteristics tied to disease. The results get analyzed alongside genetic and patient-derived data, creating a richer picture than any single dataset could offer.
This isn’t theoretical. Relation’s Osteomics project is a proprietary functional single-cell bone atlas, built from patient samples and combining single-cell and spatial omics with imaging, genomics, proteomics, and clinical phenotypes. It’s already being used to investigate osteoporosis biology and identify patient subgroups, with hospitals in the UK and Australia involved.
Why bigger datasets don’t automatically mean better AI
Here’s where things get counterintuitive. You might assume more data always helps. A June 2025 study in Nature Methods suggests otherwise.
Researchers trained 400 single-cell foundation models on a corpus of 22.2 million cells, evaluating them across 6,400 experiments. The result? Models hit performance plateaus after training on only a fraction of the available data. Unlike large language models, these systems didn’t show clear scaling laws where more data consistently led to better results.
The study concluded that model capacity, dataset size, and compute need to be balanced—not just cranked up together. Adding more biological data didn’t reliably produce better models.
A separate 2025 study in Genome Biology evaluated Geneformer and scGPT, two well-known single-cell foundation models. Neither consistently outperformed simpler approaches. The researchers also flagged batch effects and warned against assuming larger pretrained models automatically yield better biological representations.
Quality control is the real challenge
The problem isn’t just volume. It’s consistency. A 2025 review in Experimental & Molecular Medicine noted that public repositories like CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus offer vast amounts of single-cell data—CZ CELLxGENE alone has over 100 million standardized cells.
But that data comes from hundreds of labs, each with its own sampling methods, sequencing protocols, and processing pipelines. Technical noise and artifacts are everywhere. Dataset overlap is another headache: the same cells can appear in multiple resources, giving them outsized influence during training and creating data-leakage risks when training and test sets overlap.
The review’s conclusion was blunt: assembling a high-quality, non-redundant dataset matters just as much as model architecture.
Pharma’s new playbook: specialized data as a strategic asset
This is why companies like GSK are paying premium prices for proprietary data. A 2025 Nature Biotechnology analysis of AI-focused biopharma deals identified specialized dataset providers as a key trend, alongside larger upfront payments and new therapeutic modalities.
The analysis pointed to several examples:
- GSK’s separate $37.5 million agreement with Ochre Bio for human liver single-cell and perfused-organ data
- AstraZeneca and Pathos AI’s $200 million deal with Tempus in 2025, covering de-identified clinical, genomic, and imaging data from over 150,000 patients
- Relation’s own Osteomics atlas, built specifically for osteoporosis research
These deals reflect a simple reality: high-quality, disease-specific datasets are becoming a critical input for causal and generative machine-learning models. Public data has its place, but it’s not enough.
The data bottleneck won’t disappear anytime soon
A Nature research highlight on federated learning in pharma called limited access to suitable training data a major bottleneck for AI applications. Companies face restrictions on sharing proprietary information, and even when they want to collaborate, the data infrastructure often isn’t there.
That’s why AI-biopharma agreements take so many forms. Some focus on accessing AI platforms. Others cover joint development or data licensing. The GSK–Relation deal does both: it funds data generation and model development, with Relation producing human cellular datasets and using them to train target-identification models.
The message is clear. In AI drug discovery, the algorithm is no longer the differentiator. The data is. And the companies that control the best biological datasets—not the biggest ones—will likely lead the next wave of discoveries.
For more on how AI is reshaping pharma, check out how AI is shortening drug discovery timelines in China and AI foundation models in biomedicine.