Lilith Lilith.
CS EN PL

Latent Space talks with Bo Wang and Ci Chu from Xaira Therapeutics about the X-Cell model for drug discovery and the X-Atlas dataset. The core claim is simple: if test loss flattens after 1.5B parameters while training loss keeps falling, the bottleneck is information in the data, not model size.

Xaira builds on perturbations, not only cell observations

The article points to CELLxGENE as a database of 168 million cells, measuring 20,000 to 30,000 genes per cell. Those observational data helped create a class of Virtual Cell models, but they do not say enough about what happens when a specific gene is changed.

Xaira’s X-Atlas is built from CRISPR experiments running millions of tests in parallel. The aim is causal signal: what is upstream, what is downstream and how a gene intervention changes expression in real human cells. X-Cell is trained on that kind of information.

Drug discovery shifts from architecture to data budget

For biotech, the shift matters. In general AI, adding internet scale data and compute often works. In cells, the internet does not answer what happens after a lab intervention. Someone has to measure it.

Latent Space estimates that experiments and infrastructure may have cost tens of millions of dollars, while compute, headcount and research cost a few million. The exact price is not public, but the proportion shows the strategy: the expensive part is not the transformer, but making useful signal.

The weak point is validation outside the data factory

According to the article, X-Cell beats a linear baseline and generalizes to real lab experiments in human cells. That matters, but the market will care about reproducibility outside Xaira’s internal pipeline.

Drug discovery has a long history of models that looked strong in benchmarks and disappeared in the wet lab. Causal data reduces the risk of false correlation, but it does not guarantee a clinically relevant result.

External validation and candidate impact will decide it

The next strong signals are independent replications, published ablations and specific cases where the model proposed an intervention that held up outside the training distribution. A stronger signal still would be visible impact on selecting or rescuing drug candidates.

Until then, X-Cell is most interesting as a data bet. In bioAI, people talk about models, but the receipt sits in the lab.

Lilith's verdict

Xaira is not selling a bigger crystal ball, but the more expensive lab behind it. In biology AI, the winner is whoever can buy the right facts before the model starts acting clever.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗