Blog

August 7, 2026
|
5 mins

Beyond the model: Building AI systems that science can trust

The AI market rewards certainty. Science does not.

Yet much of today's discussion about AI assumes those two worlds move at the same pace. They don't.

In Zurich, our previous AI thought leader dinner explored why AI in biology must be grounded in evidence rather than persuasive outputs. By the time we gathered again in Paris, the conversation had shifted: if evidence is the foundation, how should pharma R&D build and govern the next generation of AI?

It felt like the right question to ask: what kind of AI systems should underpin the future of pharmaceutical R&D?

Before we can ask how to govern AI, we need to be much clearer about what has actually been built, what it can reliably do, and where responsibility still sits.

Part of the challenge is that AI is increasingly treated as though it were a single technology. In reality, different systems have fundamentally different capabilities, limitations, and risks. They shouldn't all be evaluated or governed the same way.

There's growing pressure to make AI sound immediate, autonomous, and inevitable. Investors want clear benchmarks, companies want compelling use cases, and governments want to move quickly. But in drug discovery, speed without scientific validation is not progress.

A frontier model can be extremely powerful, whether deployed as a general-purpose assistant or within a specialized scientific platform. But architecture alone doesn't make a system scientific. To answer meaningful biological questions, it needs the right patient and molecular data, specialized analytical tools, domain-specific models, and workflows grounded in experimental validation.

We've moved past being impressed by a great generated answer to a broad scientific question. What matters is whether AI can identify a novel biomarker, uncover a treatment-response signal, or reveal a clinically relevant patient subgroup that would otherwise go unnoticed.

This matters because not all AI creates the same level of scientific or clinical risk. Summarizing a document is fundamentally different from prioritizing a drug target or shaping clinical trial design. As AI moves closer to consequential scientific decisions, the bar for evidence and oversight must rise.

One comparison from the discussion captured this well: give a quantitative analyst better tools and they may become ten-times more effective. But they still select the strategy and own the risk.

AI can work the same way for scientists, allowing them to investigate more hypotheses, interrogate more data, and test more possible paths in parallel. It gives science more shots on goal without removing responsibility. Scientists still decide what's biologically credible, what's sufficiently validated, and what's worth taking forward. Human oversight shouldn't be a ceremonial approval tacked onto the end of an automated process. It needs to be designed into the workflow from the outset.

We also kept returning to the question of trust.

The future of biological AI depends on access to large, deeply characterized datasets. But people can't be expected to contribute health or genomic data without understanding what they're providing, who can use it, and how they stand to benefit from the models it helps create. That's not a tech issue, it’s a cultural one.

Better models require better data. Better access to data requires clear governance. And clear governance requires clarity about the value, risk, and accountability attached to each specific use case rather than a single policy applied to everything we call "AI."

If we get that right, governance becomes more than risk management. It becomes an enabler of better science: giving researchers confidence to use AI where it adds value while ensuring the underlying science remains rigorous enough to support the decisions that matter most.

By the end of the evening, the question we came to Paris with had evolved:

How do we move from possibility to proof by building systems that accelerate science while preserving scientific rigour, trust, and human responsibility at every layer?

The market may reward certainty. Biology rewards evidence, careful reasoning, and scientific proof. AI systems designed for science should do the same.

We'll keep exploring these themes and more at the next dinner in the series.

Thank you to Delphine Gény-Stephann, Romain Curunet, Constantin Landers, Pegah Esmaeili, André M. König, Kaïs Zhioua, Zoe (Ziwen) Qin, Malik Mitha, Georgios Stamatas, PhD, and Radu Vadanici for joining us and bringing perspectives from across science, technology, investment, and industry. And thank you to Lucie, Andrea, Jonas, Fayssal, and Denis from Owkin for leading the discussion, and to AWS for continuing to build this series with us.

Authors

Sepideh Shokrpour
No items found.