A-Z of AI in Healthcare

Large Language Model

A type of AI trained on large datasets that generates text or images.

What is LLM?

Large Language Models (LLMs), like ChatGPT or DALL-E, are a type of "generative artificial intelligence," or machine learning, trained on enormous datasets to produce fake data in a range of formats, including text and image data. They do this in response to a simple written instruction known as a "prompt," with the data they produce looking almost identical to real data.

How does an LLM actually work?
  1. You provide the model with a prompt, for example, "describe the uses of data in healthcare."
  2. The text is automatically divided into words and smaller units, which are encoded into a format the model can process.
  3. The model uses this code, combined with the context of the prompt (e.g., your previous conversation history), to capture the meaning of your input.
  4. The model generates text in an autoregressive manner, meaning it predicts the next word based on the context of your input and the text it has already generated, drawing on everything it learned during training.
  5. This process repeats until a stopping condition is met (e.g., a maximum length or the end of a sentence).
  6. As you keep interacting with the model, it builds a better understanding of your input by factoring in your entire conversation history, which leads to more meaningful replies.
What is generative AI capable of, more broadly?

Popularised by models like ChatGPT and DALL-E, generative AI, including LLMs, is a form of machine learning (typically deep learning) capable of generating new data across a range of formats and adapting to new tasks in real time using simple written prompts. This makes it considerably more flexible than "narrower" AI, since one model can complete many tasks without needing to be retrained.

How can generative AI be used in medical research?

Generative AI can assist researchers at almost every stage of the research process, including:

  • Summarising previous research
  • Identifying knowledge gaps
  • Generating hypotheses
  • Helping design experiments
  • Helping write analytical code
  • Drafting, editing, and disseminating research papers

A few uses are worth exploring in more depth:

Synthetic data generation
Medical research typically depends on access to patient data, both structured (e.g., electronic health records, imaging) and unstructured (e.g., free-text clinical notes). But patient data is highly sensitive, often affected by missing values, and frequently biased or unrepresentative of real clinical populations. Generative AI (particularly GANs and VAEs) can learn the statistical structure of patient data and use it to generate synthetic datasets that closely resemble the real thing, without exposing real patient information. One study found generative AI could produce synthetic retinal images indistinguishable from real patient images to expert consultants. Synthetic data can also be used for "augmentation," filling gaps caused by missing or biased data.

Data harmonisation
Building large, representative research datasets usually requires linking datasets that use different formats, a complex, expensive, and labour-intensive process. Generative AI models, including LLMs, represent all data types uniformly, typically as a list of numbers (or a vector). This process, called "embedding," has considerable potential to make data harmonisation faster and more cost-effective.

Drug discovery
Generative AI isn't limited to text and images. It can also generate novel small molecules, nucleic acid sequences, and proteins with a desired structure or function. This allows researchers to test a range of potential therapeutics quickly for a specific condition or patient profile, speeding up personalisation and optimisation.

Clinical trials
Synthetic data can simulate patient populations, treatment groups, and outcomes within clinical trials, helping researchers optimise trial design, estimate population-level efficacy and safety, and identify potential biases or confounders.

Can generative AI be used in direct patient care?

Potentially, but this is a much earlier-stage application than research. From a direct-care perspective, generative AI, and LLMs in particular, could be used as triaging or diagnostic chatbots, to help overcome language barriers in patient-doctor communication, or to take notes during consultations.

These use cases raise complex ethical and regulatory challenges. Any LLM used in direct care will need to be regulated as a medical device, but it's still unclear how medical device law will apply to generative AI, or how developers will be expected to meet requirements like "evidence of efficacy." It's likely to be a long time before generative AI meaningfully impacts direct care at scale. In contrast, it's already having an impact on medical research today.

What are the risks and limitations of generative AI in medical research?

Generative AI is a "dual-use" technology. It can be used for beneficial or harmful purposes, and it comes with real limitations alongside its benefits:

  • Synthetic data may still be biased, and may remain vulnerable to re-identification attacks (such as model inversion attacks) that threaten patient privacy.
  • Generative AI has no genuine "semantic understanding." It may summarise information incorrectly, or "hallucinate" false information, including fake citations, when generating text.

This is why it's important to approach generative AI with curiosity, but also caution.

In Practice

Curious how these principles are being put into practice?

Owkin is building agentic AI and biological reasoning models to better understand biology and advance biological superintelligence.

K Pro

Owkin's agentic AI co-pilot, applying biological reasoning to real biopharma research and decision-making.

Discover K Pro