A-Z of AI in Healthcare

Subgroup determination

Understanding which groups of patients respond to drugs and in what way.

What is a subgroup?

Drugs often have different effects on different groups of people, both in terms of how the drug affects the patient and the size of that effect. These different groups are known as "subgroups" of the population.

What is subgroup determination?

When developing a new drug or treatment, it's important to understand which patients respond best (i.e., have the best results), which have a moderate response (i.e., OK results), and which have a negative response (i.e., bad or dangerous results, such as an allergic reaction). These three types of patients are known as subgroups, since they're smaller groups within an overall population. Working out which patients belong in which group is a process known as subgroup determination.

How can AI help with subgroup determination?

AI, particularly unsupervised AI, can support this process. For example, you might give an unsupervised algorithm an unlabelled dataset containing the electronic health records of all patients enrolled in a clinical trial, including the results of the new drug or treatment (e.g., lowering blood pressure). The algorithm can then group similar patients according to their results, a task known as clustering, and use this to determine subgroups and assign patients accordingly.

Using AI this way can be valuable because it might indirectly identify potential causes for different patient responses that weren't immediately obvious. For example, it might reveal that all patients in the "good response" group were aged 18–45, those with a moderate response were 45–65, and those with a poor response were over 65. This kind of insight is useful because it could inform future clinical guidance, for example, recommending the drug as a primary treatment for younger patients with high blood pressure, but not as the primary treatment for older patients.

Why does subgroup determination matter?

Identifying subgroups and the variations in their treatment response is a crucial step in drug discovery, drug safety monitoring, and the development of personalised medicine. This can be done in two ways: prospectively (ahead of a clinical trial) or post-hoc (after the trial has started, during real-world testing of the drug).

What is prospective subgroup identification?

Prospective subgroup identification, formally known as confirmatory subgroup analysis, involves identifying a small number of predefined covariates, typically demographic patient characteristics known as biomarkers, that are listed in the registered trial protocol. For example, COVID-19 vaccine trials recruited participants and tested different vaccines across various age groups.

What is post-hoc subgroup identification?

Post-hoc, or exploratory subgroup analysis, is a more data-driven process that increasingly relies on machine learning techniques, such as unsupervised clustering algorithms, to discover new subgroups by analysing a large number of covariates (demographic, genomic, clinical, or other patient characteristics) and their impact on treatment response.

This process is important not just for identifying which patients might benefit most from a specific drug, but also for preventing overfitting and bias, and making estimates of treatment effect size more honest.

Are there established guidelines for this kind of analysis?

Not fully. There's no single agreed method for conducting post-hoc exploratory subgroup analysis. While guidelines exist for prospective confirmatory subgroup analysis, no equivalent guidelines yet exist for data-driven subgroup discovery. These are likely to develop over time, particularly as the push toward personalised or precision medicine gathers pace.

In the meantime, it's important that anyone conducting post-hoc analysis, especially using machine learning techniques, documents their process carefully, follows basic best-practice guidelines for statistical analysis (e.g., avoiding p-hacking), and considers relevant responsible AI guidelines.

In Practice

Curious how these principles are being put into practice?

Owkin is building agentic AI and biological reasoning models to better understand biology and advance biological superintelligence.

K Pro

Owkin's agentic AI co-pilot, applying biological reasoning to real biopharma research and decision-making.

Discover K Pro