When healthcare organisations discuss bias in AI, the conversation often gravitates towards training data. That is understandable because the data matters enormously. But it captures only part of the problem.
Bias can enter when we decide which clinical problem to solve, choose the outcome a model should predict, or determine who is represented in the data. It can also emerge during validation or much later, when a technically sound model meets a different hospital, workflow or patient population.
Equity, therefore, needs to be considered across the entire AI-enabled care process.
Where can bias enter the healthcare AI lifecycle?
AbrĂ moff and colleagues map equity across conception, design, development, validation, access and monitoring. Hasanzadeh and colleagues reach a similar conclusion, tracing potential bias from conception through to clinical deployment and ongoing surveillance.
The message is clear: addressing bias at one stage does not guarantee equity at the next.

Source: AbrĂ moff et al., npj Digital Medicine (2023)
Bias can start with what an AI model predicts.
One of the most important decisions happens before an AI model is trained: deciding what it should predict.
The Obermeyer study examined an algorithm designed to identify patients who might need additional care. Race was not included in the model. Instead, it used expected future healthcare spending as a measure of how much care a patient needed.
The problem was that spending did not always reflect clinical need. Black patients with similar levels of illness had historically received less healthcare and generated lower costs. As a result, the algorithm could accurately predict spending while underestimating the amount of care Black patients needed.
The bias was therefore not simply in the data or the model itself. It came from using healthcare spending as a stand-in for health needs.
Healthcare data reflects the system that created it.
Healthcare data does not exist in isolation. It reflects who received care, what clinicians recorded, which tests were ordered and how healthcare services were used.
This can influence how AI performs. Larrazabal and colleagues found that when chest X-ray datasets were imbalanced between male and female patients, model performance declined for the underrepresented group.
The STANDING Together recommendations reinforce the importance of understanding who is represented in healthcare datasets and being transparent about where gaps could affect particular populations.
But who is represented in the data is only part of the challenge.
Can healthcare AI generalise between hospitals?
AI can also learn patterns based on how and where healthcare data was collected.
Ly and colleagues analysed data from more than 200,000 patients and found that models could detect hidden differences in how the data was generated. As a result, testing a model on data from the same environment in which it was developed could overestimate its performance elsewhere by up to 20% on average.
This means a model that performs well in one hospital may not perform the same way in another, where the equipment, workflows or patient population are different.

Source: Hasanzadeh et al., npj Digital Medicine (2025)
Not every performance difference means the same thing.
Jones and colleagues show that finding a difference in model performance is only the first step. Similar differences can have very different causes, from genuine variations in disease risk to gaps or imbalances in the data.
For example, a real relationship between age and disease risk needs to be treated differently from a pattern caused by unequal access to healthcare.
This means there is no single measure or solution for every type of AI bias. Understanding why a difference exists and the clinical context behind it is essential.
Bias can emerge after deployment.
Once AI enters clinical practice, another set of variables becomes active.
Hasanzadeh and colleagues identify risks, including automation bias and feedback loops. AbrĂ moff and colleagues extend the issue further to access: what happens if an AI system correctly identifies a patient, but that patient cannot access the next stage of care?
A technically equitable prediction does not necessarily produce an equitable outcome.
Regulation is increasingly reflecting this lifecycle view. The FDA’s guidance for AI-enabled medical devices takes a Total Product Life Cycle approach, while WHO guidance similarly emphasises governance and monitoring beyond development.
What can averages hide?
There is also an emerging question about aggregate performance.
A recent preprint describes an “average patient fallacy”, where strong overall performance can coexist with weaker reliability for rare or underrepresented patients. The authors propose greater attention to individual uncertainty and reliability.
As a preprint, this should be treated as an emerging design proposition rather than an established consensus. But it raises an important question: who might be hidden by the average?
Equity needs to be assessed throughout the AI lifecycle.
For healthcare technology companies, assessing equity needs to be an ongoing part of how AI is developed and used.
Teams need to understand who the technology is designed for and what it is trying to achieve. They also need to know where the data comes from, who is represented, how the technology performs across different populations and settings, and whether that changes after deployment.
Doing this builds stronger evidence for procurement and clinical governance, while helping identify problems before they are repeated across different health systems.
Bias in healthcare AI cannot be addressed once and forgotten. It needs to be assessed throughout the lifecycle.
Bring explainable AI into clinical workflows.
SMARTSuite helps clinicians find relevant information within complex patient records, with outputs traceable to their source data and clinicians remaining in control of clinical decisions.
Explore SMARTSuite and see how AI-enabled intelligence can support healthcare teams.
Authored by Tom Varghese, Global Product Marketing & Growth Manager at Orion Health.
References
- AbrĂ moff MD, Tarver ME, Loyo Berrios N, et al. 2023. Considerations for addressing bias in artificial intelligence for health equity. npj Digital Medicine. 6:170.
- Alderman JE, Palmer J, Laws E, et al. 2025. Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations. The Lancet Digital Health. 7(1):e64–e88.
- Fard P, Azhir A, Rezaii N, Tian J, Estiri H. 2025. An N of 1 Artificial Intelligence Ecosystem for Precision Medicine. arXiv. Preprint 2510.24359.
- Hasanzadeh F, Josephson CB, Waters G, Adedinsewo D, Azizi Z, White JA. 2025. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. npj Digital Medicine. 8:154.
- Jones C, Castro DC, De Sousa Ribeiro F, Oktay O, McCradden M, Glocker B. 2024. A causal perspective on dataset bias in machine learning for medical imaging. Nature Machine Intelligence. 6:138–146.
- Larrazabal AJ, Nieto N, Peterson V, Milone DH, Ferrante E. 2020. Gender imbalance in medical imaging datasets produces biased classifiers for computer aided diagnosis. Proceedings of the National Academy of Sciences. 117(23):12592–12594.
- Ly CO, Unnikrishnan B, Tadic T, et al. 2024. Shortcut learning in medical AI hinders generalization: method for estimating AI model generalization without external data. npj Digital Medicine. 7:124.
- Obermeyer Z, Powers B, Vogeli C, Mullainathan S. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 366(6464):447–453.
- US Food and Drug Administration. 2025. Artificial Intelligence Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. Draft Guidance for Industry and Food and Drug Administration Staff. US Food and Drug Administration.
- World Health Organization. 2024. Ethics and governance of artificial intelligence for health: Guidance on large multi modal models. Geneva: World Health Organization.