Independent learning for medical-device professionals
SearchCommentaryConsulting
LearningMTL-133 · AI-ENABLED MEDICAL DEVICES

Bias, Fairness and Representativeness in Medical AI

Replace reassuring averages with clinically justified subgroup evidence, transparent uncertainty and proportionate controls.

What you will learn

By the end of this topic, you should be able to identify important sources of bias, distinguish representativeness from simple demographic balance, select clinically meaningful subgroup analyses, interpret uncertainty and implement proportionate controls when performance differs across populations or settings.

01

Fairness is a clinical risk question

The aim is not identical statistics for every group. It is to understand whether performance is sufficiently safe and effective for the people and settings covered by the intended purpose.

AI lifecycle principle

Define the claim, control the complete system, generate independent evidence and monitor the product in real use.

02

Core concepts

Selection bias

The included patients, sites or records differ systematically from the intended-use population.

Measurement bias

Sensors, workflows or documentation create different-quality inputs across people or settings.

Label bias

The reference outcome reflects inconsistent practice, access to care, observer judgement or historical inequity.

Representation

The data cover clinically relevant variation in population, prevalence, equipment, site, workflow and disease presentation.

Performance disparity

Error rates, calibration or clinical consequences differ between meaningful subgroups or environments.

Deployment bias

A model is used by different people, for a different decision or under different conditions than those evaluated.

03

A practical lifecycle

Bias assessment begins when the problem and data are chosen and continues after deployment.

1

Map affected groups

Identify biological, demographic, technical, clinical and environmental factors that could influence inputs, outputs or harm.

Typical evidence: Equity and subgroup analysis plan.
2

Examine the data-generating process

Assess who enters the data, who is absent, how labels arise and where care pathways differ.

Typical evidence: Bias-source analysis and data-flow review.
3

Set subgroup questions

Predefine clinically justified groups, metrics, minimum sample support and uncertainty reporting.

Typical evidence: Statistical analysis plan.
4

Evaluate intersections

Investigate combinations of factors where risk may be hidden by broad single-variable groups.

Typical evidence: Stratified results and confidence intervals.
5

Control the risk

Improve data, limit claims, tune thresholds, change workflow, add warnings or exclude unsupported use.

Typical evidence: Risk-control decisions and residual-risk rationale.
6

Monitor real use

Track coverage, data quality, errors and outcomes across relevant populations and sites.

Typical evidence: Post-market subgroup dashboard and signal review.
04

Controls to build in

Choose subgroup analysis from clinical and risk reasoning, not from whichever variables happen to be convenient.

  • Define representativeness against the intended-use population and environment.
  • Report denominators, uncertainty and data quality alongside subgroup performance.
  • Check calibration and predictive values where prevalence affects the clinical interpretation.
  • Investigate missing demographic data and the limitations of proxy variables.
  • Evaluate whether user workflow amplifies or mitigates model disparities.
  • Turn unsupported populations or conditions into explicit limitations and monitoring priorities.
05

Evidence to retain

Bias analysis

Plausible bias sources across collection, labelling, modelling, evaluation and deployment.

Representativeness report

Comparison between datasets and intended-use populations, sites, equipment and workflows.

Subgroup results

Predefined metrics with denominators, confidence intervals, calibration and failure review.

Control rationale

Data, model, workflow, labelling and monitoring actions tied to identified risk.

06

Common pitfalls

Aggregate success

Strong overall performance can conceal unacceptable error in a smaller population.

Demographics only

Site, device, prevalence, language, workflow and disease spectrum may be equally important.

Tiny subgroup certainty

A point estimate from few cases can look reassuring or alarming without being reliable.

Fairness as a final test

Late evaluation cannot repair a target, label or sampling strategy that encoded the problem.

07

Action checklist

  1. Define who may benefit, who may be harmed and who is missing from the data.
  2. Identify clinical, technical and workflow factors that may change performance.
  3. Predefine subgroup metrics, sample support and uncertainty reporting.
  4. Investigate disparities and their likely root causes.
  5. Apply controls through data, claims, design, workflow and user information.
  6. Continue subgroup monitoring after release.
IN DEPTH

Investigate who experiences the errors

Representation is more than a demographic count

A dataset can include several demographic groups while missing important differences in disease severity, comorbidity, acquisition equipment or care setting. Start with the intended population and plausible failure mechanisms. For a skin-image product, image quality and skin appearance may matter; for a deterioration predictor, observation frequency and treatment practices may matter. Identify clinically meaningful subgroups before evaluating the final model. Consider intersections where there is a credible risk, but recognise that dividing data repeatedly produces small cells and unstable estimates. State what the evidence cannot establish.

Different fairness measures answer different questions

Equal sensitivity asks whether cases are missed at similar rates. Equal positive predictive value asks whether an alert has similar reliability across groups. Calibration asks whether predicted risks mean the same thing. These properties need not coincide when underlying outcome rates differ. Select measures according to the device’s clinical role and harms, explain the trade-offs and assess absolute performance as well as differences. Two groups can have equally poor outcomes; parity alone does not demonstrate acceptable performance. Conversely, a numerical difference is a signal to investigate, not an automatic explanation of its cause.

Investigate before choosing a mitigation

A performance gap may arise from inadequate representation, poorer input quality, inconsistent labels, a different clinical spectrum or the model itself. Audit these possibilities before deciding that retraining will fix it. Potential actions include better collection, improved acquisition instructions, revised preprocessing, model changes or a narrower intended population with a justified access impact. Each action can introduce new harms. Group-specific thresholds require clinical, ethical and regulatory assessment, including how group membership is determined and what happens when it is missing. Re-evaluate the complete product after a change.

WORKED DECISION

A headline sensitivity conceals a weak subgroup

Teaching scenario

In a fictional evaluation, the model detects 180 of 200 positive cases in group A and 24 of 40 in group B. The report leads with pooled sensitivity of 204/240 = 85%. All numbers are illustrative, and the groups represent a prespecified clinically relevant distinction.

1. Show the denominators

Group A sensitivity is 90%; group B sensitivity is 60%. The second estimate is based on only 40 positive cases and needs an uncertainty interval. Report both alongside the pooled result. A small group must not disappear from reporting because its result is inconvenient, and a non-significant comparison would not establish equivalence.

2. Test competing explanations

Review severity, acquisition conditions, missingness and reference labels. Suppose group B was mostly imaged with a different camera in poor light. That observation suggests a mechanism but does not prove the demographic distinction is irrelevant. Design additional evaluation that can separate camera effects, setting and patient characteristics.

3. Make the limitation actionable

The team holds the broad release decision, improves acquisition controls and collects further independent evidence. It assesses whether restrictions would exclude intended patients from benefit or be impractical to enforce. The new evaluation uses prespecified criteria for overall and subgroup performance, rather than accepting the change solely because the pooled score rises.

EVIDENCE IN PRACTICE

Example subgroup analysis record

This abbreviated teaching example shows the reasoning to capture. Adapt it to the product, risk and quality-system procedures, and link to the underlying evidence.

Rationale
Subgroup selected before evaluation because a plausible acquisition or clinical failure mechanism could alter missed-case risk.
Results
Patient and positive-case counts, sensitivity, specificity, uncertainty intervals, missing group data and relevant intersections.
Investigation
Competing explanations, analyses performed, remaining confounders and limitations of causal interpretation.
Decision
Mitigation owner, independent reassessment plan and impact on access to the intended clinical benefit.
PUT IT INTO PRACTICE

Make the decision yourself

After retraining, pooled sensitivity increases from 85% to 88%, but group B falls from 60% to 55%. Can the team call the update an improvement?

Write down your decision, the missing evidence and the next action before opening the answer.

Read the model answer

It can describe the pooled metric as higher, but should not treat that as an adequate product-level benefit conclusion. Examine counts and uncertainty, compare with prespecified subgroup criteria, investigate the mechanism and reassess harm to group B. Consider the severity and distribution of missed cases, not just their total. The release decision must explain the trade-off and any unmet criteria. Selecting a more favourable aggregate after seeing the results would hide the problem rather than resolve it.

Apply this to your project

Use the example record above to document one real decision. Identify the assumption most likely to change the conclusion, the evidence needed to test it and the person responsible for the next step.

Read alongside this lesson: NIST AI Risk Management Framework — consider harmful bias and the context of use

REFERENCES

Authoritative starting points

This module provides educational guidance, not a product-specific regulatory determination. Confirm the legislation, guidance and submission expectations applicable to each intended market.

KEY TAKEAWAY

Make performance visible where risk differs

Responsible AI replaces reassuring averages with clinically justified subgroup evidence, transparent uncertainty and controls proportionate to the consequences.

Continue through the MedTechLearning AI-enabled medical-device pathway to connect this topic with the wider lifecycle.