What you will learn
By the end of this topic, you should be able to design an AI-specific post-market monitoring plan, distinguish data, concept and performance drift, define leading and lagging indicators, detect clinically relevant changes, investigate signals and connect findings to risk management, CAPA, retraining and regulatory change control.
A locked model can still drift in practice
The parameters may be unchanged while patients, prevalence, equipment, workflows, data quality and user behaviour move away from the conditions represented in validation.
Define the claim, control the complete system, generate independent evidence and monitor the product in real use.
Core concepts
Data drift
The distribution or quality of model inputs changes relative to the validated baseline.
Concept drift
The relationship between inputs and the clinical target changes over time.
Performance drift
Observed discrimination, calibration, error or clinical effectiveness changes in real use.
Proxy monitoring
Input and output indicators used when verified outcomes arrive slowly or incompletely.
Signal
A pattern that may indicate new risk, degraded performance or an ineffective control and warrants evaluation.
Trigger
A predefined condition for review, investigation, containment, correction or formal reporting.
A practical lifecycle
Plan monitoring before release so required data, baselines, permissions and response routes actually exist.
Define monitored claims and risks
Select safety, performance, equity, usability and operational questions linked to intended use and residual risk.
Typical evidence: AI post-market monitoring plan.Establish the baseline
Preserve validation distributions, metrics, thresholds, subgroups, sites and known limitations for comparison.
Typical evidence: Reference baseline and data dictionary.Collect reliable signals
Combine complaints, incidents, overrides, data quality, output distributions, outcomes and service information.
Typical evidence: Signal sources, quality controls and dashboards.Detect meaningful change
Use statistical and clinical thresholds with sufficient denominators, delay awareness and subgroup visibility.
Typical evidence: Trend analysis and trigger records.Investigate and contain
Confirm validity, identify affected configurations and populations, assess risk and take proportionate interim action.
Typical evidence: Investigation, health-hazard and containment records.Correct and learn
Feed findings into CAPA, risk, clinical evaluation, labelling, retraining, PCCP and regulatory reporting.
Typical evidence: CAPA, change record and effectiveness review.Controls to build in
Monitoring must lead to decisions; collecting telemetry without ownership and triggers is not surveillance.
- Assign owners for each metric, review cadence, investigation and escalation route.
- Preserve privacy while collecting enough context to interpret performance and subgroup signals.
- Track data-quality failures and abstentions, not only successful model outputs.
- Account for delayed, missing and biased ground truth in real-world outcome estimates.
- Separate model issues from upstream sensors, workflow, integration and user behaviour.
- Maintain fleet and version visibility so affected users and configurations can be identified.
Evidence to retain
Monitoring plan
Questions, indicators, sources, baselines, thresholds, owners, cadence and actions.
Performance dashboard
Version-aware trends for data quality, outputs, outcomes, failures and relevant subgroups.
Signal assessment
Validity, scope, clinical significance, risk, reportability and required containment.
Periodic review
Integrated conclusion across complaints, vigilance, clinical follow-up, security and model monitoring.
Common pitfalls
Accuracy without ground truth
A dashboard cannot claim real-world accuracy if reliable reference outcomes are unavailable.
Average stability
Stable fleet-wide results can conceal deterioration at one site, device type or patient subgroup.
Alert without action
Thresholds need owners, investigation methods, containment options and decision authority.
Retrain first
Changing the model before understanding the root cause can hide data or workflow failures.
Action checklist
- Link monitoring questions to claims, risks and known limitations.
- Preserve validated baselines, configurations and subgroup definitions.
- Define signal sources, data quality, thresholds, owners and cadence.
- Monitor inputs, outputs, failures, human responses and clinical outcomes.
- Investigate signals across model, system, workflow and population causes.
- Connect findings to vigilance, CAPA, risk, evidence and controlled change.
Turn monitoring signals into accountable action
Distinguish a shift from a loss of performance
Data drift means that the observed input distribution changes; concept drift concerns a changed relationship between inputs and the target. A new scanner can alter images without changing disease biology, while new treatment practice can change outcomes for patients with similar recorded features. A distribution shift is a reason to investigate, not proof that performance has fallen. Conversely, stable input summaries do not prove safety. Monitor product failures, workflow and outcomes as well as statistical summaries, and choose measures that relate to plausible failure mechanisms rather than whatever is easiest to plot.
Account for delayed and selective outcomes
The reference outcome may become available weeks later or only for patients selected for further investigation. If only flagged patients receive the reference test, monitoring may miss false negatives among unflagged patients. Define how outcome data will be obtained, linked and assessed for missingness and selection bias. Report coverage and delay alongside performance. Proxies such as score distributions, overrides and input quality provide earlier signals, but none is automatically a substitute for clinical performance. A dashboard should identify which conclusions each measure can and cannot support.
Set the response procedure before an alert occurs
Every monitored signal needs an owner, review frequency, justified trigger, investigation route and possible action. Separate an early warning from a release-stopping or containment threshold. Statistical limits depend on baseline variation, volume, repeated testing and risk; there is no universal drift threshold suitable for all products. Specify what happens when monitoring itself fails. Link confirmed problems to complaint handling, risk review, corrective action and applicable vigilance assessment. An engineer’s drift alert and a legally reportable event are different determinations, but the first may supply evidence for the second.
A scanner update changes the inputs overnight
A fictional imaging service sees rejected inputs rise from a stable baseline near 2% to 12% at one hospital after scanner software maintenance. Confirmed clinical labels take four weeks to arrive. Other sites are stable. The percentages are teaching examples, not recommended alarm limits.
1. Triage the signal
The monitoring owner checks volume, logging changes and the timing of maintenance, then compares acquisition metadata and failure reasons. The signal is localised to a changed export format. Low-quality and rejected cases remain visible in the denominator; removing them would make the remaining model metrics appear reassuring while hiding lost service availability.
2. Contain according to risk
The team assesses whether any affected inputs produced plausible but incorrect outputs. It activates the established manual workflow for the affected configuration where warranted, informs the responsible users and preserves logs. It does not wait four weeks for labels before considering containment, nor conclude that all sites must be shut down solely from this local signal.
3. Close the investigation with evidence
Engineering reproduces the format issue; clinical and quality reviewers assess affected cases and potential harm. A controlled correction is verified and evaluated before reactivation. Later outcome data check the suspected impact. The team updates compatibility controls and decides whether complaint, corrective-action or vigilance processes require further action.
Example monitoring signal record
This abbreviated teaching example shows the reasoning to capture. Adapt it to the product, risk and quality-system procedures, and link to the underlying evidence.
- Signal
- Input rejection rate by site and acquisition version, with counts, baseline and monitoring completeness.
- Assessment
- Data shift, possible availability/performance impact, label delay and potential affected patients distinguished.
- Action
- Named owner, risk-based containment decision, user communication and controlled investigation.
- Closure
- Verified correction, reactivation approval, follow-up outcomes and links to quality-system records.
Make the decision yourself
The average model score has not changed for six months. Can you conclude there is no drift or deterioration? Propose a stronger monitoring approach.
Write down your decision, the missing evidence and the next action before opening the answer.
Read the model answer
No. A stable average can conceal changes in the tails, subgroups, missing inputs, population mix or the relationship between predictions and outcomes. Review relevant input features and acquisition versions, rejection and failure rates, score distributions, subgroup performance, clinical outcomes and human responses. Record denominators, label coverage and delay. Use predefined, risk-related investigation triggers. The aim is to detect meaningful changes in the supplied product and its use, rather than to prove safety from one convenient summary statistic.
Apply this to your project
Use the example record above to document one real decision. Identify the assumption most likely to change the conclusion, the evidence needed to test it and the person responsible for the next step.
Read alongside this lesson: NIST AI Risk Management Framework — connect measurement to management and accountable response
Authoritative starting points
- FDA — Good Machine Learning Practice for Medical Device Development: Guiding Principles
- IMDRF — Machine Learning-enabled Medical Devices: Key Terms and Definitions
- NIST — Artificial Intelligence Risk Management Framework
- Regulation (EU) 2024/1689 — Artificial Intelligence Act
This module provides educational guidance, not a product-specific regulatory determination. Confirm the legislation, guidance and submission expectations applicable to each intended market.
Monitor the changing system around the model
Effective AI surveillance combines version-aware technical indicators with clinical outcomes, user behaviour, subgroup visibility and decisive response processes.
Continue through the MedTechLearning AI-enabled medical-device pathway to connect this topic with the wider lifecycle.