What you will learn
Describe potential AI functions, distinguish prediction from action, evaluate a proposed benefit and identify the evidence needed to support a bounded product claim.
Start with the clinical job
An AI opportunity starts with a problem in care: an unreliable measurement, a delayed review, a difficult decision or a treatment that needs to respond to changing conditions. Define the patient, user, environment and action before choosing a model. A technically impressive prediction has little value if nobody can act on it in time.
AI may run in an instrument, an implant, a connected platform or standalone software. Machine learning identifies patterns from examples; generative AI produces content. Conventional algorithms can also provide sophisticated automation. A closed-loop controller, alarm or robot is not automatically an AI system. Compare candidate approaches using the same task and acceptance criteria.
The examples here are application patterns, including existing uses and emerging possibilities. They do not establish authorisation, suitability or clinical benefit for any particular product. For marketed examples, examine the intended use and limitations in the public decision records linked from the FDA AI-enabled medical-device list. The list itself is not comprehensive.
Begin with MTL-130 — Is My Health Software a Medical Device? and MTL-131 — AI and Machine Learning Foundations for Medical Devices.
Decide what the output is allowed to do
These responsibilities can coexist in one product. Increasing autonomy changes the assurance problem, but an advisory label does not make a function low risk. A missed triage flag may delay care; a misleading measurement may drive an incorrect treatment even when a clinician makes the final decision. Specify how the output influences the actual workflow.
Define three states explicitly: a valid result, an uncertain or unsupported result, and no result. Users must be able to distinguish them. A system that cannot process an image must not silently report that no abnormality was found.
Acquire, reconstruct and measure
Acquisition guidance can analyse a live ultrasound image and suggest probe movement. Its input includes image frames and acquisition context; its output is positioning guidance or an assessment of image adequacy. The benefit is more consistent acquisition. Evidence must address whether intended users obtain clinically adequate views, including difficult anatomy and unfamiliar equipment. A confident instruction in the wrong direction is a foreseeable failure.
Reconstruction and enhancement can use measured signals to produce a clearer image or support shorter acquisitions. The task is to preserve information needed for interpretation. Visual smoothness alone is insufficient: processing could suppress a small lesion or introduce a misleading feature. Evaluate the relevant clinical task using representative acquisition conditions and an appropriate reference.
Quantification includes organ segmentation, tumour volume, cardiac measurements and cell counting. The user may rely on a trend rather than one value. Repeatability, boundary errors, units, calibration and variation between equipment matter. A software update that changes segmentation behaviour can create an apparent disease change without any biological change. Store the method and version alongside the measurement.
Detect, interpret and prioritise
Detection systems can highlight suspicious image regions, classify ECG segments or identify unusual cell morphology. Inputs must carry the right patient and specimen identity and sufficient quality. Outputs might be locations, scores or classes. A highlighted region is not necessarily a diagnosis; the intended claim determines what evidence and workflow are needed.
For IVD systems, potential applications include digital pathology, cell classification and interpretation of complex assay patterns. Distinguish an analytical quality flag from a clinical interpretation. Sample preparation, reagent lots, instruments, interfering substances and rare patterns can affect performance. A model cannot repair an unidentified sample or make an invalid assay valid.
Triage changes review priority. A chest X-ray assistant might alert staff to a suspected urgent finding while all images remain subject to normal reporting. Evaluate missed findings, unnecessary alerts and the effect on unflagged patients. Sensitivity alone does not show that queues improve: excessive alerts can overwhelm the intended benefit.
Forecasting differs from detecting a current condition. A wearable model might predict physiological deterioration using recent measurements. Define the prediction horizon, available intervention and acceptable alarm burden. Validate performance using information actually available at prediction time; future treatments or outcomes leaking into inputs can make retrospective results misleading. NIH’s AI overview describes examples in imaging, monitoring and prediction.
Plan treatment and assist action
Treatment planning can use images and clinical information to propose contours, estimate response or compare options. The output should expose relevant assumptions and permit review. A plausible plan is not evidence of better outcomes. Compare it with appropriate practice and consider whether the user can recognise a harmful recommendation.
Procedural assistance may identify anatomical landmarks, track instruments or support navigation. In a moving scene, timing and registration errors can matter as much as classification accuracy. Smoke, blood, occlusion, camera movement or unfamiliar anatomy may invalidate an overlay. Define when guidance disappears, how the user is notified and how the procedure continues safely.
Adaptive therapy is a potential application in drug delivery, stimulation and ventilation. A learned prediction might inform a controller without having unrestricted authority over it. Specify permitted adjustments, independent limits, stale-data handling, sensor faults and recovery. Test the complete sensing–decision–actuation system. Automatic retraining or unrestricted live adaptation must never be assumed merely because the device uses AI.
Rehabilitation and assistive devices can interpret movement, muscle signals or user intent to adapt exercises or assistance. An incorrect prediction can cause a physical movement, so evaluate everyday transitions, fatigue, unusual postures and loss of signal. Personalisation requires evidence that adaptation remains within safe and useful bounds.
Support home use, reporting and reliability
Home-use applications may assess inhaler technique, recognise movement patterns or help a patient obtain a usable recording. Homes introduce variable lighting, noise, connectivity, handling and user capability. Feedback must be understandable and actionable; the system needs a route for users who repeatedly cannot complete the task.
Generative AI can draft a structured report or explanation from device outputs. Preserve the underlying findings and make unsupported additions detectable. A fluent sentence can invent a measurement, reverse a negation or attach information to the wrong patient. Evaluate omissions, contradictions and clinically consequential errors, not just readability. See MTL-138 — Generative AI and Large Language Models in MedTech.
AI can also analyse device telemetry for calibration drift or developing faults. A maintenance prediction should identify what action is justified and how a missed warning is handled. Distinguish a product function relied on for safe use from a manufacturer’s service analytics tool. The latter belongs alongside the development and quality applications in MTL-503 — AI in Medical-device Development and Quality Management.
Worked decision: a wearable early-warning feature
Fictional proposal: a home respiratory monitor collects breathing rate, motion and signal-quality information. The team proposes an early-warning score to help a defined clinical service identify patients who need review. The initial claim is support for clinician review, with no diagnosis and no automatic treatment adjustment.
The team first measures the current review process and defines a useful warning horizon. It evaluates whether earlier review leads to an actionable decision, then specifies acceptable missed-event performance and workload. It separates training and evaluation by patient and time as appropriate, tests realistic missing-data patterns and reports relevant subgroup results.
Suppose a retrospective demonstration performs well only after low-quality recordings are removed. That does not justify home deployment. The team must determine how often the product produces no result, which patients are affected and whether users understand the fallback. It should also evaluate the complete service: an accurate alert that nobody receives or responds to cannot deliver the intended benefit.
The decision record retains the claim, workflow, data exclusions, reference outcome, performance by subgroup, alarm burden, usability findings, remaining limitations and monitoring plan. This example intentionally supplies no clinical thresholds: those require product-specific justification.
Turn an opportunity into an evidence plan
- State the clinical problem, current comparator and measurable benefit. Ask whether simpler automation would achieve it.
- Define the supported patient population, users, environments and equipment, including exclusions.
- Describe inputs, outputs, uncertainty states and the action taken because of the result.
- Identify data rights, provenance, representativeness and independent evaluation arrangements.
- Analyse incorrect, late, absent and misleading outputs, including foreseeable user reliance.
- Specify performance, workflow and usability acceptance criteria before the final evaluation.
- Assign ownership for release, updates, monitoring, incident investigation and fallback.
Link the evidence plan to MTL-132 — Data Governance for Medical AI, MTL-133 — Bias, Fairness and Representativeness in Medical AI, MTL-134 — AI Risk Management and Human Oversight and MTL-135 — Verification, Validation and Clinical Evidence for Medical AI. Changes and field behaviour connect to MTL-136 — Predetermined Change Control Plans for Medical AI and MTL-137 — Post-market Monitoring, Drift and Model Performance. Protect the full input and model pipeline using MTL-139 — AI Security, Robustness and Adversarial Resilience.
Practice: distinguish three different claims
A fictional imaging team proposes one model for three products: A highlights a suspected abnormality; B moves the image to the top of a worklist; C automatically changes a treatment setting. What changes in the intended purpose, evidence and fallback? Why is one model-accuracy figure insufficient?
Explore a model answer
A requires evidence that the finding and presentation support the intended interpretation task, including missed and distracting highlights. B requires evidence about time to review, alert workload and effects on unflagged patients. C requires evidence for the complete treatment-control system, including limits, timing, sensors, actuator behaviour and hazardous transitions.
All three need representative data, suitable reference methods and assessment of relevant subgroups. Their output actions and consequences differ, so the same test set and headline metric cannot establish suitability for all three. Fallback might mean ordinary image review for A and B, while C requires a defined safe control strategy appropriate to the therapy. A human in the workflow counts as a control only if that person can detect and manage the relevant failure in time.
Continue learning
Follow MTL-408 — Medical AI — From Claim to Controlled Use to see intended purpose, data, risk, validation, change control and monitoring connected through one fictional product. Use MTL-503 — AI in Medical-device Development and Quality Management to explore the AI tools used by the teams developing and maintaining devices.
Source review: 21 September 2026. The FDA list and NIH overview linked above provide external starting points; the worked scenarios and decision exercises are fictional educational examples. Check current product-specific evidence and market requirements before applying a concept.