Independent learning for medical-device professionals
CommentaryConsulting
LearningMTL-120 · CORE MEDICAL DEVICE TOPIC

Clinical and Performance Evaluation

How to connect intended purpose and product claims to scientifically credible evidence—and maintain that evidence as the device, state of the art and real-world experience develop.

What you will learn

By the end of this topic, you should be able to define the clinical questions created by a device’s intended purpose and claims; distinguish clinical evaluation for medical devices from performance evaluation for IVDs; plan and appraise evidence; identify gaps; decide when new investigations or studies are needed; connect endpoints and acceptance criteria to benefit-risk conclusions; and maintain the evaluation throughout the product lifecycle.

01

Clinical evaluation is a lifecycle evidence process

Clinical and performance evaluation asks whether the available evidence supports the safety, performance and benefit claims for the intended population, users and conditions of use. It is not simply a literature review and it does not begin when engineering finishes.

Clinical questions influence requirements, risk controls, verification, usability work, labelling and post-market plans. Engineering evidence also supports the evaluation by showing that the configuration used to generate clinical data represents the product being placed on the market.

The central principle

Start with the claim and the decision it requires. Collecting data without a defined evidence question produces volume, not assurance.

Keep the scope aligned with MTL-102 — Intended Purpose, Users and Use Environments.

02

Translate intended purpose into evidence questions

Every important claim should create a question that can be answered with objective evidence. Broad claims usually require broader populations, comparators, endpoints and follow-up than narrowly defined claims.

PopulationPatients, specimens, conditions, disease stage and relevant subgroups
Intervention or index testDevice, procedure, configuration, operator and use conditions
ComparatorCurrent practice, reference method, control, baseline or justified alternative
OutcomeClinical benefit, safety, diagnostic performance or patient-relevant effect
TimeFollow-up sufficient to observe benefit, harm, recurrence or durability
DecisionAcceptance criterion and conclusion needed for the intended claim
  • Separate intended purpose from promotional statements and secondary claims.
  • Identify which claims are clinical, analytical, technical, usability-related or economic.
  • Define claimed benefits in terms meaningful to patients or clinical practice.
  • Identify safety outcomes, foreseeable adverse events and residual risks.
  • Specify relevant subgroups, contraindications and limitations.
  • Trace each question to the evidence source and final conclusion.

Convert the resulting needs into controlled requirements through MTL-103 — User Needs and Design Inputs.

03

Medical devices and IVDs use related but different models

Medical-device clinical evaluation

Assesses clinical data concerning the device to verify clinical safety and performance and support the claimed clinical benefits and benefit-risk conclusion.

IVD scientific validity

Establishes the association between the analyte or marker and a clinical condition or physiological state.

IVD analytical performance

Establishes the test’s technical ability to detect or measure the analyte, including characteristics such as accuracy, precision and analytical sensitivity.

IVD clinical performance

Establishes the ability to yield results related to a particular clinical condition or physiological or pathological process in the intended population.

An IVD performance evaluation integrates scientific validity, analytical performance and clinical performance. These strands answer different questions and should not be collapsed into one generic “validation” statement.

04

Plan the evaluation before choosing evidence

The plan should define the questions, methods and decision rules before the team knows which answer is easiest to obtain. It should remain proportionate to the device, novelty, claims, risks and maturity of available knowledge.

  • Define device scope, variants, accessories, software versions and intended markets.
  • List clinical-safety, performance and benefit claims to be evaluated.
  • Describe the state of the art, alternatives and relevant clinical practice.
  • Define search methods, data sources, inclusion criteria and appraisal rules.
  • Identify endpoints, follow-up, subgroups and acceptance criteria.
  • Explain how equivalence or similar-device data will be used, if at all.
  • Define gaps and the criteria for deciding whether new data are necessary.
  • Plan PMCF or PMPF and periodic updating.
  • Assign qualified authors, reviewers and medical or scientific expertise.

The evaluation plan should be part of the connected evidence structure described in MTL-104 — Design Controls and Technical Documentation.

05

Use multiple evidence sources deliberately

Manufacturer-generated data

Clinical investigations, performance studies, bench and analytical studies, usability validation, complaints, vigilance, PMCF, PMPF and registries.

Scientific literature

Peer-reviewed studies on the device, equivalent or similar technologies, clinical condition, endpoints, alternatives and state of the art.

External real-world data

Registries, health records, claims databases, surveillance systems and independent post-market studies.

Technical and regulatory evidence

Standards, competent-authority information, safety notices, recalled-device experience and product-specific guidance.

Each source has strengths, limitations and possible bias. A database search may miss unpublished adverse information; a company study may be well controlled but narrow; real-world data may be broad but confounded; bench data may explain mechanism without proving clinical benefit.

06

Appraise relevance and quality separately

A rigorous study may still be irrelevant to the claimed device, population or outcome. A highly relevant observation may be too weak to support a decisive conclusion on its own.

IdentityIs the device or method sufficiently characterised?
PopulationDo participants or specimens represent the intended population?
MethodDoes the design control important bias and confounding?
OutcomeAre endpoints valid, measured consistently and clinically meaningful?
CompletenessAre missing data, withdrawals, deviations and adverse events explained?
ApplicabilityCan the result support this claim, configuration, user and environment?

Record both favourable and unfavourable evidence. Excluding inconvenient data without a predefined, scientifically defensible reason undermines the evaluation.

07

Define the state of the art and clinical context

The state of the art describes current knowledge, accepted clinical practice, available alternatives, relevant outcomes and known risks. It provides the context for judging whether the device’s benefits, performance and residual risks are acceptable.

  • Describe the condition, pathway, users and unmet need.
  • Identify relevant guidelines, consensus, technologies and treatment alternatives.
  • Characterise expected outcomes and complications under current practice.
  • Explain whether endpoints and comparators remain clinically relevant.
  • Identify changes in practice, epidemiology, technology or diagnostic criteria.
  • Use the same context when defining benefits and evaluating residual risks.

The comparison is not necessarily a claim that the device is superior. The purpose is to understand what safe and effective performance means in the current clinical setting.

08

Treat equivalence as a demanding evidence argument

Data from another device may contribute only when differences are understood and do not invalidate transfer of the clinical conclusions. Similarity can help identify hazards or design questions without establishing equivalence.

Technical characteristics

Design, specifications, materials, energy, software, principles of operation and critical performance.

Biological characteristics

Patient-contact materials, substances released, tissue interaction and biological response.

Clinical characteristics

Clinical condition, purpose, population, user, site, performance and relevant outcomes.

Data access

Sufficient information to assess the characteristics, methods, results, adverse events and limitations credibly.

A predicate, competitor or previous-generation product is not automatically equivalent. Document why every relevant difference does or does not affect safety, clinical performance or benefit-risk.

09

Use gap analysis to decide what evidence must be generated

A gap exists when the available evidence cannot answer an important question with sufficient relevance and quality. The response should address the specific uncertainty rather than default to the largest possible study.

  • Map each claim, GSPR, risk and residual uncertainty to available evidence.
  • Separate absence of evidence from evidence of unacceptable performance.
  • Identify whether the gap is technical, analytical, usability, clinical or post-market.
  • Consider whether bench, modelling, usability, literature or existing clinical data can answer the question.
  • Explain why a new clinical investigation or performance study is necessary when people or specimens are involved.
  • Prioritise gaps that affect patient protection, benefit-risk or major claims.
  • Agree regulatory strategy early where evidence expectations are uncertain.

Use MTL-105 — Medical-device Risk Management to connect residual uncertainty to risk decisions rather than treating every gap as equally important.

10

Clinical investigations must be ethical, necessary and scientifically sound

A clinical investigation involving human subjects should answer a question that cannot be resolved adequately by less burdensome evidence. Subject rights, safety and well-being take priority over commercial or project objectives.

  • Define the objective, hypothesis or estimand and the decision the result will support.
  • Select a population, comparator, endpoints and follow-up appropriate to the intended claim.
  • Identify device- and procedure-related risks and establish monitoring and stopping rules.
  • Obtain applicable regulatory and ethics approvals before starting.
  • Use an investigational configuration that is sufficiently controlled and representative.
  • Protect informed consent, privacy, data integrity and vulnerable populations.
  • Control investigators, sites, training, deviations, adverse events and device accountability.
  • Preserve the protocol, statistical plan, dataset, analysis, report and disclosure obligations.

ISO 14155:2026 provides the current international good-clinical-practice framework for medical-device investigations. National and regional requirements still determine approvals and legal obligations.

11

IVD performance studies need specimen and clinical-context control

IVD evidence depends on the relationship between specimen, intended user, method, clinical condition, reference standard and statistical analysis. Sample numbers alone do not make a dataset representative.

  • Define intended population, specimen type, collection, handling, storage and inclusion criteria.
  • Use an appropriate reference method or clinical truth definition.
  • Represent relevant disease stages, prevalence, interferents and confounding conditions.
  • Control reader, site, instrument, reagent-lot and operator effects where applicable.
  • Predefine invalid, indeterminate, missing and discordant-result handling.
  • Separate analytical-performance questions from clinical-performance questions.
  • Protect subjects and specimen donors and obtain the required approvals and consent.
  • Ensure the tested system represents instruments, software, reagents, calibrators and interpretation rules to be marketed.

ISO 20916:2019 remains current for good study practice in IVD clinical performance studies using specimens from human subjects.

12

Endpoints and analysis must support the intended conclusion

Endpoints should be clinically meaningful, measured consistently and linked to the claim. Surrogate, composite or technical endpoints require justification when they stand in for patient-relevant benefit.

Performance endpoint

Measures whether the device delivers the claimed function, diagnostic result or clinical effect.

Safety endpoint

Captures adverse events, complications, device deficiencies and procedure-related harm.

Patient-centred endpoint

Reflects symptoms, function, quality of life, burden, recovery or another outcome meaningful to patients.

Operational endpoint

Examines usability, workflow, training, interpretation or successful completion where it affects clinical use.

Predefine sample-size assumptions, analysis populations, multiplicity, missing-data handling, sensitivity analyses and subgroup treatment. A statistically significant result may be clinically unimportant, while an underpowered inconclusive result does not demonstrate equivalence or absence of harm.

13

Identify the device configuration behind every result

Clinical evidence cannot support a product that is materially different from the device investigated unless the impact of the differences is assessed. Hardware, software, algorithms, materials, accessories, workflow and labelling can all affect transferability.

  • Record model, variant, serial or lot, software and algorithm versions.
  • Identify accessories, consumables, reagents, calibrators and connected systems.
  • Control training, instructions, procedure and clinical workflow.
  • Document changes during a study and determine whether approval or protocol updates are required.
  • Assess whether later design changes affect the mechanism, risk, endpoint or population.
  • Use bridging evidence where appropriate and generate new data where uncertainty remains material.

Connect configuration and result traceability to MTL-106 — Verification and Validation.

14

The evaluation report must show the reasoning

A credible report lets an independent reviewer follow the questions, methods, evidence, limitations and conclusions. It should not hide the difficult parts behind a list of references.

ScopeDevice, purpose, variants, claims, populations and markets
MethodPlan, searches, selection, appraisal and analysis rules
EvidenceFavourable, unfavourable, internal and external data
LimitationsBias, uncertainty, gaps, applicability and unresolved questions
ConclusionSafety, performance, benefit-risk and claim support
MaintenancePMCF or PMPF, PMS inputs, update frequency and triggers

The report should trace to the intended purpose, risk-management file, technical documentation, labelling and post-market plan. Use MTL-307 — EU MDR and IVDR General Safety and Performance Requirements for the wider conformity-evidence structure.

15

Post-market evidence maintains—not merely confirms—the conclusion

Post-market clinical follow-up for medical devices and post-market performance follow-up for IVDs should address residual questions, verify continued performance and identify emerging risks or changes in the state of the art.

  • Define specific objectives rather than “collect more data.”
  • Use complaints, vigilance, surveys, registries, literature, studies and real-world data appropriately.
  • Track exposure or denominator information so event rates can be interpreted.
  • Look for subgroup, use-environment and long-term outcomes not resolved pre-market.
  • Compare observed performance with acceptance criteria and benefit-risk assumptions.
  • Identify new alternatives, clinical practices, standards and scientific findings.
  • Feed conclusions into risk, labelling, design, CAPA and regulatory reporting.
  • Update the evaluation after significant changes or new evidence, not only on a calendar.
16

Clinical and performance evaluation across the lifecycle

1

Define the clinical claim

Clarify intended purpose, target population, users, indications, contraindications, clinical benefits, performance and foreseeable risks.

Typical evidence: Intended-purpose statement, claims matrix, state-of-the-art review and clinical evidence questions.
2

Plan the evaluation

Define methods, data sources, appraisal criteria, endpoints, acceptance criteria, responsibilities and update triggers.

Typical evidence: Clinical or performance evaluation plan, literature protocol, data-management plan and gap-analysis method.
3

Assess existing evidence

Identify relevant internal and external data, judge quality and relevance, and explain how each source contributes to the claims.

Typical evidence: Search records, appraisal tables, equivalence analysis, complaint data, prior studies and evidence map.
4

Close important gaps

Generate proportionate bench, usability, analytical, clinical or post-market evidence when existing data cannot answer the question.

Typical evidence: Investigation or study protocol, approvals, monitoring records, datasets, analyses, deviations and report.
5

Reach a traceable conclusion

Integrate favourable and unfavourable data to conclude whether safety, performance, benefits and residual risks are adequately supported.

Typical evidence: Clinical evaluation report or performance evaluation report, benefit-risk conclusion and claim disposition.
6

Maintain the evidence

Use post-market surveillance, vigilance, PMCF or PMPF and new scientific information to challenge assumptions and update conclusions.

Typical evidence: Update schedule, PMS outputs, PMCF or PMPF reports, trend reviews, change assessments and revised evaluation.
17

Common misconceptions

“Clinical evaluation is a literature review.”

Literature is one source. The evaluation integrates all relevant clinical, technical, risk and post-market evidence.

“Every device needs a new clinical investigation.”

The need depends on claims, novelty, risk, gaps and available evidence; the decision must be scientifically and regulatorily justified.

“A predicate device is automatically equivalent.”

Regulatory predicate status does not by itself establish technical, biological and clinical equivalence for transferring evidence.

“Analytical performance proves an IVD’s clinical performance.”

Analytical and clinical performance answer different questions and both must be connected to scientific validity.

“A successful study proves every claim.”

Conclusions are limited by the studied population, device configuration, endpoints, comparator, follow-up and analysis.

“The report is finished at market release.”

New field data, scientific knowledge, risks and product changes can require the evaluation and its conclusions to be updated.

18

Practical review checklist

  • Are intended purpose, population, users, claims and benefits clearly defined?
  • Does every important claim have a specific evidence question and acceptance basis?
  • For an IVD, are scientific validity, analytical performance and clinical performance addressed separately?
  • Does the evaluation plan define methods before evidence selection?
  • Are favourable and unfavourable data identified and appraised transparently?
  • Is the state of the art current and relevant to the claimed clinical context?
  • Is any equivalence argument supported across technical, biological and clinical characteristics?
  • Are gaps connected to risk and resolved with proportionate evidence?
  • Are investigations or performance studies ethical, approved and scientifically sound?
  • Do endpoints, follow-up and statistical methods support the intended conclusions?
  • Is the evaluated product configuration traceable to the marketed device?
  • Does the report state limitations, uncertainties and unsupported claims?
  • Are PMCF, PMPF and PMS activities tied to specific residual questions?
  • Do new evidence and product changes trigger controlled updates?
19

Authoritative references

Clinical-evidence requirements depend on device type, classification, claims, novelty, risk and jurisdiction. Confirm current legislation, standards, guidance, ethics requirements and competent-authority expectations for the intended markets and study locations.

KEY TAKEAWAY

Begin with the clinical claim, then build only the evidence needed to support it credibly

A strong evaluation connects intended purpose, state of the art, risk, product configuration, data quality, limitations and post-market learning in one transparent and continuously maintained argument.