Independent learning for medical-device professionals
CommentaryConsulting
LearningMTL-128 · CORE MEDICAL DEVICE TOPIC

Statistical Methods and Measurement Assurance

How to choose proportionate statistical methods, understand variation and ensure that measurements are sufficiently trustworthy to support medical-device decisions.

What you will learn

By the end of this topic, you should be able to frame a quantitative decision, distinguish important statistical and metrological concepts, identify relevant sources of variation, select a proportionate study design and sample size, assess a measurement system, interpret uncertainty and traceability, use capability and reliability evidence appropriately, and report conclusions without overstating what the data demonstrate.

01

Statistics turn observations into decisions

Medical-device development depends on decisions made with incomplete information: whether a requirement is met, whether a risk control is effective, whether a measurement method is adequate, whether a process is stable, whether two designs differ and whether performance remains acceptable over time.

Statistical methods help quantify variation and uncertainty. Measurement assurance asks whether the observations themselves are sufficiently reliable for the intended decision. Neither discipline should be reduced to a calculation performed after testing.

The central principle

Define the decision, claim, population and measurement process before choosing the statistical method or sample size. A sophisticated analysis cannot rescue unsuitable data.

02

Start with the decision and its consequences

A useful plan begins with a plain-language decision statement. “Test ten units” is not a decision. “Demonstrate that dose accuracy meets the specified limits across the released configuration and foreseeable operating range” is.

Question

What exactly must the evidence allow the organisation to conclude?

Population

Which devices, lots, users, environments, sites, time periods or software configurations must be represented?

Errors

What are the safety, quality and business consequences of accepting an inadequate design or rejecting an adequate one?

Effect

What difference, degradation or failure rate would be practically important—not merely statistically detectable?

Evidence

Is the purpose exploration, estimation, comparison, verification, validation, monitoring or prediction?

Action

What action follows a pass, failure, inconclusive result, trend or unexpected observation?

Connect each decision to the approved requirement and protocol principles in MTL-106 — Verification and Validation.

03

Use statistical and measurement terms precisely

AccuracyCloseness of agreement, involving both trueness and precision
TruenessCloseness of the average result to a reference value
PrecisionCloseness of agreement among repeated results
RepeatabilityPrecision under the same defined conditions
ReproducibilityPrecision under changed conditions, such as operator, site or laboratory
UncertaintyQuantified doubt associated with a measurement result

Bias is a systematic difference; variability describes dispersion. Resolution is the smallest displayed or detectable increment, not proof of accuracy. Calibration establishes the relationship between indication and reference values under stated conditions; it does not by itself demonstrate that the complete measurement process is suitable for its use.

04

Model the variation that exists in real use

Apparent consistency can result from testing an unrealistically narrow set of samples. The study should expose the variation relevant to the claim while controlling variation that would obscure the decision.

  • Product units, lots, cavities, tooling, suppliers, component tolerances and ageing.
  • Users, operators, laboratories, sites, shifts, fixtures and setup methods.
  • Temperature, humidity, supply conditions, orientation, vibration and electromagnetic environment.
  • Reference materials, biological samples, matrices, interferents and specimen handling.
  • Software, algorithms, configuration, data processing and rounding.
  • Time-related effects such as drift, warm-up, wear, maintenance and recalibration.
  • Sampling, missing values, excluded observations and repeated measurements on the same unit.

Use MTL-102 — Intended Purpose, Users and Use Environments to define representative conditions and MTL-103 — User Needs and Design Inputs to translate them into measurable limits.

05

Build one connected assurance process

{assuranceSteps.map(step =>
{step.number}

{step.title}

{step.text}

Typical evidence: {step.evidence}
)}

The process should be proportionate to risk and decision importance. An exploratory engineering test may need a concise rationale; a safety-critical performance claim, clinical study or production-release decision normally needs a pre-approved and independently reviewable plan.

06

Assure the whole measurement system

The measuring instrument is only one part of the system. Fixtures, software, reference items, preparation, operator method, environment and data handling can dominate the result.

Suitability

Range, resolution, bandwidth, sensitivity and environmental capability fit the quantity and limits being assessed.

Calibration

Status, interval, reference standards and as-found results are controlled and traceable.

Repeatability

Variation when the same method is repeated under closely controlled conditions is understood.

Reproducibility

Relevant operator, instrument, site, setup or laboratory differences are represented.

Stability

Drift, warm-up, ageing and maintenance effects are monitored over the required period.

Method integrity

Fixtures, algorithms, data transformations, rounding and manual steps are verified and controlled.

A gauge repeatability and reproducibility study may be useful for some continuous production measurements, but it is not a universal template. Attribute inspection, destructive tests, automated algorithms and laboratory methods require study designs suited to their actual error structure.

07

Relate uncertainty to the acceptance limit

A result close to a specification boundary cannot be interpreted responsibly without considering measurement uncertainty. The organisation should define how uncertainty affects conformity decisions before seeing the result.

MeasurandPrecisely define the quantity intended to be measured
InfluencesIdentify reference, instrument, method, operator and environmental contributions
EvaluationEstimate contributions from data, certificates, specifications or justified models
CombinationCombine significant components using an appropriate uncertainty model
Decision ruleState how uncertainty is considered at the acceptance boundary
TraceabilityMaintain a documented chain to suitable references, with associated uncertainties

Metrological traceability does not mean that every measurement must be traceable directly to an SI unit. The reference must be appropriate to the measurand and claim; for some biological or ordinal quantities this may involve certified materials, reference procedures or agreed reference systems.

08

Justify samples from the claim—not a default number

Sample size depends on the question, expected variation, effect of interest, desired precision, confidence, statistical model, grouping and risk of an incorrect decision. The number of observations is not necessarily the number of independent experimental units.

  • Define the experimental unit and avoid treating repeated readings from one device as independent devices.
  • Use estimates of variation from relevant pilot, historical or published evidence—and record their limitations.
  • Allow for variants, lots, users, sites, conditions and interactions that the conclusion must cover.
  • For estimation, choose a sample that provides a useful confidence-interval width.
  • For comparisons, state the smallest practically important effect and selected error probabilities.
  • For reliability claims, relate exposure, failures, confidence and censoring to the stated mission or demand profile.
  • Allow for invalid or missing data without routinely replacing unfavourable observations.
  • Recalculate only through a pre-specified adaptive rule or a documented change—not after looking for a preferred outcome.

Worst-case selection can reduce testing when scientifically justified, but it must address the drivers of performance and cannot replace population evidence when variation itself is the question.

09

Select the method that matches the data and decision

Describe

Plots, distributions, ranges, proportions and summary measures reveal structure and anomalies before formal inference.

Estimate

Confidence intervals communicate the plausible range of an effect, mean, proportion, rate or reliability quantity.

Compare

Tests and models assess differences, equivalence or non-inferiority when their assumptions and margins are justified.

Relate

Regression and correlation examine relationships, while recognising confounding, non-linearity and repeated observations.

Optimise

Designed experiments efficiently investigate factors and interactions when the experimental process is controlled.

Monitor

Control charts and trend methods distinguish common variation from signals requiring investigation.

Do not select a method solely because the software offers it. Check distributional assumptions, independence, censoring, multiplicity, missingness and whether the model represents the mechanism and sampling design.

10

Stability comes before process capability

Capability indices summarise the relationship between a stable process distribution and specification limits. They are not evidence that the process is controlled, the measurement system is adequate or the specification is clinically meaningful.

  • Confirm that the process definition, subgrouping and measurement method are consistent.
  • Investigate special causes and establish statistical stability before interpreting capability.
  • Use a distribution or transformation appropriate to the observed data.
  • Distinguish short-term potential capability from longer-term performance.
  • Consider one-sided limits, non-normal data and low defect rates explicitly.
  • Link critical process characteristics to design outputs, risk controls and product acceptance.
  • Continue monitoring after validation; an initial capability result is not permanent assurance.

Process statistics support, but do not replace, controlled design transfer, process validation, maintenance, supplier controls and reaction plans.

11

Treat time, cycles and failures explicitly

Reliability evidence may involve success-run demonstration, life distributions, accelerated tests, degradation measurements, repeated demands, repairable systems or field rates. Each approach makes different assumptions.

  • Define failure, mission time or demand, environment and population before analysing reliability.
  • Record completed, failed, suspended and censored observations correctly.
  • Do not pool unlike failure mechanisms without justification.
  • Use acceleration models only when the stress-to-life relationship and failure mechanism are credible.
  • Distinguish observed test exposure from the life or reliability being claimed.
  • Report confidence bounds, not only a point estimate.
  • Compare development predictions with production, complaint and service evidence.

Apply the physical and lifecycle principles in MTL-119 — Reliability, Durability and Environmental Robustness.

12

Protect data and analysis integrity

Traceable conclusions require controlled data from acquisition to report. Preserve original observations, metadata, units, timestamps, sample identity, exclusions, transformations, software versions and analysis code where applicable.

  • Predefine data structures, naming, units, rounding and valid ranges.
  • Retain raw data and distinguish corrections from the original record.
  • Document missing, invalid, excluded and outlying observations with reasons.
  • Verify formulas, scripts, spreadsheets and statistical software used for consequential decisions.
  • Control analysis versions and reproduce the final tables and figures from retained data.
  • Protect blinding and randomisation where they are part of the study design.
  • Review important analyses independently, including assumptions and interpretation.

For clinical and IVD evidence, integrate the statistical analysis plan with MTL-120 — Clinical and Performance Evaluation.

13

Report the result, uncertainty and practical meaning

A defensible report allows a reviewer to understand what was planned, what happened, how the analysis was performed and what the data do—and do not—support.

ObjectiveDecision, hypothesis, claim and linked requirement
DesignPopulation, samples, factors, randomisation, controls and conditions
MeasurementMethod, equipment, traceability and measurement-system adequacy
AnalysisModel, assumptions, software, deviations and treatment of missing data
ResultsRaw-data location, estimates, intervals, plots, failures and anomalies
ConclusionDecision, practical importance, limitations and required follow-up

A small p-value is not a measure of clinical importance, design margin or evidence quality. Prefer effect sizes and confidence intervals alongside any hypothesis test, and explain the result in the units and context of the requirement.

14

Common misconceptions

“Thirty samples is always statistically valid.”

No universal sample number fits every distribution, claim, confidence, risk, source of variation or study design.

“More readings mean more independent evidence.”

Repeated readings can improve understanding of measurement variation, but they do not automatically increase the number of independent devices, users or lots represented.

“A calibrated instrument guarantees a valid result.”

Calibration is essential where applicable, but fixtures, methods, operators, environment, software and sampling can still make the overall measurement unsuitable.

“Passing a normality test makes a parametric analysis correct.”

Model suitability also depends on independence, study design, censoring, variance structure and whether the assumptions represent the process.

“Statistical significance proves practical importance.”

A very small effect can be statistically detectable yet irrelevant to safety or performance; a meaningful effect can remain uncertain in a weak or undersized study.

“No failures proves the product cannot fail.”

A success-run test supports only a bounded reliability statement determined by the exposure, sample size, model and confidence.

15

Practical checklist

  • Is the decision, claim and consequence of error stated?
  • Are the population, experimental unit and sources of variation defined?
  • Is the measurement method suitable for the range, tolerance and environment?
  • Are calibration, traceability, uncertainty, stability and method variation addressed?
  • Does the sample-size rationale match the statistical objective and model?
  • Are acceptance criteria approved before confirmatory data are examined?
  • Are independence, distribution, censoring, repeated measures and missing data handled correctly?
  • Are raw data, transformations, exclusions, software and analysis versions retained?
  • Does the report present estimates and uncertainty as well as pass / fail?
  • Are unexpected results investigated and connected to risk, design and process controls?
  • Will production and post-market data test whether development assumptions remain valid?
16

Authoritative external references

Use the current controlled editions and the product-specific standards, regulatory requirements and laboratory procedures applicable to the device, market, measurement and decision. This learning topic supports understanding; it does not replace qualified statistical, metrological, clinical or regulatory judgement.

KEY TAKEAWAY

Quantitative evidence is credible only when the decision, sampling, measurement process, analysis and uncertainty form one traceable argument.