What you will learn
By the end of this topic, you should be able to frame a quantitative decision, distinguish important statistical and metrological concepts, identify relevant sources of variation, select a proportionate study design and sample size, assess a measurement system, interpret uncertainty and traceability, use capability and reliability evidence appropriately, and report conclusions without overstating what the data demonstrate.
Statistics turn observations into decisions
Medical-device development depends on decisions made with incomplete information: whether a requirement is met, whether a risk control is effective, whether a measurement method is adequate, whether a process is stable, whether two designs differ and whether performance remains acceptable over time.
Statistical methods help quantify variation and uncertainty. Measurement assurance asks whether the observations themselves are sufficiently reliable for the intended decision. Neither discipline should be reduced to a calculation performed after testing.
Define the decision, claim, population and measurement process before choosing the statistical method or sample size. A sophisticated analysis cannot rescue unsuitable data.
Start with the decision and its consequences
A useful plan begins with a plain-language decision statement. “Test ten units” is not a decision. “Demonstrate that dose accuracy meets the specified limits across the released configuration and foreseeable operating range” is.
Question
What exactly must the evidence allow the organisation to conclude?
Population
Which devices, lots, users, environments, sites, time periods or software configurations must be represented?
Errors
What are the safety, quality and business consequences of accepting an inadequate design or rejecting an adequate one?
Effect
What difference, degradation or failure rate would be practically important—not merely statistically detectable?
Evidence
Is the purpose exploration, estimation, comparison, verification, validation, monitoring or prediction?
Action
What action follows a pass, failure, inconclusive result, trend or unexpected observation?
Connect each decision to the approved requirement and protocol principles in MTL-106 — Verification and Validation.
Use statistical and measurement terms precisely
Bias is a systematic difference; variability describes dispersion. Resolution is the smallest displayed or detectable increment, not proof of accuracy. Calibration establishes the relationship between indication and reference values under stated conditions; it does not by itself demonstrate that the complete measurement process is suitable for its use.
Model the variation that exists in real use
Apparent consistency can result from testing an unrealistically narrow set of samples. The study should expose the variation relevant to the claim while controlling variation that would obscure the decision.
- Product units, lots, cavities, tooling, suppliers, component tolerances and ageing.
- Users, operators, laboratories, sites, shifts, fixtures and setup methods.
- Temperature, humidity, supply conditions, orientation, vibration and electromagnetic environment.
- Reference materials, biological samples, matrices, interferents and specimen handling.
- Software, algorithms, configuration, data processing and rounding.
- Time-related effects such as drift, warm-up, wear, maintenance and recalibration.
- Sampling, missing values, excluded observations and repeated measurements on the same unit.
Use MTL-102 — Intended Purpose, Users and Use Environments to define representative conditions and MTL-103 — User Needs and Design Inputs to translate them into measurable limits.
Build one connected assurance process
{step.title}
{step.text}
Typical evidence: {step.evidence}The process should be proportionate to risk and decision importance. An exploratory engineering test may need a concise rationale; a safety-critical performance claim, clinical study or production-release decision normally needs a pre-approved and independently reviewable plan.
Assure the whole measurement system
The measuring instrument is only one part of the system. Fixtures, software, reference items, preparation, operator method, environment and data handling can dominate the result.
Suitability
Range, resolution, bandwidth, sensitivity and environmental capability fit the quantity and limits being assessed.
Calibration
Status, interval, reference standards and as-found results are controlled and traceable.
Repeatability
Variation when the same method is repeated under closely controlled conditions is understood.
Reproducibility
Relevant operator, instrument, site, setup or laboratory differences are represented.
Stability
Drift, warm-up, ageing and maintenance effects are monitored over the required period.
Method integrity
Fixtures, algorithms, data transformations, rounding and manual steps are verified and controlled.
A gauge repeatability and reproducibility study may be useful for some continuous production measurements, but it is not a universal template. Attribute inspection, destructive tests, automated algorithms and laboratory methods require study designs suited to their actual error structure.
Relate uncertainty to the acceptance limit
A result close to a specification boundary cannot be interpreted responsibly without considering measurement uncertainty. The organisation should define how uncertainty affects conformity decisions before seeing the result.
Metrological traceability does not mean that every measurement must be traceable directly to an SI unit. The reference must be appropriate to the measurand and claim; for some biological or ordinal quantities this may involve certified materials, reference procedures or agreed reference systems.
Justify samples from the claim—not a default number
Sample size depends on the question, expected variation, effect of interest, desired precision, confidence, statistical model, grouping and risk of an incorrect decision. The number of observations is not necessarily the number of independent experimental units.
- Define the experimental unit and avoid treating repeated readings from one device as independent devices.
- Use estimates of variation from relevant pilot, historical or published evidence—and record their limitations.
- Allow for variants, lots, users, sites, conditions and interactions that the conclusion must cover.
- For estimation, choose a sample that provides a useful confidence-interval width.
- For comparisons, state the smallest practically important effect and selected error probabilities.
- For reliability claims, relate exposure, failures, confidence and censoring to the stated mission or demand profile.
- Allow for invalid or missing data without routinely replacing unfavourable observations.
- Recalculate only through a pre-specified adaptive rule or a documented change—not after looking for a preferred outcome.
Worst-case selection can reduce testing when scientifically justified, but it must address the drivers of performance and cannot replace population evidence when variation itself is the question.
Select the method that matches the data and decision
Describe
Plots, distributions, ranges, proportions and summary measures reveal structure and anomalies before formal inference.
Estimate
Confidence intervals communicate the plausible range of an effect, mean, proportion, rate or reliability quantity.
Compare
Tests and models assess differences, equivalence or non-inferiority when their assumptions and margins are justified.
Relate
Regression and correlation examine relationships, while recognising confounding, non-linearity and repeated observations.
Optimise
Designed experiments efficiently investigate factors and interactions when the experimental process is controlled.
Monitor
Control charts and trend methods distinguish common variation from signals requiring investigation.
Do not select a method solely because the software offers it. Check distributional assumptions, independence, censoring, multiplicity, missingness and whether the model represents the mechanism and sampling design.
Stability comes before process capability
Capability indices summarise the relationship between a stable process distribution and specification limits. They are not evidence that the process is controlled, the measurement system is adequate or the specification is clinically meaningful.
- Confirm that the process definition, subgrouping and measurement method are consistent.
- Investigate special causes and establish statistical stability before interpreting capability.
- Use a distribution or transformation appropriate to the observed data.
- Distinguish short-term potential capability from longer-term performance.
- Consider one-sided limits, non-normal data and low defect rates explicitly.
- Link critical process characteristics to design outputs, risk controls and product acceptance.
- Continue monitoring after validation; an initial capability result is not permanent assurance.
Process statistics support, but do not replace, controlled design transfer, process validation, maintenance, supplier controls and reaction plans.
Treat time, cycles and failures explicitly
Reliability evidence may involve success-run demonstration, life distributions, accelerated tests, degradation measurements, repeated demands, repairable systems or field rates. Each approach makes different assumptions.
- Define failure, mission time or demand, environment and population before analysing reliability.
- Record completed, failed, suspended and censored observations correctly.
- Do not pool unlike failure mechanisms without justification.
- Use acceleration models only when the stress-to-life relationship and failure mechanism are credible.
- Distinguish observed test exposure from the life or reliability being claimed.
- Report confidence bounds, not only a point estimate.
- Compare development predictions with production, complaint and service evidence.
Apply the physical and lifecycle principles in MTL-119 — Reliability, Durability and Environmental Robustness.
Protect data and analysis integrity
Traceable conclusions require controlled data from acquisition to report. Preserve original observations, metadata, units, timestamps, sample identity, exclusions, transformations, software versions and analysis code where applicable.
- Predefine data structures, naming, units, rounding and valid ranges.
- Retain raw data and distinguish corrections from the original record.
- Document missing, invalid, excluded and outlying observations with reasons.
- Verify formulas, scripts, spreadsheets and statistical software used for consequential decisions.
- Control analysis versions and reproduce the final tables and figures from retained data.
- Protect blinding and randomisation where they are part of the study design.
- Review important analyses independently, including assumptions and interpretation.
For clinical and IVD evidence, integrate the statistical analysis plan with MTL-120 — Clinical and Performance Evaluation.
Report the result, uncertainty and practical meaning
A defensible report allows a reviewer to understand what was planned, what happened, how the analysis was performed and what the data do—and do not—support.
A small p-value is not a measure of clinical importance, design margin or evidence quality. Prefer effect sizes and confidence intervals alongside any hypothesis test, and explain the result in the units and context of the requirement.
Common misconceptions
“Thirty samples is always statistically valid.”
No universal sample number fits every distribution, claim, confidence, risk, source of variation or study design.
“More readings mean more independent evidence.”
Repeated readings can improve understanding of measurement variation, but they do not automatically increase the number of independent devices, users or lots represented.
“A calibrated instrument guarantees a valid result.”
Calibration is essential where applicable, but fixtures, methods, operators, environment, software and sampling can still make the overall measurement unsuitable.
“Passing a normality test makes a parametric analysis correct.”
Model suitability also depends on independence, study design, censoring, variance structure and whether the assumptions represent the process.
“Statistical significance proves practical importance.”
A very small effect can be statistically detectable yet irrelevant to safety or performance; a meaningful effect can remain uncertain in a weak or undersized study.
“No failures proves the product cannot fail.”
A success-run test supports only a bounded reliability statement determined by the exposure, sample size, model and confidence.
Practical checklist
- Is the decision, claim and consequence of error stated?
- Are the population, experimental unit and sources of variation defined?
- Is the measurement method suitable for the range, tolerance and environment?
- Are calibration, traceability, uncertainty, stability and method variation addressed?
- Does the sample-size rationale match the statistical objective and model?
- Are acceptance criteria approved before confirmatory data are examined?
- Are independence, distribution, censoring, repeated measures and missing data handled correctly?
- Are raw data, transformations, exclusions, software and analysis versions retained?
- Does the report present estimates and uncertainty as well as pass / fail?
- Are unexpected results investigated and connected to risk, design and process controls?
- Will production and post-market data test whether development assumptions remain valid?
Authoritative external references
- ISO 13485:2016 — Medical devices — Quality management systems — Requirements for regulatory purposes
- ISO 5725-1:2023 — Accuracy (trueness and precision) of measurement methods and results — General principles and definitions
- ISO / IEC 17025:2017 — General requirements for the competence of testing and calibration laboratories
- BIPM / JCGM — Guides in metrology, including the GUM and International Vocabulary of Metrology
- NIST / SEMATECH e-Handbook of Statistical Methods
- FDA — Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests
Use the current controlled editions and the product-specific standards, regulatory requirements and laboratory procedures applicable to the device, market, measurement and decision. This learning topic supports understanding; it does not replace qualified statistical, metrological, clinical or regulatory judgement.