Risk files are full of numbers that were never measured. An occurrence rating of 2. A probability of “remote.” A residual risk judged acceptable because thirty units were built and none of them failed.

ISO 14971 defines risk as the combination of the probability of occurrence of harm and the severity of that harm. Severity is a clinical and engineering judgment. Probability is not — it is a number your product and process data can estimate, and estimate with stated uncertainty.

This article is about that half: where statistics legitimately supplies the probability side of a risk file, and — just as importantly — what your evidence does not support.

Which half of risk is statistical

The split matters, because conflating them is how risk files end up with numbers nobody can defend.

Severity is not statistical. How badly a patient is harmed if a failure occurs is a clinical question, informed by the intended use, the clinical literature, and complaint history. No amount of process data speaks to it. A statistician who offers you a severity rating is out of their lane.

Probability is statistical — and it decomposes. ISO/TR 24971, the guidance that accompanies ISO 14971, describes splitting the probability of harm into two parts:

  • P1 — the probability that the hazardous situation occurs
  • P2 — the probability that, given the hazardous situation, harm actually results
P(harm)=P1×P2P(\text{harm}) = P_1 \times P_2

This decomposition is where the work divides cleanly:

  • P1 is a product and process question. How often does the seal leak, the dose run high, the weld fail, the reading drift out of tolerance? That is nonconformance rate, capability, and reliability — measurable from data you already collect.
  • P2 is largely a clinical question. Given a leak, how often does a patient come to harm? That is informed by clinical knowledge and post-market experience, not by your process capability study.

So the honest statement of scope is this: statistics can put a defensible number on P1, and quantify how uncertain that number is. The rest of the risk estimate depends on judgments that belong to other disciplines. Saying so explicitly makes the file stronger, not weaker.

Estimating P1 from process data

For a characteristic with a specification, P1 is the probability of falling outside it — which is exactly what a capability study estimates. Under a normal model:

P1=Pr ⁣(outside spec)=Φ(3Cpk)P_1 = \Pr\!\left(\text{outside spec}\right) = \Phi(-3 C_{pk})

That gives the familiar translation:

CpkDistance to specNonconforming (one-sided)
0.672.0σ22,216 PPM (≈ 1 in 45)
1.003.0σ1,350 PPM (≈ 1 in 740)
1.334.0σ33 PPM (≈ 1 in 30,000)
1.675.0σ0.27 PPM (≈ 1 in 3.7 million)
2.006.0σ0.001 PPM

This table is the bridge between a process study and a risk file. It is also the point at which two cautions must be stated, because the numbers get quoted far beyond what they can bear.

First, these are model extrapolations, not observations. A Cpk of 1.67 implying 0.27 PPM is a statement about a normal distribution five sigma out in the tail — a region where you almost certainly have no data at all. A capability study of 100 parts contains no information about what happens one in 3.7 million times; that figure comes entirely from the assumed shape of the curve. If the real distribution is even slightly heavier-tailed — and real manufacturing data often is — the true rate can be orders of magnitude higher. Verify normality, and treat deep-tail PPM as indicative rather than precise.

Second, capability itself is an estimate. The Cpk you computed has a confidence interval, and at typical sample sizes it is wider than people expect — see how much data a capability index needs. A P1 derived from a point estimate inherits all of that uncertainty. Use the lower confidence bound on Cpk, which gives you a conservative (higher) P1 — the direction that protects the patient.

Zero failures is not zero risk

Here is the most common statistical error in risk files, and it is worth being blunt about.

A team builds 30 units. All 30 pass. The occurrence is recorded as “Improbable.”

But zero failures in a sample does not mean the failure rate is zero — it means the failure rate is low enough that 30 units probably wouldn’t reveal it. How low? There is a clean answer, often called the rule of three: with zero failures in nn trials, the upper 95% confidence bound on the failure rate is approximately

pupper3np_{\text{upper}} \approx \frac{3}{n}

Which yields a sobering table:

Units tested (zero failures)Upper 95% bound on the true failure rate
1025.9%
309.5%
1003.0%
3001.0%
1,0000.30%
3,0000.10%
30,0000.010%

Thirty successful units is consistent with a failure rate as high as roughly 1 in 10. That is not “improbable” by any scale in use. It is barely evidence of anything.

This single fact reframes a great many risk files. The occurrence rating was not measured; it was assumed. Thirty units with zero failures is equally consistent with a true rate of one in a million and a true rate of nearly one in ten — the test simply cannot tell those apart.

The occurrence rating your evidence actually supports

Turn it around, the way a protocol should. If you want to claim a given occurrence rate, how much zero-failure evidence does that take?

Claimed rateTypical occurrence languageUnits required (zero failures)
< 1 in 100Occasional299
< 1 in 1,000Remote2,995
< 1 in 10,000Very remote29,956
< 1 in 100,000Improbable299,572
< 1 in 1,000,000Improbable / negligible2,995,731

(Occurrence scales vary by organization; the labels are illustrative of common five- and ten-point scales, not a standard.)

Read the bottom row again. Demonstrating a one-in-a-million failure rate by attribute testing alone would take about three million units. No verification build will ever do that.

This is not a counsel of despair — it is the reason the standard is built the way it is. Three legitimate routes out:

  1. Measure instead of count. Variables data extracts far more information per unit. A capability study on a measured characteristic can support a low P1 from tens of units where pass/fail would need thousands — the same trade explored in sample size justification. This is the single largest lever most teams leave unused.
  2. Use accumulated production and field data. Your risk estimate does not have to rest on the verification build alone. Tens of thousands of units of production history, properly analyzed, are real evidence — which is precisely what the standard’s production and post-production clause is for.
  3. Rely on risk control, not on demonstration. If a low occurrence cannot be demonstrated, the honest response is to reduce the risk by design or by control, not to assert a rating the data cannot carry.

Why this makes post-production surveillance load-bearing

ISO 14971 requires an active system for collecting and reviewing production and post-production information, and feeding it back into the risk assessment. It is often treated as an administrative afterthought. The arithmetic above shows it is nothing of the kind.

Pre-market data is almost never sufficient to establish a low probability of occurrence. Only accumulated production and field experience reaches the sample sizes that low occurrence claims require. Post-production surveillance is not a formality bolted onto the end of the file — for low-probability, high-severity hazards, it is the only place the evidence can come from.

That reframes the statistical tools accordingly:

  • Control charts and continued process verification are the mechanism that detects when P1 has shifted away from the value in your risk file. A process that drifts has silently invalidated a risk estimate.
  • Trending capability over time turns a one-off P1 into a monitored parameter.
  • Complaint and failure rates, with confidence bounds, either confirm the assumed occurrence or trigger a re-evaluation. A rate estimated from field data should carry an interval like any other estimate.

The question worth asking of your own system: if the true occurrence rate doubled, how long would it take us to notice? That is a statistical property of your monitoring plan, and it can be designed deliberately rather than left to chance.

Verifying that a risk control actually works

Risk controls must be verified for effectiveness — and “effectiveness” is a claim about a rate, which means a sample size question. The chain is exactly the one described in sample size justification:

the severity of the harm → the confidence/reliability the control must demonstrate → the number of units required

A control mitigating a high-severity hazard warrants a stronger demonstration than one addressing a minor annoyance. When that link is written down in a sampling policy, every individual sample size in your files becomes traceable to a risk judgment. When it isn’t, each number has to be defended on its own — and usually can’t be.

What a defensible probability estimate looks like

Pulling it together, a risk file that will hold up states:

  • What was measured, and on how many units — with the acceptance criterion set in advance.
  • The estimated P1 with an interval or bound, not a bare point estimate. ”≤ 1% at 95% confidence, based on 299 units with zero failures” is defensible. “Improbable” is not.
  • The distributional assumption behind any capability-derived rate, and how it was checked.
  • The separation of P1 from P2, with the clinical basis for P2 cited rather than assumed.
  • What would change the estimate — the monitoring signal that would trigger a re-evaluation, closing the loop into post-production activities.

None of this requires exotic statistics. It requires refusing to write down a number that the evidence does not support — and being explicit about how much uncertainty remains.

The occurrence column of a risk table looks like data. Very often it is opinion wearing the costume of data. The difference is visible to anyone who asks the sample size.


Need help putting defensible numbers behind a risk file, or sizing the evidence a risk control demands? See our sample size and power services or book a call.

ISO 14971Risk ManagementFMEAOccurrenceSample SizeProcess Capability