Every measurement you record is the true value plus measurement error. When the error is large relative to what you’re trying to see, your data describes your gauge more than your product — and every downstream conclusion inherits that noise.
A Gage R&R study separates the two. The trouble is that it returns several numbers that can point in different directions, and the standard 10%/30% thresholds get applied to whichever one the software printed first.
What the study decomposes
A crossed Gage R&R (multiple operators measuring the same parts, several times each) partitions total variation into:
- Repeatability — equipment variation. The same operator measuring the same part twice and getting different answers.
- Reproducibility — appraiser variation. Different operators getting systematically different answers on the same part. Often reveals technique or fixturing differences.
- Part-to-part — genuine product variation. This is the signal you want.
Gage R&R is repeatability and reproducibility combined:
Note that variances add, not standard deviations. A common error is adding the percentages — they don’t sum, which is why the components in a variance table look inconsistent with the percentage column unless you square them.
The three ratios, and why they disagree
Here’s the source of most confusion. The same study reports GRR against different denominators.
%Study Variation (%SV)
Measurement error as a fraction of the total variation observed in your study. This answers: can this gauge distinguish the parts I gave it?
The catch: the denominator depends on the parts you selected. Choose parts spanning a wide range and %SV looks excellent. Choose parts from a single tight lot and the same gauge looks terrible — because there’s little part variation to compare against.
This is the single most misread statistic in MSA. A bad %SV may mean a bad gauge, or it may mean you sampled parts that were too similar.
%Tolerance (%P/T)
Measurement error as a fraction of the specification width. This answers: can this gauge decide whether a part conforms?
Because the denominator is fixed by the specification rather than by your part selection, %Tolerance is usually the more meaningful criterion for inspection and release decisions. It’s also the number that matters for guard-banding.
Number of Distinct Categories (ndc)
How many distinct groups the measurement system can reliably tell apart across the part range. AIAG guidance calls for ndc ≥ 5; below 4, the system is effectively sorting parts into “big” and “small” and cannot support variables SPC or capability work.
An ndc of 2 means your gauge is functioning as an attribute gauge no matter what decimals it displays.
Acceptance guidelines
The conventional AIAG bands, applied to %SV or %Tolerance:
| %GRR | Verdict |
|---|---|
| < 10% | Acceptable |
| 10–30% | Conditionally acceptable — justify based on the importance of the characteristic, cost of the gauge, and cost of a wrong decision |
| > 30% | Unacceptable. Improve the measurement system before trusting the data |
These are guidelines, not requirements — the same risk-based logic that applies to capability targets applies here. For a high-severity characteristic, 15% may be unacceptable; for a low-risk one, 25% may be fine with a rationale.
Judge against the decision you’re making. For process monitoring and capability, %SV (with well-chosen parts) is the relevant view. For conformance decisions, %Tolerance is. When they disagree, both facts are true — they answer different questions.
Designing a study that gives a real answer
Most bad Gage R&R results come from study design, not from the gauge.
- Parts must span the real range of process variation — ideally the full expected spread, not a convenience sample from one shift. Ten parts is typical, and they should be deliberately selected to represent the range. This is the fix for a misleading %SV.
- Operators must be the people who actually do the job, using the documented method. A study run by three engineers who all trained together underestimates reproducibility in production.
- Randomize and blind the order. If operators can see the previous reading — or recognize the part — you measure their memory, not the gauge.
- Three operators × ten parts × three trials (90 measurements) is the common design. Fewer trials weakens the repeatability estimate; fewer parts weakens the part-to-part estimate.
- Confirm resolution first. The classic rule is that the gauge increment should be no more than one tenth of the process spread or the tolerance. A gauge that reads to 0.001” on a 0.002” tolerance cannot produce a useful study.
Beyond R&R: bias, linearity, and stability
Gage R&R measures precision only. A gauge can be beautifully repeatable and consistently wrong. A complete measurement systems analysis also covers:
- Bias — systematic offset from a reference standard.
- Linearity — whether bias changes across the measurement range. A gauge accurate mid-range may drift at the extremes, exactly where conformance decisions get made.
- Stability — whether performance drifts over time, which is what calibration intervals exist to control.
For medical device test method validation (TMV), these are typically expected alongside R&R. A validation package that reports only %GRR is incomplete.
What to do with a failing system
In rough order of effectiveness:
- Improve the method and fixturing. Most reproducibility problems are how the part is located and held, not the instrument. This is usually the cheapest large gain.
- Clarify and retrain to the procedure. If operators interpret the method differently, the SOP is ambiguous.
- Average repeated measurements. Averaging readings reduces the repeatability component by . Effective, but it costs inspection time on every part forever.
- Buy a better gauge. Sometimes necessary, but it doesn’t fix a fixturing or method problem — and often people start here.
- Widen the tolerance — only with a genuine engineering and risk basis, never to make a measurement problem disappear on paper.
Attribute systems have their own method — an attribute agreement analysis rather than variables R&R — but the same principle applies: quantify the error before trusting the data.
Need an MSA or test method validation designed, analyzed, or defended? See our measurement systems analysis services or book a call.