Somewhere in most manufacturing files there is a control chart with the specification limits drawn across it as two red lines. Every point sits comfortably between them, and the chart is signed off as evidence the process is in control.

It is evidence of no such thing. A chart drawn that way is answering a question nobody asked, and it is silently trained to stay quiet.

Two limits, two questions

The confusion is understandable — both are horizontal lines on a plot of the same characteristic — but they originate in different places and answer different questions.

Control limitsSpecification limits
Come fromThe process, computed from dataDesign, the customer, or a standard
Answer”Is this process behaving as it has been?""Is this part acceptable?”
Apply toThe statistic being plottedIndividual units
Change whenThe process genuinely changesThe requirement changes

Neither constrains the other. A process can sit perfectly inside its control limits and produce scrap on every unit. A process can be wildly out of control and still, for now, fall inside a generous tolerance. Control limits tell you whether the process is predictable; specification limits tell you whether the output is acceptable. You need both answers, and you cannot get them from one pair of lines.

Where control limits actually come from

Control limits are estimated from the process itself — specifically from within-subgroup variation. For an Xˉ\bar X chart with subgroups of size nn, the limits are

xˉˉ±3 σ^n,σ^=Rˉd2\bar{\bar{x}} \pm 3\,\frac{\hat\sigma}{\sqrt{n}}, \qquad \hat\sigma = \frac{\bar R}{d_2}

where d2d_2 depends only on subgroup size. For n=5n = 5, d2=2.3259d_2 = 2.3259. If you estimate from subgroup standard deviations instead, the unbiasing constant is c4=0.9400c_4 = 0.9400 at the same nn.

The detail that matters: σ^\hat\sigma is built from variation inside subgroups. That is a deliberate choice, and it is what makes the chart a comparison of short-term noise against longer-term behaviour. It is also the assumption that most often gets quietly violated — see rational subgrouping below.

Why specification limits never belong on a control chart

The decisive objection is not philosophical. It is that they are not on the same scale.

An Xˉ\bar X chart plots subgroup means, and means vary less than individuals by a factor of n\sqrt{n}. Expressed in the units of an individual part, the control limits sit at:

Subgroup sizeControl limits, in individual σ\sigma
n=2n = 2±2.12σ\pm 2.12\sigma
n=4n = 4±1.50σ\pm 1.50\sigma
n=5n = 5±1.34σ\pm 1.34\sigma
n=10n = 10±0.95σ\pm 0.95\sigma

So with subgroups of five, your control limits live at about ±1.34σ\pm 1.34\sigma, while a specification giving you a Cpk of 1.33 sits at ±4σ\pm 4\sigma. Draw both on one axis and the specification lines land three times further out than the control limits. The process looks gloriously comfortable no matter what the chart is trying to say.

That is the practical damage. Operators learn, correctly, that the red lines are the ones that matter and that points near the control limits are nothing to act on. The chart stops being a detection tool and becomes decoration. If you want to show capability, show a capability study — a histogram against specification, or Cpk with its interval. Do not overload the control chart with a second, louder question.

The one case where the scales do coincide is an individuals chart, where n=1n = 1. Even there the lines answer different questions, and the habit is worth breaking everywhere.

The arithmetic of a false alarm

Why three sigma? Assume the process is stable and the plotted statistic is approximately normal. Then a point falls outside ±3σ\pm 3\sigma with probability

P=2 Φ(−3)=0.0027P = 2\,\Phi(-3) = 0.0027

That is 0.27%, or roughly one false alarm every 1/0.0027≈3701/0.0027 \approx 370 points. That number — the in-control average run length, or ARL — is the honest cost of running the chart.

The choice of three is economic, not a law of nature. Two-sigma limits would catch shifts sooner, at a false-alarm rate of 4.55% — an ARL of about 22. A chart that cries wolf every 22 points is a chart nobody investigates, and an ignored chart detects nothing at all. Three sigma is the compromise that keeps people responding to signals.

What three-sigma limits will and will not catch

Sensitivity is where teams are most often surprised. For a sustained shift in the mean, with only the outside-the-limits rule in play:

ShiftIndividuals (n=1n=1)Xˉ\bar X chart (n=5n=5)
0.5σ0.5\sigma155 points33 subgroups
1.0σ1.0\sigma44 points4.5 subgroups
1.5σ1.5\sigma15 points1.6 subgroups
2.0σ2.0\sigma6.3 points1.1 subgroups
3.0σ3.0\sigma2.0 points1.0 subgroups

An individuals chart takes about 15 points on average to notice a shift of one and a half sigma. If you sample once per shift, that is a fortnight of drift before the chart speaks. Read that as a design constraint, not a defect: it tells you the sampling frequency your chart needs in order to catch the size of shift you actually care about.

It is also the real argument for subgrouping. Moving from individuals to subgroups of five takes the same 1.5σ1.5\sigma shift from 15 points to under 2 subgroups. Subgrouping is not an administrative convenience — it is how you buy detection power.

Runs rules buy sensitivity, and you pay for it

Because a single point beyond three sigma is slow to catch small shifts, most software adds runs rules — the Western Electric set being the usual default:

  1. One point beyond 3σ3\sigma
  2. Two of three consecutive points beyond 2σ2\sigma, same side
  3. Four of five consecutive points beyond 1σ1\sigma, same side
  4. Eight consecutive points on one side of the centre line

Each is a separate opportunity to signal, so the false alarms accumulate. Rule 4 alone fires with probability 2(1/2)8=0.00782(1/2)^8 = 0.0078 at any given point — an ARL of 128 on its own. Running all four together brings the in-control ARL down to roughly 92 — I simulated 50,000 runs and got 92.3, against the published value of 91.75.

That is the trade, stated plainly: switching on the full Western Electric set makes your chart about four times more likely to raise a false alarm, in exchange for catching small sustained shifts far sooner. That may well be the right trade. It should be a decision you made, recorded in the control plan, rather than whatever your software shipped with — and your operators should know how often the chart is expected to cry wolf.

Rational subgrouping decides what the chart can see

Since the limits are built from within-subgroup variation, the way you form subgroups determines what the chart is capable of detecting. The principle: a subgroup should contain only common-cause variation, so that anything you want to detect shows up between subgroups rather than inside them.

The classic failure is a multi-cavity mould. Take one part from each of five cavities, call that a subgroup, and the cavity-to-cavity difference is now inside the subgroup. It inflates Rˉ\bar R, which inflates σ^\hat\sigma, which widens the control limits. The chart becomes blind to precisely the variation you built it to catch — and it looks well-behaved while doing so.

The same trap catches subgroups spanning shifts, lots, or raw material batches. If a source of variation matters, keep it between subgroups: chart the cavities separately, or subgroup within a single cavity and use a second chart for cavity differences. Wide, comfortable control limits are more often a symptom of bad subgrouping than of a well-behaved process.

Stable and capable are different failures

Because they answer different questions, the two limits give you a two-by-two rather than a single verdict:

  • In control, capable. Predictable and meeting requirements. Monitor.
  • In control, not capable. Predictable and consistently wrong. There is nothing to troubleshoot — no special cause exists. You must re-centre the process or reduce its common-cause variation, which usually means changing the process, not policing it.
  • Out of control, apparently capable. The most dangerous cell, and the one those red lines on the chart create. Today’s parts pass; the process is unpredictable, so tomorrow’s are a coin toss.
  • Out of control, not capable. Find and remove the special causes before anything else.

That second cell is worth dwelling on, because it is where teams waste the most effort — launching investigations into a process that is behaving exactly as designed and simply is not good enough. And the third is why capability indices computed on an unstable process are meaningless: as covered in Cpk vs. Ppk, you cannot estimate a parameter of a distribution that will not hold still. Stability is a prerequisite for capability, not a parallel check.

When the chart signals

  1. Investigate before adjusting. A signal is evidence of a special cause, not an instruction to turn a dial. Adjusting in response to common-cause variation increases variation — the lesson of Deming’s funnel experiment.
  2. Record what you found, including when you found nothing. Roughly one signal in every 370 points is expected to be a false alarm, and a documented “investigated, no cause identified” is a legitimate outcome.
  3. Do not recalculate the limits to make the signal disappear. This is the most common sin in practice and it destroys the chart’s purpose. Limits are recalculated after a deliberate, documented process change — not after an inconvenient point.
  4. Check whether the measurement system is the special cause before chasing the process. A gauge that has drifted produces signals indistinguishable from a process shift; see interpreting Gage R&R.

Control limits are a claim about what your process does. Specification limits are a claim about what it must do. Keeping the two apart is what lets a chart tell you something you did not already know.


Need control charts that detect what actually matters — the right chart, the right subgroups, and limits you can defend? See our statistical process control services or book a call.

SPCControl ChartsProcess CapabilityProcess ValidationRational Subgrouping