M9V.comBook a walkthrough

Tolerance stacks: over-specified in three places, loose in one

A worked account of how tolerance budgets drift away from the physics, why worst-case thinking hides the problem, and what a variance-driven redistribution actually changes on the floor.

In short

Why is our tolerance stack causing scrap even though every part measures in-spec?

Most production tolerance stacks are inherited from a first drawing and never re-derived against measured variation. Worst-case arithmetic weights every contributor equally, so it offers no ranking, and engineers tighten whatever is cheapest rather than whatever dominates the output variance. The signature is three features held far tighter than needed and one, often a locating datum or an unbudgeted thermal effect, carrying most of the variation. Reallocating the budget from measured standard deviations cuts scrap and inspection cost without changing design intent.

Key points

  • A tolerance stack calculated with worst-case arithmetic assumes every contributor sits at its limit simultaneously, an outcome whose probability falls sharply as the number of contributors grows.
  • Statistical stack-up combines contributors in quadrature, so a feature contributes to the output variance in proportion to the square of its standard deviation, so halving the largest contributor removes more variation than tightening three small ones.
  • Parts can all measure within their individual tolerances and still produce an out-of-tolerance assembly whenever the stack was never validated against the assembly's critical dimension.
  • Thermal expansion and fixture repeatability are frequently absent from the drawing stack entirely, which is why the dominant contributor is so often something that was never assigned a tolerance.
  • Redistribution is cost-neutral by construction: the budget released from over-specified features pays for tightening the dominant one.

A tolerance stack is a budget. Like every budget, it is usually inherited rather than derived: someone drew the first version of the part, assigned limits that looked reasonable, and every revision since has adjusted the geometry without revisiting the arithmetic underneath it. The drawing gets more detailed. The budget stays where it was in 2011.

This matters because the budget encodes an assumption about which features control the outcome. When that assumption is wrong, the plant pays twice: once for precision it does not need, and again for the scrap it did not prevent.

The symptom: everything is in spec and the assembly still fails

The clearest sign of a mis-allocated stack is an argument between two departments that are both right.

Quality reports that every component measures within its drawing limits. Assembly reports that a measurable fraction of finished units miss the critical dimension. Neither is mistaken. What they are describing is a stack that was never validated against the dimension that actually matters.

Individual feature tolerances are a statement about parts. The critical dimension is a statement about the assembly. The two are connected by the stack-up, and if that connection was never computed, or was computed once for a geometry that has since changed, then in-spec parts producing out-of-spec assemblies is not a contradiction. It is the expected result.

Why worst-case arithmetic hides the allocation problem

Most stacks are still calculated the way they are taught: sum the tolerances along the dimensional chain and check the total against the requirement.

Worst-case analysis assumes every contributor sits at its extreme, in the direction that hurts, at the same moment. For a two-part stack that is a reasonable thing to guard against. For a fourteen-part stack it describes an event that will not occur in the life of the programme, because it requires fourteen independent variables to land simultaneously at the same tail.

The arithmetic has a specific consequence for allocation. In a worst-case sum, every contributor enters linearly and with equal weight. A tenth of a millimetre from a bore diameter counts exactly as much as a tenth from a fixture location. The method therefore offers no guidance about which tolerance to tighten. Faced with a total that exceeds the requirement and no way to rank the contributors, engineers tighten whatever is cheapest and least disruptive to tighten.

That is a rational response to a method that provides no ranking. It is also how a stack ends up over-specified in exactly the places where tightening was easy.

Statistical stack-up ranks the contributors

Statistical tolerance analysis combines contributors in quadrature rather than linearly. Under the root-sum-square approach, the variance of the output is the sum of the variances of the contributors, each weighted by the square of its sensitivity coefficient, which is how much the output moves per unit of movement in that input.

The practical consequence is the whole point: contribution scales with the square of the standard deviation. A feature with twice the variation of another contributes four times as much to the output variance. In a stack with one dominant contributor and several small ones, tightening the small ones is close to pointless, and this is arithmetically visible rather than a matter of judgement.

This is what produces the pattern in the title. Run a sensitivity analysis on a stack that has never had one and the distribution of contributions is rarely even. A small number of features dominate, several contribute almost nothing, and the ones that dominate are frequently not the ones carrying the tightest limits.

The three over-specified features are usually over-specified because they were easy to tighten: a machined face, a reamed hole, a ground surface, all on operations with capable processes and cheap adjustment. The one loose feature is loose because tightening it was awkward, expensive, or belonged to somebody else.

The contributor that is not on the drawing

The most common dominant contributor is one that has no tolerance at all, because it is not a feature.

Three recurring examples:

Fixture repeatability. The stack accounts for part geometry and says nothing about how repeatably the part is located. If the fixture’s clamping sequence produces a two-hundredth of a millimetre of variation in datum contact, that variation propagates into every downstream dimension referenced from that datum, and it appears nowhere in the drawing arithmetic.

Thermal state. Aluminium has a coefficient of thermal expansion around 23 micrometres per metre per kelvin. A 400 mm feature measured at 20 °C and machined at 28 °C differs by roughly 74 micrometres. On a stack budgeted in tenths of a millimetre, that is not a rounding error. Stacks are typically written as if everything is at the 20 °C reference temperature; production is not.

Measurement error. The gauge is part of the system. If the measurement system’s variation is a meaningful fraction of the tolerance band, some of the observed spread is the gauge rather than the process, and tightening the process cannot remove it. This is what a gauge repeatability and reproducibility study is for, and the NIST/SEMATECH handbook’s treatment of measurement process characterisation is the standard reference for separating the two.

Contributor Varies with On the drawing?
Fixture repeatability Clamping sequence, wear, operator No
Thermal state Shop temperature, machine warm-up, material Rarely
Measurement error Instrument, fixturing of the instrument, operator No

A stack that omits all three is not describing the assembly. It is describing an idealised drawing of the assembly, at reference temperature, measured perfectly, in a fixture that locates identically every time.

What redistribution actually looks like

Redistribution is not “tighten everything”. It is a reallocation of a fixed budget, and the arithmetic is what makes it affordable.

The sequence:

  1. Measure the contributors as they actually behave. Not the drawing limits but the observed distributions, sampled across shifts, tool states and material lots. A sample taken from one shift immediately after a tool change describes that shift, not the process.
  2. Compute sensitivity coefficients. How much does the critical dimension move per unit of movement in each contributor? For a simple linear chain these are ±1. For anything involving angles, offsets or rotation about a datum, they are not, and assuming they are is a common way to get a confident wrong answer.
  3. Rank by contribution to variance. Standard deviation times sensitivity, squared. This is the list that tells you where the money is.
  4. Reallocate. Loosen the features contributing negligibly. Spend the released budget on the dominant contributor.
  5. Re-verify against the requirement, and confirm the new allocation is capable. A tolerance the process cannot hold is not a budget, it is a wish.

Step four is where the objection arrives, and it is worth taking seriously. Loosening a tolerance feels like a reduction in quality even when the arithmetic says the assembly is insensitive to that feature. The counter is that the assembly’s behaviour is the thing being controlled, and the assembly does not know what the drawing says. If a feature contributes two percent of the output variance, holding it to a tenth of what it needs is not quality. It is expenditure.

A worked example

Numbers make the asymmetry concrete. Take a five-contributor stack on a critical dimension with a requirement of ±0.30 mm, all sensitivity coefficients ±1 for simplicity, and measured standard deviations rather than drawing limits:

Contributor σ (mm) σ² Share of variance
A: bore diameter 0.020 0.000400 4%
B: face location 0.025 0.000625 6%
C: slot width 0.018 0.000324 3%
D: fixture repeatability 0.090 0.008100 78%
E: thermal, 8 K swing 0.030 0.000900 9%
Total 0.010349

The combined standard deviation is the square root of the total variance, roughly 0.102 mm. At three standard deviations that is ±0.305 mm against a ±0.30 mm requirement. That is marginal, and the line will produce excursions.

Now look at where the variance lives. Contributors A, B and C together account for thirteen percent of it. Halve all three, an expensive programme touching three separate processes, and the combined standard deviation falls to about 0.097 mm. A five percent improvement for three separate process changes.

Halve contributor D alone, and the total variance drops to 0.004274, giving a combined standard deviation near 0.065 mm. That is a thirty-six percent improvement from one fixture rebuild.

The arithmetic is not subtle, and that is the point: it is invisible under worst-case summation, where A, B, C, D and E all look like tolerances of a similar size. Only when each contributor is squared does the fixture separate from the rest of the field. The three “over-specified” features in the title are A, B and C, probably already held tighter than this table implies, because they were the ones easy to tighten in previous rounds.

Note also what happens to E. An eight-kelvin swing contributes more variance than any of the three machined features, and it appears on no drawing. It is not a tolerance; it is the shop floor in August.

The second-order savings

The direct saving is scrap. Two others are usually larger and less visible.

Inspection. This is where a redistribution meets inspection strategy directly. A feature held far tighter than the assembly requires is usually inspected at a frequency that matches its stated criticality rather than its actual influence. Re-ranking the contributors also re-ranks the inspection plan. Fewer gauges, better placed, sampled at rates derived from the actual variance structure rather than from convention.

Process capability headroom. A feature running at a capability index near the acceptance threshold is a feature that will generate excursions whenever anything drifts. Loosening a non-dominant feature from marginal capability to comfortable capability removes a recurring source of line stoppages and containment activity that never appeared in the scrap number at all, because the parts were caught.

Where this goes wrong

Three failure modes worth naming, because each produces a redistribution that looks right and fails in production.

Assuming independence that does not exist. Root-sum-square assumes contributors are independent. Two features machined in the same setup on the same tool are not independent; they share the tool’s wear state and the setup’s location error. Treating correlated contributors as independent understates the output variance, and the stack passes on paper.

Sampling too narrowly. A distribution estimated from a single shift, a single lot and a fresh tool describes the best case. The redistribution derived from it will be tighter than reality supports, and it will fail when the tool is halfway through its life.

Redistributing without changing the inspection plan. If the loosened features remain on the same inspection frequency and the tightened one is not measured any more closely, the plant carries the new risk profile with the old detection profile. The saving is real and the exposure is new.

The engineering argument, not the software argument

None of this requires a particular tool. It requires that the stack be treated as a derived result rather than an inherited constant, and that the derivation use measured variation rather than drawing limits.

The reason it so rarely happens is not ignorance of the method. It is that re-deriving a stack means measuring a lot of features across a lot of conditions, computing sensitivities for a geometry that may not be documented in a convenient form, and then having a conversation about loosening a tolerance, which is the least popular sentence in a quality review.

That is a workload problem, not a knowledge problem. It is also the specific workload that reading the line’s existing measurement history can absorb: the distributions are already in the historian and the CMM exports, and the sensitivity analysis is arithmetic once the geometry is expressed properly. What the exercise returns is a ranked list of where the variance actually lives, which is usually the first time anyone has seen that list for the assembly in question.

The three tight features and the one loose one are not a universal law. They are what a budget looks like when it has been adjusted for a decade by people optimising locally, each of whom made a defensible decision, none of whom had the ranked list.

A last note on sequencing. Re-derive the stack before buying capability, not after. The usual order is reversed: an excursion appears, a capital request follows for a more precise machine on the operation that seems responsible, and the analysis is written afterwards to justify it. Run the ranking first and that request is sometimes cancelled outright, because the operation under suspicion turns out to contribute four percent of the variance and the fixture nobody costed contributes most of the rest.

Questions

Should we use worst-case or statistical tolerance analysis?
Use worst-case when an out-of-tolerance assembly is a safety or total-loss event and the contributor count is small. Use statistical analysis when contributors are numerous and independent, and when you can measure their actual distributions. Many teams use worst-case by default because it feels conservative, but on a long stack it is so conservative that it forces tolerances nobody can hold economically, and the response is usually an undocumented deviation rather than a tighter process.
How many parts do we need to measure before trusting a redistribution?
Enough to estimate a standard deviation with a usable confidence interval, and drawn across the sources of variation you care about: different shifts, different tool states, different material lots. A sample from a single shift on a freshly dressed tool will understate variation and produce a redistribution that fails in week three. The number matters less than the spread of conditions it covers, because a narrow sample gives a confident estimate of the wrong distribution.
Does loosening a tolerance ever reduce quality?
Loosening a feature that contributes little to the critical dimension does not degrade the assembly, because the assembly was never sensitive to it. What it does change is the inspection story: a feature with a wider band needs a documented rationale, because the next engineer will read the open tolerance as carelessness rather than as a derived result.
What if the dominant contributor is a purchased part?
Then the budget conversation moves to the supplier, and the useful artefact is the sensitivity analysis rather than a tighter print. A supplier asked to halve a tolerance will quote a price; a supplier shown that their feature drives 60 percent of your assembly variance can often propose a cheaper process change that achieves the same effect.

Sources

  1. ASME Y14.5 — Dimensioning and TolerancingASMEThe governing standard for geometric dimensioning and tolerancing in North American practice.
  2. ISO 1101:2017 — Geometrical product specifications (GPS)ISOGeometrical tolerancing of form, orientation, location and run-out.
  3. Dimensional Metrology programmeNISTMeasurement uncertainty and traceability, the basis for trusting a measured standard deviation.
  4. NIST/SEMATECH e-Handbook of Statistical Methods — Measurement Process CharacterizationNIST/SEMATECHGauge R&R and measurement system analysis, used here for the measurement-error argument.