# Tolerance stacks: over-specified in three places, loose in one

Source: https://m9v.com/en/blog/tolerance-stack-up-over-specified
Published: 2026-08-24
Updated: 2026-08-24

**Question answered:** Why is our tolerance stack causing scrap even though every part measures in-spec?

**Summary:** Most production tolerance stacks are inherited from a first drawing and never re-derived against measured variation. Worst-case arithmetic weights every contributor equally, so it offers no ranking, and engineers tighten whatever is cheapest rather than whatever dominates the output variance. The signature is three features held far tighter than needed and one, often a locating datum or an unbudgeted thermal effect, carrying most of the variation. Reallocating the budget from measured standard deviations cuts scrap and inspection cost without changing design intent.

A tolerance stack is a budget. Like every budget, it is usually inherited rather than derived: someone drew the first version of the part, assigned limits that looked reasonable, and every revision since has adjusted the geometry without revisiting the arithmetic underneath it. The drawing gets more detailed. The budget stays where it was in 2011.

This matters because the budget encodes an assumption about which features control the outcome. When that assumption is wrong, the plant pays twice: once for precision it does not need, and again for the scrap it did not prevent.

## The symptom: everything is in spec and the assembly still fails

The clearest sign of a mis-allocated stack is an argument between two departments that are both right.

Quality reports that every component measures within its drawing limits. Assembly reports that a measurable fraction of finished units miss the critical dimension. Neither is mistaken. What they are describing is a stack that was never validated against the dimension that actually matters.

Individual feature tolerances are a statement about parts. The critical dimension is a statement about the assembly. The two are connected by the stack-up, and if that connection was never computed, or was computed once for a geometry that has since changed, then in-spec parts producing out-of-spec assemblies is not a contradiction. It is the expected result.

## Why worst-case arithmetic hides the allocation problem

Most stacks are still calculated the way they are taught: sum the tolerances along the dimensional chain and check the total against the requirement.

Worst-case analysis assumes every contributor sits at its extreme, in the direction that hurts, at the same moment. For a two-part stack that is a reasonable thing to guard against. For a fourteen-part stack it describes an event that will not occur in the life of the programme, because it requires fourteen independent variables to land simultaneously at the same tail.

The arithmetic has a specific consequence for allocation. In a worst-case sum, every contributor enters linearly and with equal weight. A tenth of a millimetre from a bore diameter counts exactly as much as a tenth from a fixture location. The method therefore offers no guidance about *which* tolerance to tighten. Faced with a total that exceeds the requirement and no way to rank the contributors, engineers tighten whatever is cheapest and least disruptive to tighten.

That is a rational response to a method that provides no ranking. It is also how a stack ends up over-specified in exactly the places where tightening was easy.

## Statistical stack-up ranks the contributors

Statistical tolerance analysis combines contributors in quadrature rather than linearly. Under the root-sum-square approach, the variance of the output is the sum of the variances of the contributors, each weighted by the square of its sensitivity coefficient, which is how much the output moves per unit of movement in that input.

The practical consequence is the whole point: **contribution scales with the square of the standard deviation.** A feature with twice the variation of another contributes four times as much to the output variance. In a stack with one dominant contributor and several small ones, tightening the small ones is close to pointless, and this is arithmetically visible rather than a matter of judgement.

This is what produces the pattern in the title. Run a sensitivity analysis on a stack that has never had one and the distribution of contributions is rarely even. A small number of features dominate, several contribute almost nothing, and the ones that dominate are frequently not the ones carrying the tightest limits.

The three over-specified features are usually over-specified because they were easy to tighten: a machined face, a reamed hole, a ground surface, all on operations with capable processes and cheap adjustment. The one loose feature is loose because tightening it was awkward, expensive, or belonged to somebody else.

## The contributor that is not on the drawing

The most common dominant contributor is one that has no tolerance at all, because it is not a feature.

Three recurring examples:

**Fixture repeatability.** The stack accounts for part geometry and says nothing about how repeatably the part is located. If the fixture's clamping sequence produces a two-hundredth of a millimetre of variation in datum contact, that variation propagates into every downstream dimension [referenced from that datum](https://www.asme.org/codes-standards/find-codes-standards/y14-5-dimensioning-tolerancing), and it appears nowhere in the drawing arithmetic.

**Thermal state.** Aluminium has a coefficient of thermal expansion around 23 micrometres per metre per kelvin. A 400 mm feature measured at 20 °C and machined at 28 °C differs by roughly 74 micrometres. On a stack budgeted in tenths of a millimetre, that is not a rounding error. Stacks are typically written as if everything is at the [20 °C reference temperature](https://www.nist.gov/pml/sensor-science/dimensional-metrology); production is not.

**Measurement error.** The gauge is part of the system. If the measurement system's variation is a meaningful fraction of the tolerance band, some of the observed spread is the gauge rather than the process, and tightening the process cannot remove it. This is what a gauge repeatability and reproducibility study is for, and the [NIST/SEMATECH handbook's treatment of measurement process characterisation](https://www.itl.nist.gov/div898/handbook/mpc/mpc.htm) is the standard reference for separating the two.

| Contributor | Varies with | On the drawing? |
| --- | --- | --- |
| Fixture repeatability | Clamping sequence, wear, operator | No |
| Thermal state | Shop temperature, machine warm-up, material | Rarely |
| Measurement error | Instrument, fixturing of the instrument, operator | No |

A stack that omits all three is not describing the assembly. It is describing an idealised drawing of the assembly, at reference temperature, measured perfectly, in a fixture that locates identically every time.

## What redistribution actually looks like

Redistribution is not "tighten everything". It is a reallocation of a fixed budget, and the arithmetic is what makes it affordable.

The sequence:

1. **Measure the contributors as they actually behave.** Not the drawing limits but the observed distributions, sampled across shifts, tool states and material lots. A sample taken from one shift immediately after a tool change describes that shift, not the process.
2. **Compute sensitivity coefficients.** How much does the critical dimension move per unit of movement in each contributor? For a simple linear chain these are ±1. For anything involving angles, offsets or [rotation about a datum](https://www.iso.org/standard/66777.html), they are not, and assuming they are is a common way to get a confident wrong answer.
3. **Rank by contribution to variance.** Standard deviation times sensitivity, squared. This is the list that tells you where the money is.
4. **Reallocate.** Loosen the features contributing negligibly. Spend the released budget on the dominant contributor.
5. **Re-verify against the requirement**, and confirm the new allocation is capable. A tolerance the process cannot hold is not a budget, it is a wish.

Step four is where the objection arrives, and it is worth taking seriously. Loosening a tolerance feels like a reduction in quality even when the arithmetic says the assembly is insensitive to that feature. The counter is that the assembly's behaviour is the thing being controlled, and the assembly does not know what the drawing says. If a feature contributes two percent of the output variance, holding it to a tenth of what it needs is not quality. It is expenditure.

## A worked example

Numbers make the asymmetry concrete. Take a five-contributor stack on a critical dimension with a requirement of ±0.30 mm, all sensitivity coefficients ±1 for simplicity, and measured standard deviations rather than drawing limits:

| Contributor | σ (mm) | σ² | Share of variance |
| --- | ---: | ---: | ---: |
| A: bore diameter | 0.020 | 0.000400 | 4% |
| B: face location | 0.025 | 0.000625 | 6% |
| C: slot width | 0.018 | 0.000324 | 3% |
| D: fixture repeatability | 0.090 | 0.008100 | 78% |
| E: thermal, 8 K swing | 0.030 | 0.000900 | 9% |
| **Total** | | **0.010349** | |

The combined standard deviation is the square root of the total variance, roughly 0.102 mm. At three standard deviations that is ±0.305 mm against a ±0.30 mm requirement. That is marginal, and the line will produce excursions.

Now look at where the variance lives. Contributors A, B and C together account for thirteen percent of it. Halve all three, an expensive programme touching three separate processes, and the combined standard deviation falls to about 0.097 mm. A five percent improvement for three separate process changes.

Halve contributor D alone, and the total variance drops to 0.004274, giving a combined standard deviation near 0.065 mm. That is a thirty-six percent improvement from one fixture rebuild.

The arithmetic is not subtle, and that is the point: it is invisible under worst-case summation, where A, B, C, D and E all look like tolerances of a similar size. Only when each contributor is squared does the fixture separate from the rest of the field. The three "over-specified" features in the title are A, B and C, probably already held tighter than this table implies, because they were the ones easy to tighten in previous rounds.

Note also what happens to E. An eight-kelvin swing contributes more variance than any of the three machined features, and it appears on no drawing. It is not a tolerance; it is the shop floor in August.

## The second-order savings

The direct saving is scrap. Two others are usually larger and less visible.

**Inspection.** This is where a redistribution meets [inspection strategy](/#drawings) directly. A feature held far tighter than the assembly requires is usually inspected at a frequency that matches its stated criticality rather than its actual influence. Re-ranking the contributors also re-ranks the inspection plan. Fewer gauges, better placed, sampled at rates derived from the actual variance structure rather than from convention.

**Process capability headroom.** A feature running at a capability index near the acceptance threshold is a feature that will generate excursions whenever anything drifts. Loosening a non-dominant feature from marginal capability to comfortable capability removes a recurring source of line stoppages and containment activity that never appeared in the scrap number at all, because the parts were caught.

## Where this goes wrong

Three failure modes worth naming, because each produces a redistribution that looks right and fails in production.

**Assuming independence that does not exist.** Root-sum-square assumes contributors are independent. Two features machined in the same setup on the same tool are not independent; they share the tool's wear state and the setup's location error. Treating correlated contributors as independent understates the output variance, and the stack passes on paper.

**Sampling too narrowly.** A distribution estimated from a single shift, a single lot and a fresh tool describes the best case. The redistribution derived from it will be tighter than reality supports, and it will fail when the tool is halfway through its life.

**Redistributing without changing the inspection plan.** If the loosened features remain on the same inspection frequency and the tightened one is not measured any more closely, the plant carries the new risk profile with the old detection profile. The saving is real and the exposure is new.

## The engineering argument, not the software argument

None of this requires a particular tool. It requires that the stack be treated as a derived result rather than an inherited constant, and that the derivation use measured variation rather than drawing limits.

The reason it so rarely happens is not ignorance of the method. It is that re-deriving a stack means measuring a lot of features across a lot of conditions, computing sensitivities for a geometry that may not be documented in a convenient form, and then having a conversation about loosening a tolerance, which is the least popular sentence in a quality review.

That is a workload problem, not a knowledge problem. It is also the specific workload that [reading the line's existing measurement history](/#method) can absorb: the distributions are already in the historian and the CMM exports, and the sensitivity analysis is arithmetic once the geometry is expressed properly. What the exercise returns is a ranked list of where the variance actually lives, which is usually the first time anyone has seen that list for the assembly in question.

The three tight features and the one loose one are not a universal law. They are what a budget looks like when it has been adjusted for a decade by people optimising locally, each of whom made a defensible decision, none of whom had the ranked list.

A last note on sequencing. Re-derive the stack before buying capability, not after. The usual order is reversed: an excursion appears, a capital request follows for a more precise machine on the operation that seems responsible, and the analysis is written afterwards to justify it. Run the ranking first and that request is sometimes cancelled outright, because the operation under suspicion turns out to contribute four percent of the variance and the fixture nobody costed contributes most of the rest.
