Numbers that were measured

A parameter the model has not got

The machine's tracing point is 0.198 units from where the model says it is, and the model has only four lengths with which to say so. It absorbs the discrepancy: the error over the measured half-turn falls by a factor of thirty-three, the error over the other half falls by twenty-one, and the rocker comes back eight per cent short.

Assumes Reading a residual.

The setup: a four-bar with a tracing point on its coupler, watched by a coordinate machine. The model has four lengths and takes the tracing point to be at (0.45, 0.62) in the coupler’s own frame. The machine has it at (0.49, 0.58), which in the plane is 0.198 units away — about five per cent of the coupler’s length, which is what an unremarkable machining error looks like, and well inside the range the tolerance field treats as ordinary.

The model has no parameter for that. It cannot express it, so it cannot fit it. What it does instead is the subject of this essay.

A calibration that improves the machine and reports the wrong one. The truth's tracing point is 0.198 units from where the model says it is, and the model has only four lengths with which to say so. Fitted over half a turn, it reduces the error there by a factor of 33 and over the other half by a factor of 21, so every practical test says the calibration worked. It got there by moving the rocker by -0.2474 — 8.2% — and the coupler by -0.0317. The residual it cannot drive away, 4.18e-3, is the only signal that anything is missing, and it is the signal a practitioner is most likely to read as instrument noise.
Fig. 1 The error before and after the fit, over the half-turn that was measured and over the half-turn that was not.

What happens

The nominal model, with the drawing’s four lengths and the drawing’s tracing point, predicts the tracer’s position with a root-mean-square error of 0.140 units over the measured range — which is the offset, as it must be, since a rigid displacement of the tracing point is a rigid displacement of everything it predicts.

The fit is started from the nominal model and given the four lengths to move. It converges. The final residual is 4.18 × 10⁻³, which is thirty-three times better than the nominal model and is not zero and never will be.

The lengths it returns:

as built        g 4.00000    a 1.00000    b 3.50000    c 3.00000
identified      g 4.00090    a 1.00021    b 3.46825    c 2.75257

The ground length and the crank have barely moved. The coupler is down by 0.032 and the rocker is down by 0.247 — 8.2 per cent.

The fit is doing its job

Nothing has gone wrong with the arithmetic and it is important to be clear about that.

The fit’s task is to find the parameters that minimise the disagreement between the model’s predictions and the readings. It found them. There is no set of four lengths that does better, and the residual it settled at is the best the model can do.

What has gone wrong is upstream. The model was asserted to be a description of the machine and it is not: the machine has six geometric parameters and the model has four. A fit cannot repair a model, it can only make the best of one, and the best of a model missing a parameter is a set of numbers that partly stands in for the missing one.

The rocker is what moved because the rocker’s column of the identification Jacobian happens to be the one most nearly parallel to the direction the missing parameter would have pushed in. Nothing about the rocker is special otherwise; on a machine with different proportions a different length would take the strain.

And it improves the machine

Here is the part that makes this dangerous rather than merely wrong.

Over the half-turn the readings came from, the nominal model predicts to 0.140 and the fitted model predicts to 0.0042. Thirty-three times better. Over the other half-turn, which the fit never saw, the nominal model predicts to 0.140 and the fitted one to 0.0067 — twenty-one times better, and worse than inside by a factor of 1.6.

So every practical test of the calibration says it worked. The machine is more accurate. It is more accurate where it was measured and also where it was not. A user who calibrated this machine and then used it would find that the calibration helped, substantially, exactly as advertised.

And the rocker bolted to the bench is 3.000 units long, not 2.753. Anybody who takes the calibration’s four numbers as measurements of the machine — and puts them into a stack-up, or an interference check, or a drawing revision — has been handed a number that is out by eight per cent and looks like it came from an instrument.

The improvement is real and repeatable

Before the complaint, the achievement, because it would be easy to read this as saying the calibration is worthless and it is not.

A machine whose tracing point is predicted to 0.140 units is a machine whose predictions are useless for anything needing better than that. After the fit it predicts to 0.0042 inside the measured range, and a second machine of the same design, given the same treatment, improves by the same order. The improvement is not an artefact of the particular numbers and it does not decay: refit the same machine a month later and the same four lengths come back.

So as a device for making a machine predict better, the calibration works exactly as advertised, and a user who wanted that is entirely satisfied. The correction it applies is a real correction to a real error, and the fact that the correction is attributed to the wrong link does not make it a worse correction.

What is worthless is the attribution. The improvement is evidence about the model’s predictions and not about its parameters, and those are separate claims that the same procedure delivers in one package.

Two different things a calibration can be for

That result forces a distinction which is usually left implicit.

A predictive calibration wants a model that reproduces what the machine does. The parameters are a means; their individual values are of no interest; the test is prediction error, and by that test this calibration is a success.

A metrological calibration wants to know what the machine is made of. The parameters are the answer; the test is whether they match the machine, and by that test this calibration is a failure with an eight per cent error in it.

The same procedure produces both and the report does not usually say which was wanted. The residual tests the first and nothing in the fit tests the second, which is the whole of why the distinction has to be made before the measurement rather than after it.

A fit converging onto the machine. The sum of squared residuals through a calibration of a four-bar built 3% long on the coupler and 1% short on the rocker, started from the nominal dimensions. It falls by a factor of 5.1e+28 in 5 steps and then stops at 9.86e-31, which is the solver's own floor. With perfect readings the shape comes back exactly. The last two steps fall faster than the ones before them, which is what a Gauss–Newton descent does when the Jacobian has full rank on the directions it is allowed to move in.
Fig. 2 For contrast, a fit whose model does match its machine: the residual falls to the solver’s floor and keeps going until it cannot.
A calibration that improves the machine and reports the wrong one. The truth's tracing point is 0.099 units from where the model says it is, and the model has only four lengths with which to say so. Fitted over half a turn, it reduces the error there by a factor of 35 and over the other half by a factor of 22, so every practical test says the calibration worked. It got there by moving the rocker by -0.1236 — 4.1% — and the coupler by -0.0184. The residual it cannot drive away, 2.01e-3, is the only signal that anything is missing, and it is the signal a practitioner is most likely to read as instrument noise.
Fig. 3 The same failure at half the offset. The residual halves, the lengths move about half as far, and the improvement factor is much the same.

Why the rocker and not the coupler

The rocker moved by 0.247 and the coupler by 0.032 — a factor of eight — and the asymmetry is not arbitrary.

The missing parameter is a displacement of the tracing point, so its effect on the observed position is a nearly-constant vector in the coupler’s own frame, rotating with the coupler as the machine moves. The fit has to approximate that with a combination of the four lengths’ effects, and it will lean hardest on whichever length’s effect most nearly resembles it.

On this four-bar that is the rocker. Lengthening the rocker moves B outward along the rocker, which carries the coupler with it and displaces the tracing point in a direction that rotates with the mechanism in roughly the way the missing offset does. The ground length’s effect is a translation of a fixed pivot and does not rotate with anything; the crank’s is dominated by a motion of A.

So the answer to which parameter absorbs the missing one is: whichever column of the identification Jacobian is most nearly parallel to the missing column, and that is computable before any measurement. On a machine with different proportions the coupler might take the strain, or the ground length.

There is a use for that. If a particular parameter is the one the report cares about, it is worth checking in advance which unmodelled effects would land on it — and the check is a set of inner products between columns, which costs nothing.

The signal, and how weak it is

There is exactly one signal that anything is wrong and it is the residual that will not fall.

The readings were given no noise at all in this run, so the residual should have gone to 10⁻¹⁶. It settled at 4.2 × 10⁻³, thirteen orders above. On a bench with a coordinate machine good to 10⁻³ that same residual would be indistinguishable from the instrument’s own error, and the calibration would look perfect.

The strength of the signal depends entirely on how good the instrument is, which is an unpleasant property for a diagnostic to have. A better instrument does not make the model better; it makes the model’s inadequacy visible. A worse one hides it.

That is worth turning into an instruction. The residual has to be compared against the instrument’s independently measured repeatability — ten readings at one pose, which takes a minute — rather than against a general sense of what a small number looks like. Without that comparison there is no test at all, and the residual’s distribution over the poses is the other half of the same check.

What the distribution says

Plotted against crank angle, this residual is not flat and not random. It rises and falls smoothly, roughly once per turn, with a shape that follows the coupler’s orientation.

That is the signature: smooth and periodic in the configuration means an unmodelled geometric property, because geometry is what varies smoothly with configuration. Random means noise; constant means an offset; one bad pose means one bad reading.

The shape narrows down the culprit further. A residual with one cycle per turn points at something attached to a link that makes one revolution; the coupler of a crank rocker rocks rather than revolving, and the pattern here follows its swing. That is not a proof and it is a strong hint, and it costs nothing to look at.

The general rule is that a model missing a parameter leaves behind a residual shaped like that parameter’s own column of the identification Jacobian — because that column is exactly the direction the fit could not move in. The leftover is a picture of what is missing.

Adding the parameter

The repair is to put the missing parameter in the model, and here that means fitting six numbers rather than four: the four lengths plus the two that say where the tracing point sits.

Done that way the residual falls to the solver’s floor, all six parameters come back as the machine’s, and there is nothing left to explain. The cost is the conditioning: six parameters on this instrument come back at a condition number of 162 against four at 9.8, so each of them is more sensitive to the instrument’s error.

That trade is the general one. A model with a redundant direction is the failure at the other end, and a model rich enough to express the machine has more parameters, more parameters are worse conditioned, and the extra sensitivity is the price of not being systematically wrong. It is nearly always worth paying, because a random error of a known size is a better thing to have than a systematic error of an unknown one.

The version of the trade that is not worth paying is a model so rich that some of its parameters are unidentifiable, which returns arbitrary numbers along a flat direction and is the failure at the other end of the same axis.

Every length wrong, every reading right. A four-bar was built to the dimensions in the upper bar of each pair and its output angle read at 30 positions. A calibration started from the nominal dimensions returns the lower bar. It reproduces every one of those readings to 1.81e-16 radians and not one of its four numbers is the machine's: they are the machine's multiplied by 0.992289, every one of them, to 3.0e-16. The shape is recovered exactly — the distance in Freudenstein's three invariants is 6.3e-16 — and the size is a free parameter the damping happened to leave near where it started. A machinist handed these numbers would build a machine that works and is not this one.
Fig. 4 And the failure at that other end: a model with a direction the data cannot see, whose residual is perfect and whose fourth number is a starting guess.

Inside and outside the measured range

The two error figures — 0.0042 inside the measured half-turn and 0.0067 outside it — deserve more attention than a factor of 1.6 usually gets.

They say the fitted model extrapolates, which is not obvious and is not always true. A model that had absorbed the missing parameter by contorting itself would predict well where it was fitted and badly elsewhere; this one predicts well in both places, because the contortion it found happens to be a decent approximation to the missing effect over the whole turn rather than only over half of it.

That is luck of a specific kind. The missing parameter here is a rigid offset of the tracing point, whose effect on the observed position is nearly the same at every configuration, and a change in the rocker’s length has an effect that is also fairly uniform. Two nearly-uniform effects are approximable by one another everywhere.

A missing parameter whose effect varies strongly with configuration would not behave like that. The fit would match it over the measured range and diverge outside, and the ratio between inside and outside would be ten rather than 1.6 — which is itself a diagnostic worth having. Measure a few poses outside the range that was fitted and compare. A model that is genuinely right predicts them as well as the ones it was fitted to; one that has absorbed something does not.

That check costs three readings and it is the only test in this essay that distinguishes a fitted model from a correct one using data alone.

The instrument's error, multiplied. The error in the recovered shape against the error in each reading, over four decades, each point the mean of six independent calibrations and the open marks the worst of the six. The slope is 0.9994 — the error is linear in the noise, with no threshold and no saturation — and the constant is 1.90. So a protractor good to a milliradian gives a shape good to about 1.9 milliradians' worth, and the factor belongs to the mechanism and the poses rather than to the instrument. The bound from the smallest singular value is 2.16, which the measurement sits under, as it must.
Fig. 5 And the error this failure is not: the random part, which shrinks when the instrument improves and when the readings are repeated.
One set of lengths, two machines. Three measured input–output pairs, marked, and the linkage Freudenstein's relation returns from them — which is the truth's four lengths to fourteen figures. The relation is a statement about the two angles and it holds on both assembly branches, because it was derived by squaring and that is the step that forgets which one the mechanism is on. So the identified linkage assembled the way the data was taken passes through every reading, to 2.53e-14 radians, and assembled the other way misses them by up to 268° — at the first precision point it reads -111.6° where 98.8° was wanted. That is not a near miss and not a failure either. It is the other answer.
Fig. 6 Nor is it this one: a model whose parameters are exactly right and whose assembly is not.

Which parameters to include

The decision that produces all of this is made before any measurement, and the rule that follows from the two failure modes is short.

Include every parameter the machine plausibly has that the instrument can see. Choosing that list is a decision the matrix can inform and cannot make. A parameter left out is absorbed and produces a systematic error with no signal but the residual. A parameter included that the instrument cannot see is a flat direction and produces an arbitrary number with no signal at all.

Which parameters the instrument can see is computable in advance: build the identification Jacobian with the candidate parameter in it and look at the singular values. If the extra column is well outside the span of the others, include it. If it is nearly inside, including it makes everything worse and leaving it out makes it a systematic error — and that is the genuinely hard case, where neither choice is right and the answer is to change the measurement rather than the model.

A model that is too rich, for comparison

Putting the two failures side by side is the clearest way to see that they are ends of one axis rather than unrelated problems.

Too poor. The machine has six geometric parameters and the model has four. The fit absorbs the difference, the residual sits above the noise, the answer is systematically wrong by an amount the residual hints at and does not measure. Detectable, weakly, and only against an independent estimate of the instrument’s precision.

Too rich. The model has a parameter the instrument cannot see — the scale of a four-bar read by a protractor is the standing example. The fit returns whatever the damping preferred along that direction, the residual is perfect, and the answer contains a number that is not a measurement. Undetectable from the residual at all; detectable instantly from the rank.

In between there is a band of models that express the machine and whose parameters the instrument can distinguish, and finding it is the real work of setting up a calibration. It is also the work that gets skipped, because a model arrives with a machine’s documentation and the question of whether it suits the measurement is rarely asked.

The two diagnostics are different and both are cheap. A residual compared against a measured repeatability catches the first. A rank computed from the identification Jacobian catches the second. Neither catches the other, and reporting one without the other leaves a whole failure mode unchecked.

The one that got away in this run

Worth naming, since the essay’s own example is not the hardest version.

The unmodelled parameter here is a rigid offset of the tracing point, which produces a large and structured residual and is easy to spot once anybody looks. Real unmodelled parameters are often subtler — a joint whose axis is not quite perpendicular, a link with a small bend, a bearing with play — which the practice field models as a short link and which no model here carries. Each of those produces a residual with its own shape, and the shapes are not all as distinctive as this one.

There is also a class that produces almost no residual at all, and it is the one to be afraid of: a missing parameter whose column is nearly parallel to a column already in the model. The fit absorbs it almost perfectly, the residual falls to the noise, every diagnostic passes, and the parameter it was absorbed into is wrong by an amount nothing reveals. That is not detectable from the data, by any method, and the only defence is having thought about the model beforehand.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 13 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCoupler pointIdentification jacobianLeast-squaresMeasurement residualNoise amplificationStructural identifiabilityUnmodelled parameter