Numbers that were measured

Reading a residual

A residual that falls to the instrument's noise and stops means the model is right. A residual that stops above it means something is missing, and which something can be read off how the leftover is distributed over the poses — as a constant, as a pattern in the crank angle, or as one bad reading.

Assumes The instrument's error, multiplied.

A calibration produces two things: a set of parameters and a set of leftovers. Almost all the attention goes to the first and almost all the information is in the second.

The leftover at each pose is the difference between what the fitted model predicts and what was read. Collectively they are the residual, usually reported as one root-mean-square number, and reduced to one number they say much less than they could.

A fit converging onto noise. The sum of squared residuals through a calibration of a four-bar built 3% long on the coupler and 1% short on the rocker, started from the nominal dimensions. It falls by a factor of 1.2e+3 in 19 steps and then stops, at 4.05e-5 — which is the noise, not the machine. The readings were given 1.00e-3 radians of error each, and no fit can go below what its data contains. The last two steps fall faster than the ones before them, which is what a Gauss–Newton descent does when the Jacobian has full rank on the directions it is allowed to move in.
Fig. 1 A residual falling to the noise and stopping there, which is what a fit does when the model is right and the data is honest.

The good case

Readings carrying independent errors of 10⁻³ radians, a model that matches the machine, thirty poses. The fit converges in six steps and the final root-mean-square residual is 1.16 × 10⁻³.

That is the noise, to two figures. A fit drives the residual down to the level of its data’s own inconsistency and stops, and when it stops there, three things are true at once: the model can express the machine, the fit reached its minimum, and the readings are as good as they claim.

The number is also directly useful. It is an estimate of the instrument’s error made without knowing anything about the instrument — better, often, than the instrument’s own specification, because it is measured in situ with everything else that goes wrong included.

What it is not is an estimate of the answer’s error. Those differ by the amplification, which on this machine is 1.90 and has to be computed separately.

The residual that stops too high

Now the interesting case. The fit converges, the steps get small, and the residual settles at 4.2 × 10⁻³ when the readings are good to 10⁻⁵.

Nothing is wrong with the fit. It has reached the best its model can do, which is what it is for. What is wrong is that the model cannot reproduce the machine, and the gap between the two is the residual it cannot drive away.

The number itself says how big the discrepancy is and nothing about what it is. Three quite different causes produce a residual stuck above the noise:

A parameter the model has not got. The machine has a property — a tracing point somewhere the model does not think, an extra offset, a joint that is not where the drawing says — that no combination of the model’s parameters can express.

A systematic error in the instrument. A zero offset, a scale factor, a drift with temperature. Nothing about the machine is unmodelled; the readings are simply not what they claim.

And a mistake in one or two readings. A transcription error, a pose taken at the wrong crank angle, a machine that was knocked between readings.

All three give the same single number. They do not give the same distribution.

Look at it pose by pose

The distribution of the leftover across the poses is where the three separate, and plotting it costs nothing.

A constant offset — every pose’s residual the same sign and about the same size — is an instrument’s zero error, or a datum that is not where the model thinks. It is the easiest to see and the easiest to fix: add the offset as a parameter and refit, and if the residual collapses, that was it.

A smooth pattern in the crank angle — a residual that rises and falls once or twice through the turn — is an unmodelled geometric property, because geometry is what varies smoothly with configuration. This is the signature of a missing parameter, and the shape of the pattern often says which: a residual with one cycle per turn points at something on the crank, two cycles at something on the coupler or rocker.

One or two poses much worse than the rest is a bad reading. Nothing about a mechanism is discontinuous in the crank angle, so a leftover that is thirty times the neighbouring ones did not come from the machine.

And a leftover that looks like nothing — no pattern, no outliers, uniform in magnitude — with a size above the instrument’s stated precision, means the instrument is worse than it claims.

A calibration that improves the machine and reports the wrong one. The truth's tracing point is 0.198 units from where the model says it is, and the model has only four lengths with which to say so. Fitted over half a turn, it reduces the error there by a factor of 33 and over the other half by a factor of 21, so every practical test says the calibration worked. It got there by moving the rocker by -0.2474 — 8.2% — and the coupler by -0.0317. The residual it cannot drive away, 4.18e-3, is the only signal that anything is missing, and it is the signal a practitioner is most likely to read as instrument noise.
Fig. 2 A residual that will not fall: the model is missing a parameter, and the fit is doing the best that four lengths can do about something that is not one of them.
A fit converging onto noise. The sum of squared residuals through a calibration of a four-bar built 3% long on the coupler and 1% short on the rocker, started from the nominal dimensions. It falls by a factor of 1.3e+1 in 23 steps and then stops, at 4.05e-3 — which is the noise, not the machine. The readings were given 1.00e-2 radians of error each, and no fit can go below what its data contains. The last two steps fall faster than the ones before them, which is what a Gauss–Newton descent does when the Jacobian has full rank on the directions it is allowed to move in.
Fig. 3 And the ordinary case at a high noise level, for comparison: the same shape of descent, stopping two orders higher because the data contains less.

The residual that is enormous

The opposite failure and it is more alarming and less dangerous.

Identify a four-bar from three measured angle pairs. The lengths come back as the truth’s to fourteen figures. Assemble the identified machine the other way round — the second branch, which the same lengths permit — and compare it against the readings: it misses them by up to 268°, and at the first pose it reads −111.6° where 98.8° was wanted.

A residual of 268° is not a near miss and not a subtle failure. It is loud, it is obvious, and it is the easiest thing in this field to diagnose: the parameters are right and the assembly is wrong, and one photograph settles which.

The reason it happens at all is that Freudenstein’s relation was derived by squaring, and squaring is the step that forgets which branch a mechanism is on. So the relation is satisfied by both assemblies and the identification cannot prefer either.

It is worth keeping this case in mind precisely because it is so unlike the others. The dangerous residuals are the small ones, and a residual of four and a half radians is a gift.

How small is small enough

The instruction the residual should fall to the noise needs a number attached, because a residual never falls to exactly the noise and deciding when it is close enough is a real judgement.

The honest test is a comparison of two independently obtained quantities. The fit’s root-mean-square residual is one estimate of the reading error. The instrument’s own repeatability — take the same pose ten times and look at the spread — is another, and it is obtained without any model at all. If the first is within a factor of about 1.2 of the second, the model is not contributing anything the readings do not.

On the runs in this field the two agree to two figures at every noise level, because the noise was generated and is therefore known. On a real bench the second number has to be measured, and measuring it is ten readings at one pose and takes a minute.

A calibration without that second number cannot say whether its residual is good. It can only say the residual is small, which is a comparison against nothing. That is the single most common gap in a calibration report and it is the cheapest to close.

A residual is not a scalar

The habit of reducing the leftovers to one root-mean-square is so ingrained that it is worth naming what it discards.

Thirty poses give thirty leftovers. That is thirty numbers, and the summary keeps one. Everything this essay is about — the constant, the pattern, the outlier, the drift with measurement order — lives in the twenty-nine that were thrown away.

The summary is useful for one thing: comparing two fits of the same data. It is nearly useless for diagnosing a single fit, and it is what nearly every fitting routine returns by default.

So the recommendation is not to compute anything extra. It is to keep the vector the fit already produced and look at it, which costs nothing and is skipped because the routine’s return value is a number.

What the fit’s shape says

Beside the final value, the path the residual took is worth reading.

A clean Gauss–Newton descent on a well-posed problem accelerates: two orders in the first step, four in the second, nine in the third. The number of correct digits doubles each step once the answer is close, which is what quadratic convergence looks like — the same behaviour the solver shows on a position solve, and seeing it is evidence that the Jacobian has full rank on the directions the fit is allowed to move in.

A descent that crawls — small improvements for many steps — means the objective has a long shallow valley, which is a poorly conditioned direction. That is a warning about the answer’s reliability rather than about the model’s correctness, and the number it corresponds to is the condition number, which should have been computed before the fit rather than inferred from its behaviour.

A descent that stalls and then jumps means the damping is fighting: the Levenberg term grew because a step made things worse, then shrank when one worked. Occasional, it is normal. Repeated, it usually means the model has a discrete alternative nearby.

Two causes with one signature

Of the causes listed above, two are genuinely hard to separate and it is worth saying so rather than pretending the taxonomy is complete.

A systematic instrument error that varies with configuration — a protractor whose scale is slightly non-linear, an encoder with a once-per-turn eccentricity — produces a smooth pattern in the crank angle, which is exactly the signature attributed above to a missing geometric parameter. Both are smooth, both are periodic in the crank angle, and both are largest where the mechanism is moving fastest.

There is one test that separates them and it needs a second measurement rather than a cleverer plot. Move the instrument. A geometric property of the machine produces the same pattern wherever the instrument is mounted; an eccentricity in the encoder rotates with the encoder. Remount it a quarter turn round and see whether the pattern follows.

That is the general shape of every honest resolution here. A residual’s distribution narrows the candidates to a class, and separating within the class needs an experiment that varies one candidate and not the others. Reading the plot harder does not do it.

An outlier is not automatically wrong

The advice to look for one bad reading is right and the action that usually follows it is not.

Discarding the worst residual and refitting always improves the residual, whether or not the discarded point was bad, because the fit then has one fewer constraint. Doing it twice improves it more. A procedure that drops points until the residual looks acceptable will always terminate and will always produce a good-looking answer, and the answer is worth nothing.

So an outlier is a prompt to investigate a specific reading, not a licence to remove it. The investigation is usually cheap: retake that pose. If the retake agrees with the model, the first was a slip and dropping it is justified by evidence rather than by its inconvenience. If the retake agrees with the first, the machine really does do that there, and the model is what is wrong.

The one case where a reading can be dropped without a retake is when it is impossible rather than merely surprising — a crank angle outside the machine’s travel, an output angle outside the rocker’s swing. Those are refusals rather than errors, and they carry their own information.

Every length wrong, every reading right. A four-bar was built to the dimensions in the upper bar of each pair and its output angle read at 30 positions. A calibration started from the nominal dimensions returns the lower bar. It reproduces every one of those readings to 1.81e-16 radians and not one of its four numbers is the machine's: they are the machine's multiplied by 0.992289, every one of them, to 3.0e-16. The shape is recovered exactly — the distance in Freudenstein's three invariants is 6.3e-16 — and the size is a free parameter the damping happened to leave near where it started. A machinist handed these numbers would build a machine that works and is not this one.
Fig. 4 And the residual’s blind spot, drawn: an answer with a perfect residual, three of whose four numbers are measurements and one of which is not.
The instrument's error, multiplied. The error in the recovered shape against the error in each reading, over four decades, each point the mean of six independent calibrations and the open marks the worst of the six. The slope is 0.9994 — the error is linear in the noise, with no threshold and no saturation — and the constant is 1.90. So a protractor good to a milliradian gives a shape good to about 1.9 milliradians' worth, and the factor belongs to the mechanism and the poses rather than to the instrument. The bound from the smallest singular value is 2.16, which the measurement sits under, as it must.
Fig. 5 What a residual does tell you, converted: the reading error it estimates, multiplied by the amplification, is the answer’s error.

What a residual cannot see

The most important limitation, and it is the reason this essay is not the last word on diagnostics.

A residual cannot see an unidentifiable direction. The scale of a four-bar read by protractor is arbitrary, the fit returns whatever the damping preferred, and the residual is 1.8 × 10⁻¹⁶ — as good as a residual gets. Every diagnostic in this essay passes with distinction on an answer one of whose four numbers means nothing.

That is worth stating flatly because the whole of this essay is about reading residuals and residuals have this blind spot exactly. The instrument that sees a flat direction is the rank of the identification Jacobian, which is not part of a fit at all, and the cost of computing it is one decomposition of a small matrix.

So the pair of diagnostics a calibration needs is: the residual, distributed over the poses, and the singular values. The first says whether the model fits the data. The second says whether the answer is determined. Neither implies the other, and reporting only the first is the ordinary practice.

One set of lengths, two machines. Three measured input–output pairs, marked, and the linkage Freudenstein's relation returns from them — which is the truth's four lengths to fourteen figures. The relation is a statement about the two angles and it holds on both assembly branches, because it was derived by squaring and that is the step that forgets which one the mechanism is on. So the identified linkage assembled the way the data was taken passes through every reading, to 2.53e-14 radians, and assembled the other way misses them by up to 268° — at the first precision point it reads -111.6° where 98.8° was wanted. That is not a near miss and not a failure either. It is the other answer.
Fig. 6 The loudest residual in the field, drawn: the identified lengths are the truth’s, and one of the two ways of assembling them agrees with no reading at all.

The residual of a synthesis is a different animal

A digression that prevents a confusion, because this site already has a quantity called an error that looks like a residual and is not one.

The structural error of a function generator is how far a synthesised linkage departs from the demanded function between its precision points. It is a property of the mechanism — of the fact that a four-bar cannot compute a logarithm exactly — and it is present in a perfectly made machine measured by a perfect instrument.

A calibration’s residual is a property of the model and the data. It is zero when the model matches the machine and the readings are exact, whatever the machine happens to compute.

The two get confused because both are called an error and both are plotted against the input angle. They can even coexist: calibrating a function generator gives a residual that says how well the model describes the machine, while the machine’s structural error says how well the machine approximates the function it was designed for. A small residual and a large structural error is a perfectly-understood machine that does not do what was wanted, and it is a common and entirely coherent state of affairs.

What to do when the residual is fine and the answer is wrong

The situation this essay’s title cannot address, stated so that a reader does not go looking for it here.

If the model is right, the data is honest, the residual falls to the noise, and the answer is still wrong, then the answer is not determined by the data — a flat or nearly-flat direction, or a second exact solution somewhere else. Nothing about the leftovers detects either.

The instruments for those are the singular values and, for the second, running the fit from several different starts. Both are cheap, neither is part of a fitting routine’s output, and both belong in a report for the same reason the residual does: they are checks the answer could fail.

A residual is a test of the model against the data. It is not a test of the answer against the machine. Everything in this essay is about the first, and the second needs different instruments.

An honest report

Four items, and between them they answer every question this essay raises.

The residual’s root-mean-square, which estimates the data’s inconsistency. The residual plotted against pose, which says what kind of inconsistency. The singular values of the identification Jacobian at the answer, which say what was determined. And the amplification, their ratio, which turns the first into a statement about the parameters.

That is one number, one small plot and a short list, and none of it requires a measurement that was not already taken. It is more than most calibrations report and the reason is not laziness — it is that the standard output of a fitting routine is the parameters and a residual, and everything else has to be asked for.

The one that looks like all of them

A closing caution, because there is an arrangement that produces every symptom above at once and none of the diagnoses is right.

A machine that changed during the measurement — a pin that worked loose, a fixture that shifted, a bearing that warmed and grew — is not one machine. The early readings and the late ones come from different mechanisms, and the fit returns something in between with a residual that will not fall, distributed with a pattern that follows the order of measurement rather than the crank angle.

The tell is that ordering, and it is invisible in every plot in this essay because every one of them is against crank angle. Plot the residual against the order the readings were taken in as well, which costs nothing and is the only way that failure ever shows itself.

There is a cheaper insurance against it too and it is worth adopting as a habit: take the poses in a scrambled order rather than in sequence. A machine that drifts then contributes a residual scattered across the crank angle rather than one correlated with it, which prevents a drift from masquerading as a geometric property — the one confusion in this essay that is genuinely hard to resolve after the fact. Scrambling the order costs nothing at the bench and converts a systematic error into a random one, which is the only trade in this field that is free.

And a second habit that costs one reading: repeat the first pose last. If the two agree to the instrument’s repeatability, nothing drifted. If they do not, the difference is the drift, measured, and the whole set can be interpreted knowing its size. That is the cheapest experiment in this field and it settles the one failure mode none of the plots can see.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Branch ambiguityCalibrationIdentification jacobianLeast-squaresMeasurement residualNoise amplificationStructural errorUnmodelled parameter