The instrument's error, multiplied
Everything so far has been about what a calibration can recover in principle. This is the essay about how well.
The experiment is simple and it is run rather than argued. Take a four-bar built out of true, generate thirty readings of its output angle, add independent errors of a stated size to each, run the whole calibration, and measure how far the recovered shape is from the truth’s. Repeat six times at each of four noise levels three decades apart.
The numbers
reading error error in the shape ratio
1 × 10⁻⁵ 1.900 × 10⁻⁵ 1.900
1 × 10⁻⁴ 1.900 × 10⁻⁴ 1.900
1 × 10⁻³ 1.899 × 10⁻³ 1.899
1 × 10⁻² 1.891 × 10⁻² 1.891
The ratio is the same to three figures across three decades. A least-squares fit of log error against log noise gives a slope of 0.99938.
That is linear, with no threshold below which the noise stops mattering and no saturation above which it stops getting worse, over a range from a hundredth of a milliradian to ten milliradians. The tiny droop in the last row — 1.891 rather than 1.900 — is the first-order model beginning to notice that ten milliradians is not infinitesimal, and it is a per cent.
Where the factor comes from
The bound is one line of linear algebra and it is worth having.
A perturbation δr in the readings produces a perturbation in the parameters of at most ‖δr‖ divided by the smallest singular value of the identification Jacobian, because the smallest singular value is by definition how little the matrix can stretch a unit vector.
On this machine with thirty poses that smallest value, on the identifiable subspace, is 0.4634. So the bound on the amplification is 1/0.4634 = 2.158.
The measurement is 1.90. It sits under the bound, as it must, and not far under — which is the bound doing its job rather than being vacuous. It is under rather than at the bound because the bound is the worst case over all directions the noise could point in, and random noise points in the worst direction only occasionally; the worst of the six repeats at each level runs about 6.1 × 10⁻⁵ against a mean of 1.9 × 10⁻⁵, which is a factor of 3.2 and comfortably outside the bound’s ratio — because the worst of six is a different statistic from the mean and the bound is on neither.
That last point matters and is easy to slide past. The bound is on the worst possible perturbation of a given size, not on the typical one and not on the worst of six draws. What it guarantees is that no reading error of size ε moves the answer by more than 2.158ε; what the measurement reports is what a typical one does.
Six repeats, and why the worst is quoted too
Each point on the curve is the mean of six independent calibrations, and the open marks are the worst of the six. Both are printed because they answer different questions and because the gap between them is a result.
At every noise level the worst run is about 3.2 times the mean. That factor is stable across three decades, which says it is a property of the distribution rather than an accident of one draw — six samples from a distribution whose tail behaves this way will typically have a worst member a few times the mean.
The practical reading is that a single calibration is not a measurement of the amplification, it is one draw from it. A report quoting an accuracy derived from one run is quoting a number that will be exceeded roughly half the time and exceeded threefold occasionally.
Six repeats is few and it is what the figure carries; the mean of six is itself uncertain by about forty per cent of its own scatter. That uncertainty is smaller than the effect being measured — the ratios agree to three figures across three decades — so it does not threaten the conclusion, and it would if the conclusion were about a difference of ten per cent between two pose sets.
The factor belongs to the mechanism
This is the reading worth carrying and it is not obvious.
The amplification is computed from the identification Jacobian, which is built from the mechanism’s geometry and the poses measured. Nothing about the instrument enters it. So 1.90 is a property of this four-bar measured at these thirty positions, and it would be 1.90 for a protractor, a laser, an optical encoder or a person with a set square.
Two consequences. A better instrument buys proportionally and nothing more: halve the reading error and halve the answer’s error, exactly, with the factor unchanged. And a better pose set changes the factor, which is the only lever available that does not cost money.
Pose selection is therefore not a refinement. Eight chosen poses give an observability of 0.2957 against twelve evenly spaced ones’ 0.2935, and the amplification is one over those — so choosing is worth about a third of the measurements, or equivalently a third of the instrument’s price, depending on which is scarcer.
What the residual says and what it does not
At a noise level of 10⁻³ the fit’s final root-mean-square residual is 1.16 × 10⁻³ — the noise, to two figures. That is what a well-posed fit does: it drives the residual down to the level of the data’s own inconsistency and stops.
So the residual is an excellent estimate of the instrument’s error. It is not an estimate of the answer’s error, and reading it as one is the mistake this essay exists to prevent.
The two differ by the amplification. A residual of 1.16 × 10⁻³ radians on this machine means a shape error of about 2.2 × 10⁻³ in the invariants; on a machine with a worse-conditioned pose set the same residual would mean ten times that, with nothing in the fit’s output to say so.
The residual measures the data. The amplification converts it into a statement about the answer, and the amplification is not in the residual. It has to be computed, from the identification Jacobian, and printed alongside.
What is being measured as an error
A word about the vertical axis, because the choice hidden in it is the same one the whole field keeps making.
The error plotted is the distance between the recovered shape and the truth’s shape, measured in Freudenstein’s three invariants. It is not the distance between the recovered lengths and the truth’s lengths, and it could not be: the lengths are recovered only up to a factor, so their distance from the truth’s is dominated by an arbitrary quantity and would measure the damping rather than the noise.
That is not a technicality dodged, it is the reason the plot is clean. Plot the length error instead and the points scatter, because each run’s arbitrary factor lands somewhere slightly different — a spurious variation with nothing to do with the instrument.
Measure the error in the coordinates the measurement determines. The recoverable part of the answer has a well-defined error that is linear in the noise; the unrecoverable part has an error that is whatever the regulariser did. Mixing them into one norm produces a number that is neither.
Where the linearity stops
Two places, both worth knowing.
Near a singularity of the mechanism. The amplification is one over a singular value, and if the pose set puts the machine close to a configuration where the constraint Jacobian is singular then the derivatives themselves blow up, the first-order model stops describing the behaviour, and the error stops being linear in the noise. A pose selection avoids those configurations anyway, because the rows there are short.
And at noise large enough to change the answer’s basin. Everything above assumes the fit converges to the same minimum every time. With enough noise it need not: a badly conditioned direction plus a large perturbation can land the fit somewhere else entirely, and there is no linear relationship between the perturbation and the resulting jump. At the levels measured here that never happens — all twenty-four runs converge to the same shape — and at a hundred times the noise it would.
Between those, three decades of clean linearity is exactly what a first-order theory promises and it is pleasant to see it delivered.
Measuring the amplification rather than bounding it
The bound is one line and the measurement takes twenty-four calibrations, and it is worth saying why both are here.
A bound is a promise about the worst case and it is computed from the matrix alone. It cannot be wrong and it can be loose: on a matrix whose smallest singular value is much smaller than the rest, the bound is dominated by a direction the noise rarely points along, and the typical behaviour is much better than the promise.
The measurement is what actually happens. It runs the whole pipeline — the noise, the fit, the damping, the convergence test — so it includes every effect the bound’s linear algebra leaves out, and it produces a number a report can quote as the expected error rather than as the worst one.
Having both is this site’s habit and it does a specific job here. The measurement must sit under the bound, and if it did not, something in the pipeline would be wrong: a sign error in the Jacobian, a mis-scaled row, a convergence test firing too early. It sits at 88% of the bound, which is close enough to say the bound is tight and far enough to say the two were computed independently.
Why the ratio is so nearly constant
Three figures across three decades is a stronger result than linearity alone and it deserves a sentence.
The fit is a non-linear one, so there is no reason in principle for a perturbation of the data to move the answer by exactly a constant times its size. What makes it so is that the perturbations are small compared with the curvature of the objective — the answer stays in the region where the first-order model is accurate — and in that region the map from data to parameters is a linear map, whose gain is the pseudo-inverse’s norm.
The droop at the largest noise level, 1.891 against 1.900, is the first sign of leaving that region. Extrapolating, the amplification would depart from constancy by a per cent somewhere around ten milliradians of reading error, which is a very poor protractor, and by a lot somewhere around a hundred.
So the linear model is not merely adequate over the range that matters, it is adequate over a range considerably wider than any real instrument occupies. That is the useful practical statement: the amplification can be computed once and applied to any reading precision.
Averaging, and why it is a different lever
The amplification is fixed by the mechanism and the poses. The noise is fixed by the instrument. There is a third thing, and it is the only one that improves without either.
Repeat the readings. With n independent repetitions of the same pose set the effective noise falls as 1/√n, so the answer’s error falls as 1/√n with the amplification unchanged. Four times the measurements for half the error, a hundred times for a tenth.
That is a real lever and a slow one, and putting it beside the other two is the useful comparison. A pose set chosen rather than evenly spaced is worth about a third — which four repeats would also buy. A better instrument by a factor of two is worth a factor of two, which sixteen repeats would buy. Choosing poses is the cheapest of the three and it is the one most often skipped, because it requires arithmetic before the measurement rather than patience during it.
The one error that does not shrink
Everything in this essay is about the random part of the reading error. There is another part and it behaves completely differently.
A systematic error — a protractor with a zero offset, an encoder with a scale factor, a machine that was warm for the first ten readings and cold for the last twenty — does not average away. Repeating the measurement a thousand times reproduces it a thousand times, the residual settles at whatever the systematic error is rather than at the instrument’s random noise, and the answer is displaced by an amount that does not shrink.
The signature is a residual that will not fall to the instrument’s stated precision, and the same signature is produced by a model missing a parameter. Those two causes are hard to tell apart from the residual alone and easy to tell apart by looking at how the residual is distributed over the poses — a subject of its own essay.
The amplification applies to both kinds. A systematic reading error of size ε displaces the answer by up to 2.158ε on this machine, exactly as a random one of that size would, and the displacement stays.
The amplification of one particular parameter
The number above is the worst direction’s. If the report is about one particular parameter, the worst direction is the wrong thing to quote and a better number is available from the same decomposition.
The error in a single parameter is bounded by the length of the corresponding row of the pseudo-inverse, which is a sum over the singular values weighted by how much that parameter appears in each right singular vector. A parameter that lives mostly in the well-conditioned directions comes back much better than 1.90 would suggest; one that lives in the weak direction comes back worse.
On this four-bar read by protractor the three invariants are recovered at 1.4, 1.6 and 1.9 respectively — so the number quoted for the set is the worst of the three and two of them do better. On a coordinate machine’s six-parameter problem the spread is far larger, because the two coupler-point parameters are recovered an order better than the four lengths.
A single amplification for a whole calibration is a summary of a list, exactly as an observability index is, and the same caution applies: the list is short and printing it settles more than any summary of it does.
An error the amplification does not describe
One kind of reading error is not a perturbation of a reading at all, and it behaves in a way none of this covers.
A pose recorded at the wrong crank angle — an encoder off by a count, a stage that did not quite arrive, a transcription slip — is not noise on the output angle. It is a row of the identification Jacobian built at the wrong place, so both the row and its right-hand side are wrong, and the two errors are correlated in a way that is not small.
The effect is a residual that is large at one pose and normal at the others, which is a distinctive signature and is why a residual should be looked at pose by pose rather than as a single root-mean-square. Averaged into one number it looks like slightly worse noise; plotted against pose it looks like exactly what it is.
That is the subject of the next essay, and the reason it follows this one: the amplification converts a residual into an answer’s error only when the residual is what it looks like.
What to print
Three numbers and they cost one decomposition between them.
The final residual, which measures the data. The smallest singular value on the identifiable subspace, which converts it. And their ratio, which is the answer’s error and is the number anybody downstream actually needs.
Reporting the first alone is the near-universal practice and it under-states the answer’s error by the amplification — by 1.90 on a well-conditioned four-bar and by a great deal more on anything harder. Reporting all three costs a line and makes the claim checkable, which is the same argument this site makes about every other number it prints.
What this makes readable
Essays that name this one as a prerequisite.
- Reading a residual Numbers that were measured
About the same objects
Not linked from either essay — found by the objects both name.
- A dimension is a measurement calibration · identification jacobian · least-squares · measurement residual · singular value
- The matrix a calibration inverts calibration · identification jacobian · least-squares · measurement residual · singular value
- A calibration is a synthesis with more equations calibration · identification jacobian · least-squares · measurement residual
- Six things a measurement cannot tell you calibration · identification jacobian · measurement residual · observability index
- The pose the machine cannot reach calibration · identification jacobian · measurement residual · pose selection
- An arm's parameters and its poses calibration · identification jacobian · pose selection
What links here
Essays that link to this one from their own argument.
- Four indices, four answers Numbers that were measured
- How many poses are enough Numbers that were measured
- Two instruments disagree about the worst Numbers that were measured
- What another measurement is worth Numbers that were measured
- Reading a residual Numbers that were measured
- A machine that measures itself Numbers that were measured
- Six per joint is two too many Numbers that were measured
The objects this essay names
Each one links to every other essay that touches it.
CalibrationIdentification jacobianLeast-squaresMeasurement residualNoise amplificationObservability indexPose selectionSingular value