Numbers that were measured

What another measurement is worth

The observability of a four-bar's worst-recovered parameter goes 0.118, 0.162, 0.188, 0.208 — and then keeps going up by less and less until adding a pose changes the fourth decimal place. The flattening is not diminishing returns on accuracy. It is a space of fixed dimension being filled.

Assumes Where a calibration should measure.

Take poses evenly round the crank and watch the observability of the worst-recovered parameter as the count goes up.

3   0.1183      8   0.2398     13   0.3055     18   0.3595
4   0.1619      9   0.2543     14   0.3171     19   0.3694
5   0.1884     10   0.2680     15   0.3282     20   0.3790
6   0.2077     11   0.2811     16   0.3390
7   0.2244     12   0.2935     17   0.3494

Three poses to six doubles it. Ten poses to twenty adds forty per cent. Whatever is happening in the first half of that table is not what is happening in the second.

What another measurement is worth. The smallest singular value of the identification Jacobian — the observability of the worst-recovered parameter — against how many poses were measured. The lower curve takes poses evenly round the crank; the upper one chooses each next pose to make this number as large as it can. Both rise steeply and then flatten: the 13 evenly-spaced poses already give four fifths of what 20 give. Chosen poses reach at 4 what even spacing needs 5 to reach. The flattening is not diminishing returns on accuracy — it is the information in a pose being a direction in a three-dimensional space, which a handful of well-spread poses already spans. Everything after that is averaging noise, which is a different gain and improves as the square root.
Fig. 1 The two halves of the same curve, and the chosen set above it.

Two different gains

The rise splits into two things that happen to add up, and separating them is the whole essay.

Spanning. Three parameters need three independent directions in the row space, and the first few poses are what supply them. A pose whose row lies outside the span of the rows already taken raises the smallest singular value a great deal; one whose row lies inside it does not.

Averaging. Once the directions are spanned, another row in a direction already covered still adds weight there. With n independent readings each carrying error ε, the parameter error falls as ε/√n. That is a real gain and it is a slow one.

The first is over quickly and the second never ends. What the table shows is the first finishing somewhere around six poses and the second taking over: from ten to twenty the observability rises by a factor of 1.41, which is √2 to three figures, which is what averaging looks like when nothing else is happening.

The tell is the condition number

There is a cleaner way to see the split than staring at the increments, and it is to watch a quantity that averaging cannot change.

The condition number over the recovered directions is 6.5 at three poses, 5.5 at four, 5.2 at five, and 5.2 at every count from six to twenty. Flat, to the digits printed.

That is because averaging scales the whole matrix rather than reshaping it. Adding a row in a direction already spanned makes every singular value grow together, so their ratio does not move. The condition number stops changing exactly when the spanning finishes, and it stops changing at five poses on this machine — which is the answer to how many poses does the shape of this problem need with no reference to noise at all.

So the two gains are separable by measurement rather than by argument: watch σ_min for the total and κ for the spanning half.

Why the dimension is what matters

The number of poses at which the flattening happens is set by the number of parameters and by nothing else about the machine.

A row of the identification Jacobian is a vector in as many dimensions as there are parameters. Three parameters means three-dimensional rows, and a handful of well-spread three-dimensional vectors span the space. There is no way for a fourth direction to be needed, because there is no fourth direction.

This is why the flattening is not “diminishing returns” in the usual sense. Diminishing returns is a smooth economic statement about effort against reward. This is a dimension being filled, which is an abrupt and structural thing, and it happens at a count that can be computed before the first reading is taken.

For six parameters the knee moves out. For a serial arm’s thirty it moves out further. Always at a small multiple of the parameter count, never at the hundreds a routine procedure asks for, and always independent of how big the machine is or how long its travel.

What another measurement is worth. The smallest singular value of the identification Jacobian — the observability of the worst-recovered parameter — against how many poses were measured. The lower curve takes poses evenly round the crank; the upper one chooses each next pose to make this number as large as it can. Both rise steeply and then flatten: the 12 evenly-spaced poses already give four fifths of what 18 give. Chosen poses reach at 4 what even spacing needs 5 to reach. The flattening is not diminishing returns on accuracy — it is the information in a pose being a direction in a three-dimensional space, which a handful of well-spread poses already spans. Everything after that is averaging noise, which is a different gain and improves as the square root.
Fig. 2 Six parameters instead of three, same machine, same rule: the same shape with the knee further out.
The order the poses are chosen in. Each next pose chosen to make the worst-recovered parameter as observable as possible, numbered in the order it was taken, on the crank's own dial. The first three land 150° apart at the widest — they have to, since three parameters need three independent rows and nearby poses give nearly the same row — and every one after that bisects a gap. Nothing told the routine to spread them; it maximises a singular value and spreading is what that turns out to mean. The eighth pose is worth 11.4% more observability than the seventh.
Fig. 3 The six poses that do the spanning, in the order they earn their place.

Chosen poses reach the plateau sooner

The upper curve is the same measurement with the poses chosen instead of taken in turn, and its shape is the same with the knee moved left.

At eight chosen poses the observability is 0.2957. At twelve evenly-spaced ones, 0.2935. The chosen set reaches at eight what even spacing reaches at twelve, which is a third fewer measurements for the same answer.

The two curves meet at twenty, where both read 0.3790 and the pool is saturated: twenty poses out of thirty-six candidates covers the turn whatever order they were picked in. The gap is widest around eight to ten poses, which is where a calibration actually lives.

There is a reading of that which is more useful than the numbers. Choosing does not raise the plateau — it cannot, since the plateau is set by the pool and the mechanism — it reaches the plateau with fewer measurements. A chosen set is not a better answer, it is the same answer sooner.

What the marginal pose is worth

Turn the table into differences and the practical question answers itself.

The fourth pose is worth 0.0435 of observability. The eighth is worth 0.0154. The fifteenth is worth 0.0111, the twentieth 0.0096. In relative terms: the eighth pose adds 6.9%, the fifteenth 3.5%, the twentieth 2.6%.

At some point that stops being worth the trouble of moving the machine, and where that point is depends on what a measurement costs. What is worth knowing is that the curve is shallow rather than flat — the increments keep coming and they keep shrinking, so there is no natural stopping point in the data. The stopping point comes from outside.

The one thing the increments never do is go negative. Another reading never makes the answer worse, which is worth saying because it is not true of everything in a calibration: another parameter can make the answer much worse, and another instrument can too if its rows are in units nobody scaled.

The first pose is worth infinity

A pedantic point that turns out to be the useful way to see the whole curve.

With no poses at all the observability is zero and the parameters are entirely undetermined. With three it is 0.1183 and they are determined. Nothing on the curve between the second pose and the third is a gradual improvement — it is the difference between an answer and no answer, and putting it on the same axis as the difference between the nineteenth and twentieth pose flatters the latter enormously.

So the curve has three regions rather than two. Below the parameter count, nothing. From there to the knee, the spanning, where each pose is worth a large and roughly equal amount. Past the knee, averaging, where each pose is worth a shrinking amount that never reaches zero.

Drawing all three on one axis is honest and it is misleading, in the way that any plot with a structural discontinuity in it is misleading. The three regions want three different decisions: whether the plan is viable at all, whether it is well conditioned, and whether it has enough repeats. Only the third is a question about how long to keep measuring.

Repeats and poses are different operations

Two ways to spend an hour of measuring, and they do different things to the same matrix.

Measure twenty distinct poses once. Twenty rows, well spread, conditioning 5.2, each row carrying one reading’s worth of noise.

Measure five distinct poses four times each. Twenty rows in five directions, conditioning 5.2 — the same, since the spanning was already done at five — and each direction carrying a quarter of the noise variance.

The second plan has the same conditioning as the first and half the noise on each of its directions. It is better, on this machine, for a report that carries all three parameters. It is not better in general: on a machine whose spanning is not finished at five, the first plan is spending its rows on directions the second never visits.

So the decision rule is the one the condition number gives. Spend poses until κ stops moving, then spend repeats. That is a rule with a measurement attached rather than a habit, and it takes about a second of arithmetic to apply.

Six poses that carry three poses' information. The same machine at six poses crowded into 52° of crank travel. Every configuration is distinct and every reading is real, and the six of them together determine the parameters barely better than three would: the rows they contribute to the identification Jacobian are nearly parallel, so the sixth measurement is mostly a repeat of the first. This is the picture behind the flattening of the observability curve, and it is why where a calibration measures matters more than how often.
Fig. 4 And the plan that spends its measurements on neither: six distinct poses whose rows are nearly one row, which buys no spanning and no averaging.

The plateau is not a limit on accuracy

A reader could take the flattening as bad news — that a calibration hits a ceiling and cannot be improved past it. It is not, and the distinction is worth drawing carefully.

What flattens is the observability, which is a property of the pose set and the mechanism. What does not flatten is the accuracy, which is the observability divided into the instrument’s error and keeps improving as long as measurements keep being taken. There is no ceiling on accuracy from this direction at all.

The ceiling that does exist is somewhere else entirely: it is the model. Averaging drives the random part of the error to nothing and leaves the systematic part untouched, and the systematic part is whatever the model gets wrong about the machine. A calibration with ten thousand readings of a model that is missing a parameter converges to a wrong answer with a very small standard error, which is the failure worth being afraid of and is not what the plateau is about.

So the two halves of the curve have two different endings. The spanning half ends and stays ended. The averaging half continues until the systematic error dominates, and where that is depends on the model rather than on the pose count.

The one that is worth nothing

The failure worth naming is the crowded plan, because it is what happens when a measurement is convenient rather than designed.

Six poses inside fifty degrees of crank travel are six genuinely different configurations and six real readings. Their rows are nearly parallel, so they contribute one direction rather than six; and they are distinct rows rather than repeats, so they do not average either. The plan buys neither gain and takes as long as a plan that buys both.

It is easy to arrive at. A machine that is awkward to reposition gets measured wherever it happens to be; a fixture that only fits one way constrains the crank to a range; a technician told to take six readings takes six readings. Nothing about the result looks wrong — the fit converges, the residual is at the noise, four numbers come back.

The signal that it happened is in the matrix and nowhere else: a condition number far above the machine’s plateau value. On this four-bar the plateau is 5.2 and a crowded six-pose set reads in the tens. That number should be printed beside every calibration result, and it costs one decomposition.

The same curve on a machine that rocks

Everything above is measured on a crank rocker whose crank turns all the way round, so all thirty-six candidate poses exist. A machine that rocks is the more instructive case.

A double rocker at 4, 3.2, 1.4, 3.0 reaches seven of the twenty-four positions in this field’s standard sweep. Its curve therefore stops at seven, and the seven poses it does reach are confined to an arc rather than spread round a circle.

The instinct is that the plateau it reaches must be lower. Measured, it is higher: those seven poses give σ₃ = 0.649 and a condition number of 2.77, against the crank rocker’s twenty-four at 0.415 and 5.21.

Two things follow. The spanning finishes just as fast, because three directions are three directions and seven rows are more than enough. And a short travel is not the same as a bad pose set — what decides the plateau is how much the mechanism’s derivatives vary over the poses available, and a double rocker’s vary a great deal over its arc.

What is genuinely bad is crowding, which is a different thing: six poses inside 0.9 radians of the crank rocker give σ₃ = 0.033 against a spread six’s 0.207, and a condition number of 47.2 against 5.22. That is the case a plan has to avoid, and it is about how fast the rows turn rather than about how far the machine goes.

Reading the curve backwards

The curve is usually read left to right — how much is gained by measuring more. Read right to left it answers a question that comes up more often: how much is lost by measuring fewer.

Dropping from twenty poses to twelve costs 23% of the observability. Dropping to eight costs 37%. Dropping to five costs half. Those are the numbers to have when a measurement session is being cut short, and they are considerably less alarming than the instinct that says a shortened calibration is a spoiled one.

Halving the observability doubles the amplification, so a plan cut from twenty poses to five turns a shape good to two milliradians into one good to four. Whether that matters is a question about the report, and it is a question with an answer rather than a reason for anxiety.

The place where the reading backwards goes badly wrong is below the parameter count, and there it goes wrong completely rather than by a factor. Two poses for three parameters is not half as good as four; it is an under-determined system that returns a confident answer, and the curve has nothing to say about it because the quantity it plots is zero.

Four indices, four answers. Four of the five observability indices in use, each divided by its own value at 16 poses so their shapes can be compared: their absolute sizes differ by ten orders and a shared axis would draw three flat lines. They are five different questions about one list of singular values — the geometric mean, the smallest alone, the reciprocal condition number, and two normalisations of the smallest — and at eight poses they disagree by a factor of 1.41 about how much of the job is done. A pose set chosen to maximise one is not the set that maximises another, and the literature quotes the choice as a preference.
Fig. 5 And a warning about the curve itself: four ways of scoring the same growing pose set, each normalised to its own final value, disagreeing about where the knee is.
The order the poses are chosen in. Each next pose chosen to make the worst-recovered parameter as observable as possible, numbered in the order it was taken, on the crank's own dial. The first three land 150° apart at the widest — they have to, since three parameters need three independent rows and nearby poses give nearly the same row — and every one after that bisects a gap. Nothing told the routine to spread them; it maximises a singular value and spreading is what that turns out to mean. The eighth pose is worth 4.9% more observability than the seventh.
Fig. 6 Ten chosen poses, which is past the knee: the last few are bisecting gaps that were already small.

Where the numbers in the table come from

The table at the top is a measurement rather than a formula, and it is worth saying how it was made, because a reader who wanted to make it for their own machine can.

For each pose count n, take n crank angles evenly spaced through the turn starting at 0.3 radians. Build the identification Jacobian at the nominal dimensions — one row per pose, four columns, each entry the derivative of the output angle with respect to one length. Decompose. Record the third singular value, which is the smallest on the identifiable subspace, and the ratio of the first to the third.

That is the whole procedure and it needs no readings at all. The curve is available before the machine is touched, from the drawing, because the identification Jacobian depends on the parameters and not on the measurements. It is a property of the mechanism and the pose plan, which is exactly what makes it a design tool rather than a diagnostic.

The one approximation in it is that the derivatives are taken at the nominal parameters rather than at the machine’s, and the machine’s are what is being looked for. That is circular in principle and harmless in practice: a machine out of true by a per cent has singular values out by a per cent, so a plan promising 0.2957 delivers between 0.293 and 0.298.

Which curve is being plotted

One caution about all of the above, and it is the reason the next essay exists.

Every number here is the smallest singular value over the recovered directions. That is one of five quantities in use for scoring a pose set, and the five do not agree — normalised to their own final values and compared at eight poses, they differ by a factor of 1.5 about how much of the job has been done — and two of the four have already reached their final value there while two are at two thirds of it.

So the knee’s position is not quite a property of the mechanism. It is a property of the mechanism and of which number is being watched, and a reader who takes the shape of this curve as the shape of “how much a calibration has learned” is taking one opinion for a measurement.

The shape is robust — all five rise steeply and then flatten — and the position of the knee is not. Four of the five are drawn together so the disagreement can be seen rather than described.

What the plateau does not settle

The plateau says how many poses. It says nothing about how good the answer is, and the two are easy to run together.

An observability of 0.3790 is a number in whatever units the rows are in. What a report needs is how far the instrument’s error travels into the parameters, and that is a ratio: the amplification, which on this machine and this pose set is 1.90 and is bounded by 1/σ_min = 2.16. Reaching the plateau means the amplification has stopped improving, not that it is small.

Nor does the plateau say anything about whether the model is right. A calibration on twenty well-chosen poses of a model missing a parameter reaches its plateau, converges, and returns numbers that are not the machine’s. The conditioning is a statement about the columns that are there.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationIdentifiableIdentification jacobianLeast-squaresNoise amplificationObservability indexPose selectionSingular value