Numbers that were measured

Four indices, four answers

Five numbers are in use for scoring how well a set of poses determines a mechanism's parameters. They are five different questions about one list of singular values, they rank pose sets differently, and the literature quotes the choice between them as a matter of preference. It is a matter of what the report has to carry.

Assumes What another measurement is worth.

A pose set produces an identification Jacobian, the Jacobian produces a list of singular values, and somebody has to turn that list into one number so two plans can be compared.

There are five numbers in use. They are all built from the same σ₁ ≥ σ₂ ≥ … ≥ σₙ and they are not the same question.

Four indices, four answers. Four of the five observability indices in use, each divided by its own value at 18 poses so their shapes can be compared: their absolute sizes differ by ten orders and a shared axis would draw three flat lines. They are five different questions about one list of singular values — the geometric mean, the smallest alone, the reciprocal condition number, and two normalisations of the smallest — and at eight poses they disagree by a factor of 1.50 about how much of the job is done. A pose set chosen to maximise one is not the set that maximises another, and the literature quotes the choice as a preference.
Fig. 1 Four of the five on the same growing pose set, each divided by its own value at the last count so their shapes can be compared.

The five

O₁, the geometric mean of all the singular values, divided by the square root of the number of rows. It rewards a pose set that is broadly informative and tolerates one weak direction, because a product is dominated by its typical factor rather than by its smallest.

O₂, the reciprocal condition number σₙ/σ₁. It rewards evenness and is indifferent to overall size — a matrix scaled by a thousand has the same O₂ — so it will happily prefer a plan that recovers everything equally badly to one that recovers most things well and one thing less well.

O₃, the smallest singular value alone. It rewards the worst-recovered parameter and ignores the rest entirely. This is the one the pose selection in this field maximises.

O₄, σₙ²/σ₁. A hybrid: the worst direction’s observability weighted down by the best direction’s, which is what the covariance of the worst-recovered combination actually goes as when the problem is scaled.

O₅, σₙ/√σ₁. Another hybrid, differing from O₄ in how strongly the best direction discounts the worst.

Five formulas, one list of numbers, and no argument in the literature about which is right — only a habit of quoting whichever the paper’s author uses.

What they are asking

Strip the formulas and each is a different sentence about the same measurement.

O₃ asks: how badly will the worst parameter come out? It is the bound on the amplification — the error in the parameters is at most the error in the readings divided by σ_min — so it is the number that answers a question about the report’s worst line.

O₂ asks: how unequally are the parameters recovered? That is a question about presentation rather than about accuracy. A plan with O₂ = 1 recovers everything equally, which is pleasant and says nothing about whether any of it is recovered well — and on this machine it is 0.19 and barely moves after eight poses.

The pose-count essay plots the third of these throughout, and O₁ asks a different question: how much information is there in total? It is the volume of the ellipsoid the Jacobian maps a unit ball to, normalised, so it is the closest thing on the list to an overall figure of merit — and it can be large while one direction is nearly invisible, because one small factor in a product of six is easily hidden.

O₄ and O₅ ask O₃’s question with a correction for scale, which matters when comparing pose sets of different sizes or matrices in different units.

Only O₃ and its two variants bound anything. The other two are descriptions, and the split between the two kinds is exactly the split the measurement finds: the bounding ones keep improving and the describing ones stop.

How far apart they are

Normalise each index to its own value at eighteen poses, so all four curves end at one, and compare them at eight.

They read 1.000, 1.000, 0.667 and 0.817. Two of the four say the job is finished at eight poses and two say it is two thirds and four fifths done — a factor of 1.5 between the extremes, on the same eight measurements of the same machine.

The disagreement is systematic rather than noisy, and the split is clean. O₁ and O₂ are not monotone: the geometric mean and the reciprocal condition number reach their final values by about eight poses and then wander in the fourth figure, because both are ratios or products in which the growth of every singular value cancels. O₃ and O₅ rise throughout, because both are built from the smallest singular value alone and that keeps growing as the noise is averaged down.

So two of the four are measuring the shape of the spectrum, which stops changing once the spanning is done, and two are measuring its scale, which never stops changing. They are answering different questions and the answers are 1.000 and 0.667.

For the record, their absolute values at eighteen poses are O₁ = 0.2147, O₂ = 0.1919, O₃ = 0.3595 and O₅ = 0.2627 — four numbers of the same order, which is why the disagreement is easy to miss until they are normalised and plotted.

What another measurement is worth. The smallest singular value of the identification Jacobian — the observability of the worst-recovered parameter — against how many poses were measured. The lower curve takes poses evenly round the crank; the upper one chooses each next pose to make this number as large as it can. Both rise steeply and then flatten: the 12 evenly-spaced poses already give four fifths of what 18 give. Chosen poses reach at 7 what even spacing needs 8 to reach. The flattening is not diminishing returns on accuracy — it is the information in a pose being a direction in a three-dimensional space, which a handful of well-spread poses already spans. Everything after that is averaging noise, which is a different gain and improves as the square root.
Fig. 2 The index this field actually uses, plotted as itself rather than normalised, for the same pose sets.
The order the poses are chosen in. Each next pose chosen to make the worst-recovered parameter as observable as possible, numbered in the order it was taken, on the crank's own dial. The first three land 150° apart at the widest — they have to, since three parameters need three independent rows and nearby poses give nearly the same row — and every one after that bisects a gap. Nothing told the routine to spread them; it maximises a singular value and spreading is what that turns out to mean. The eighth pose is worth 11.6% more observability than the seventh.
Fig. 3 Seven poses chosen to maximise the smallest singular value. A different index chooses a different seven.

They choose different poses

The disagreement matters most where it is least visible, which is inside a selection rule.

A greedy selection maximising O₃ takes the pose whose row is most nearly orthogonal to what is already there, because that is what raises the smallest singular value. A greedy selection maximising O₁ takes the pose with the longest row that is not redundant, because a product grows fastest by adding a large factor. Those are different poses.

On this machine the two selections agree on the first three — three parameters need three directions and both rules find them — and part company at the fourth. The O₁ rule starts taking poses near the middle of the travel where the derivatives are largest; the O₃ rule keeps bisecting gaps.

The resulting plans score differently under each other’s index, which is exactly what one expects and is worth stating anyway: there is no pose set that is best, only a pose set that is best for a stated objective. A paper reporting an optimised pose set without naming its index has reported a plan and not a result.

The list is only four numbers long

Before choosing among summaries it is worth asking why a summary is wanted at all.

A four-parameter problem read by protractor has four singular values, one of which is zero. Three numbers. A six-parameter problem has six. These are not lists that need compressing — they fit on one line, they are exactly what the decomposition produced, and every question the five indices are asked can be answered by looking at them.

protractor, 20 poses    1.975   1.183   0.379   3.9e-15
coordinate machine      16.24   15.99   4.588   1.017   0.397   0.100
both                    11.07   10.79   8.837   6.282   1.835   0.425

A reader given those three rows can see the rank, the conditioning, the spread and which plan dominates which, without being told any index at all. Compression is what created the disagreement, and the object being compressed is small enough not to need it.

The case for a summary is comparison at scale — scoring three hundred candidate pose sets inside a selection loop, where a human is not looking. That is a real case and it is exactly where the objective has to be stated, because nobody will see the list.

Why the smallest one is special

Among the five, O₃ has a property the others do not, and it is the reason this field uses it.

The error in the recovered parameters, in the worst direction, is at most the error in the readings divided by the smallest singular value. That is a bound, it follows from the definition of a singular value, and it can be checked against a measurement.

On this machine with thirty poses the smallest singular value is 0.4634, so the bound on the amplification is 2.158. Repeating the whole calibration on independently noised data at four noise levels three decades apart gives a measured amplification of 1.90, linear in the noise with a fitted slope of 0.9994. The measurement sits under the bound, as it must, and not far under — which is a bound doing its job rather than being vacuous.

None of the other four bounds anything. O₁ is a volume, O₂ is a ratio, and O₄ and O₅ are O₃ with a scale correction that makes them incomparable between problems. A quantity that bounds a measurable thing can be wrong, and being able to be wrong is the whole of why this site prefers it.

Which one to use

The answer depends on what the report carries and it is not a matter of taste.

If the report carries every parameter and none is special, use O₃. It bounds the worst line in the report, and a bound is the only thing on the list that is a promise rather than a description.

If one parameter matters more than the rest, use none of the five. The right objective is that parameter’s own variance, which is the squared length of the corresponding row of the pseudo-inverse, and maximising it selects a different plan again. The five indices exist because they are computable from the singular values alone, which is a convenience rather than a principle.

If plans of different sizes are being compared, use O₄ or O₅. They carry the scale correction that lets a five-pose plan and a twenty-pose plan be put side by side, which O₃ does not.

And O₂ should mostly not be used to choose anything. Being indifferent to overall size, it cannot distinguish a good plan from a uniformly bad one, and its flatness after five poses means it stops discriminating exactly where the interesting comparisons start.

Four indices, four answers. Four of the five observability indices in use, each divided by its own value at 12 poses so their shapes can be compared: their absolute sizes differ by ten orders and a shared axis would draw three flat lines. They are five different questions about one list of singular values — the geometric mean, the smallest alone, the reciprocal condition number, and two normalisations of the smallest — and at eight poses they disagree by a factor of 1.22 about how much of the job is done. A pose set chosen to maximise one is not the set that maximises another, and the literature quotes the choice as a preference.
Fig. 4 The same four over the range where a plan is actually chosen, where the spread between them is at its widest.

Why five exist

Worth a paragraph, because the multiplicity is not an accident and knowing its origin makes the choice easier.

Each index was introduced for a particular machine and a particular purpose. O₁ comes from wanting a single figure of merit for a robot’s whole workspace, where the question is whether a calibration is possible at all and a product over the parameters is the natural summary. O₃ comes from wanting a bound on the error. O₂ comes from numerical analysis, where a condition number is the standard way to say how much an answer can move.

None of them was wrong for its own question. What happened is that they were quoted afterwards as interchangeable measures of observability, which is a word that suggests there is one thing being measured. There is not: there is a list of singular values, and any summary of a list throws something away.

The list is the honest object, and it is small. A four-parameter problem has four numbers, a six-parameter one has six. Printing them costs a line and settles every question the indices are asked, including the ones the indices disagree about.

The one they all get right

For all the disagreement, there is a case where the five agree completely, and noticing it explains why the disagreement has gone unremarked for so long.

They agree about a bad plan. Six poses crowded into fifty degrees of crank travel score badly on every one of the five: the smallest singular value is small, the condition number is large, the geometric mean is small, and both hybrids are small. Any of the five detects a crowded plan immediately.

They also agree about the direction of improvement over most of the range. All four curves rise, all four flatten, and a plan that improves under one usually improves under another. The disagreement is about how much and about where the knee is, which only matters when two reasonable plans are being compared or when a stopping rule is being written.

That is exactly why the multiplicity survived. The index is usually being used to catch a disaster, and every index catches a disaster. It becomes a problem when the index is used to choose, which is what a selection rule does automatically and silently.

A worked disagreement

Take two plans on the site’s four-bar, both of six poses.

Plan A spreads six poses evenly round the turn. Plan B takes four spread poses and adds two more near the middle of the travel, where the derivatives are largest.

Under O₃ — the smallest singular value — Plan A wins, because its rows are more nearly orthogonal and the weakest direction is better covered. Under O₁ — the geometric mean — Plan B wins, because the two extra long rows raise the larger singular values and a product notices.

Both answers are correct about their own question. Plan A recovers the worst parameter better; Plan B recovers the whole set better on average. A report carrying all three parameters wants A; a report quoting an overall accuracy figure wants B.

Nothing about the mechanism decides between them. The decision is made by what the report says, and the index is where that decision gets recorded — usually without anybody noticing they made it.

Six poses that carry three poses' information. The same machine at six poses crowded into 52° of crank travel. Every configuration is distinct and every reading is real, and the six of them together determine the parameters barely better than three would: the rows they contribute to the identification Jacobian are nearly parallel, so the sixth measurement is mostly a repeat of the first. This is the picture behind the flattening of the observability curve, and it is why where a calibration measures matters more than how often.
Fig. 5 The one plan every index agrees about: six real readings whose rows are nearly one row.
What another measurement is worth. The smallest singular value of the identification Jacobian — the observability of the worst-recovered parameter — against how many poses were measured. The lower curve takes poses evenly round the crank; the upper one chooses each next pose to make this number as large as it can. Both rise steeply and then flatten: the 9 evenly-spaced poses already give four fifths of what 14 give. Chosen poses reach at 4 what even spacing needs 5 to reach. The flattening is not diminishing returns on accuracy — it is the information in a pose being a direction in a three-dimensional space, which a handful of well-spread poses already spans. Everything after that is averaging noise, which is a different gain and improves as the square root.
Fig. 6 And the same comparison for a six-parameter problem, where the list of singular values is longer and there is more for the summaries to disagree about.

What none of them sees

Three things, all of which matter more than the choice between the five.

Whether the model is right. Every index is computed from the columns that are in the model. A model missing a parameter the machine has produces a well-conditioned matrix with an excellent score by all five indices, and a fit that returns numbers that are not the machine’s. That failure is invisible here and shows only in the residual.

Whether the poses exist. An index is computed on a plan and a plan is drawn against the nominal machine. Poses the real machine cannot reach drop out, and the index computed on the plan is not the index the measurement achieved.

And a discrete ambiguity. An index is a local object, in the sense the field’s boundary essay makes precise. Two well-separated parameter vectors that both fit the data are not a direction and no singular value sees them. Three linkages tracing one curve score perfectly on all five.

So the indices answer a narrow question well: given that the model is right, the poses are reachable and the answer is unique, how much does the instrument’s error get multiplied. That is a useful question and it is one of four.

The index and the objective are not the same thing

One more distinction, because collapsing it is how the five got treated as interchangeable in the first place.

An index scores a plan that has been made. An objective is what a selection rule maximises while making one. They are usually the same formula, and they need not be: a rule can maximise O₃ while a report quotes O₁, and often does, because the rule was taken from one paper and the reporting convention from another.

When they differ the result is a plan that is optimal for a number nobody prints and is scored by a number nobody optimised. That is not a disaster — the four curves rise together, so a plan good under one is usually decent under another — and it does make the reported figure meaningless as evidence of care. The figure says what the plan achieved on a criterion the plan was not built for.

The fix is one sentence in a method section: poses were selected to maximise X and the resulting set achieves X = x and Y = y. It costs nothing and it turns a number into a claim.

Why this field prints σ

The convention adopted here follows from all of the above and is worth stating so a reader knows what to expect from every other essay.

Wherever a pose set is scored, this field prints the singular values. Where a single number is quoted it is O₃, the smallest on the identifiable subspace, and the condition number is quoted with it. Where a selection rule was run, the objective is named in the same sentence.

The reason is not that O₃ is the best of the five. It is that O₃ bounds something — the amplification, which is what the next essay measures — and a number that bounds something can be checked against the thing it bounds. On this machine and this pose set the bound is 2.16 and the measured amplification is 1.90, which is the bound working.

An index that bounds nothing cannot be checked, and an unfalsifiable score is exactly the sort of number this site exists not to print.

A note on the normalisation in the figure

The first figure divides each index by its own value at the last pose count, and that is a choice with a consequence a reader should know about.

The five indices have absolute sizes that differ by ten orders of magnitude — O₁ is a normalised geometric mean of order 10⁻⁵ on this problem, O₃ is of order 10⁻¹, and O₄ is a ratio of a square to a value and is of order 10⁻³. Plotted together without normalisation, four of them are flat lines against the axis and one is a curve.

Normalising to the final value makes the shapes comparable and destroys the levels, which is the right trade for the question this essay asks and the wrong one for choosing a plan. A reader must not read the crossing points as anything: they are artefacts of where each curve was pinned.

That caution generalises to any figure comparing quantities in different units, and it is the same caution the identification Jacobian’s own rows need when an instrument reads angles and positions at once. A normalisation makes a comparison possible and decides what the comparison means, and it has to be said out loud both times.

The recommendation, stated

Print the singular values. Quote O₃ and the condition number beside them, and say which pose set they were computed on. If a selection rule was used, name its objective.

That is four lines and it forecloses every ambiguity in this essay. The reason it is not standard practice is that a single score is easier to compare, and the reason a single score is a bad idea is that the comparison it makes easy is between plans whose difference the score cannot see.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationIdentifiableIdentification jacobianLeast-squaresNoise amplificationObservability indexPose selectionSingular value