The other end of the bracket is not a number
Worth reading first: The error bar that would be needed · The pair that is not a tie.
Pricing a claim by its measurement error solves a data problem by turning the question round. The compiled measurements carry no uncertainties, and inventing them would make every answer a property of the invention — so instead of asking whether a claim survives its errors, it asks how large the errors would have to be for the claim to fail, and reports the answer as a fraction of the spread of the data the claim is drawn from.
That produced five prices. The most fragile claim in the collection is a discordant pair in the ionisation predictor — hexatriene against naphthalene — which needs an error of 3.4 per cent of its set’s range on each of two measurements. The most robust is a slope floor in the ring-strain predictor, which needs 33.8 per cent. In between sit three others: the same ionisation predictor’s slope floor at 10.0 per cent, a discordant pair in ring strain — cyclopentane against cyclohexane — at 16.0 per cent, and the spin-only predictor’s slope floor, vanadium(III) against cobalt(II), at 17.4 per cent.
Every one of those numbers has a √2 in it, and the √2 is an assumption. Two measurements each with standard error σ have a difference with standard error σ√2 only if their errors are independent. The other extreme is easy to name: measurements of one property across a series usually come from one source and share a calibration, so a shared error cancels out of a difference and the difference is better determined than independence allows. Recomputing under perfect correlation should give the other extreme, with the true value between two numbers that can be produced in one pass.
It does not. There is no second number.
The factor, and its pole
Two measurements with standard errors σ and rσ, correlated with coefficient ρ, have a difference whose standard error is σ√(1 + r² − 2ρr). The independent price divides by √2, so the corrected price is the old one times
√2 / √(1 + r² − 2ρr).
At ρ = 0 and r = 1 that is exactly one, which is the independent case and is how this calculation is checked against it. At r = 1 it simplifies to 1/√(1 − ρ), and as ρ approaches one it grows without bound.
At ρ = 1 exactly, with r = 1, the denominator is zero. The difference of two equally precise, perfectly correlated measurements is exact — the shared error subtracts out and nothing is left. No measurement error of any size overturns any claim about such a difference, so the price is infinite. Not large: infinite, and returning Infinity rather than a very large number is the point of the arithmetic rather than a defect in it.
So the hoped-for bracket runs from the independent price to unbounded. That is a true statement about the model and it is not a bracket.
It is worth being clear that the infinity is not an artefact of the one-sigma convention or of any other choice in the price. It is the statement that a difference of two quantities carrying the same error is a quantity carrying no error — which is why chemists measure differences in the first place, and why a series measured on one instrument is more useful than the same series assembled from four papers. The model is behaving correctly; the question was asking it for something it does not have.
And every price moves together
The second thing is sharper than the first, and it turns a negative result into something usable.
Look at what the factor depends on: the correlation and the precision ratio. Not the size of the difference, not the range of the predictor, not which two molecules are being compared, not whether the claim is a discordance or a slope floor. Under any single assumption about the correlation structure, every price in the collection is multiplied by the same number.
Which means the ordering cannot change. At ρ = 0 the five prices are 0.0342, 0.1004, 0.1599, 0.1738 and 0.3382. At ρ = 0.75 they are exactly twice that. At ρ = 0.99 they are exactly ten times that. At ρ = 1 they are all unbounded together. The claim that is most fragile under independence is most fragile under every correlation there is.
So the ranking — which is the thing actually reported, and the thing anyone would use — survives the extreme that could not be computed. What does not survive is the absolute reading of any individual price, and that was never safe anyway.
The rigidity has a one-line reason. The correction factor depends on the correlation and on the ratio of the two precisions, and on nothing about the claim itself, so when every claim draws its pair of measurements from the same kind of source every price is multiplied by the same positive number — and multiplying a list by a positive constant cannot reorder it.
What the numbers mean, then
If the ordering is fixed and the scale is not, the question worth asking is where each claim crosses from fragile to robust, and how much correlation it would take.
Taking one tenth of a predictor’s own range as the line — a choice, and one every number in the column moves with — four of the five claims are already past it under independence. Only the ionisation discordance is not, and it needs a correlation of 0.883 to get there.
That is not an implausible correlation, and it is a good deal less than perfect. Two ionisation energies measured in one laboratory on one instrument against one calibration standard are exactly the kind of pair whose errors are largely shared. The most fragile claim is fragile under an assumption that is likely to be wrong in the direction that would rescue it, and nobody records the number that would say.
Which is a familiar shape of answer: the calculation is fine, the ranking is fine, and the one quantity that would turn a ranking into a verdict is not written down anywhere. Two systems a model cannot tell apart had the same structure — an underdetermination that is a property of what was recorded rather than of what is true.
There is a second reading of the same figure, and it is about the set of claims rather than about any one. Four of five priced claims sitting above a tenth of range under the most pessimistic assumption available is a stronger statement than the independent pricing could make. The three discordant pairs are the most robust claims — 51.6, 43.0 and 16.0 per cent — which reads as a surprise. It is more than a surprise: those three are robust under every correlation assumption too, because the ordering is rigid and they are already at the top of it.
A finite ceiling, and why it does not help
The infinity comes from assuming the two measurements are equally precise. Relax that and the ceiling is finite: at ρ = 1 the factor is √2/|1 − r|, which is a number as soon as r ≠ 1.
Errors differing by a factor of two give ×1.41. By half, ×2.83. By a quarter, ×5.66. By a tenth, ×14.14. And by nothing at all, unbounded.
So a finite upper end exists and it is worthless as a bound, because it is set by a ratio nobody has quoted either and it diverges as that ratio approaches the most natural assumption about it. Two measurements of the same property on two similar molecules from one source are precisely the case where one expects r near one — so the parameter that rescues the bracket is the parameter that is least likely to help.
The honest answer to the question is therefore: the price of a claim is at least the independent one, has no useful upper bound, and the ordering of prices is unaffected by any of it.
One further consequence follows and it changes how such prices should be quoted. A price is meaningless without its assumed correlation, and five prices quoted without one carry the assumption in the arithmetic rather than in the text. A reader told that a claim needs “3.4 per cent” is being told a number that is right only for independent errors and that may be several times larger for the measurements the claim is actually about. The quantity that should have been quoted is the ratio between two prices, since that is the part which is assumption-free.
The overtake that cannot happen
There is one more thing worth ruling out explicitly, because it is the version of the question that would matter most if it were possible.
Suppose the correlation structure differed between predictors — the ionisation energies sharing a calibration and the ring strains not, say. Then the factors would differ and the ordering could change. Could the second most fragile claim overtake the first?
No. The first sits at 0.0342 and the second at 0.1004, a ratio of 0.341, and overtaking would need the first’s price multiplied by 2.93 relative to the second’s — which means the second’s factor exceeding the first’s by that much. But raising a correlation only ever raises a factor, and raising the first claim’s correlation moves it further from the second. Only lowering a factor below one would help, and there is no such factor: the minimum over the whole two-parameter family is exactly one, at zero correlation and equal precision.
So a search for the correlation that would produce the swap should return that there is none, rather than a number outside the range where a bisection would have put one.
The refusal is worth having for a reason beyond tidiness. A bisection over the interval [0, 1] asked for a root that lies outside it returns an endpoint, and an endpoint of exactly 1 in this problem would have read as “perfect correlation would reverse the ranking” — a striking and completely false finding, produced by a search that had no root to find. The same failure appears in a search over a temperature range, where the bracket had to be tested before the bisection was trusted, and it is the same guard.
What was computed, and how
The prices come from the independent pricing unchanged: for each of four predictors, the discordant pair whose reversal needs the smallest error, and the pair whose separation sets the slope floor. Five of those have both a defined pair and a non-zero range. Each price is the difference divided by √2 and by the set’s measured range.
The correction is one closed-form factor, evaluated at seven correlations and five precision ratios. Nothing is fitted and nothing is quoted.
The check requires eight things: that there are at least two priced claims to order; that the ordering is identical at every correlation; that the factor is unbounded at perfect correlation and equal precision, and reported as unbounded; that it rises monotonically with the correlation; that it is exactly one at independence, checked to a part in 10¹²; that any difference in precision restores a finite ceiling; that a ten per cent difference nonetheless gives a ceiling above ten times the independent price, so the ceiling is not a useful bound; and the refusal — that the swap is reported as impossible rather than as a correlation.
Where the model stops
A single correlation coefficient between two errors is a caricature. Real measurements of a series share some sources of error and not others, and the effective correlation depends on which comparison is being made — two molecules measured on the same day are more alike than two measured a decade apart. The five claims here compare molecules that in some cases were certainly not measured together: the pair that is not a tie is a comparison across a whole series, and the assumption of one correlation across such a set is doing work that a real error analysis would not let it do. The one-parameter family here brackets that, and it brackets it with an infinite end.
The prices are also one-sigma statements throughout, chosen deliberately: the answer wanted is an order of magnitude rather than a p-value, and multiplying everything by 1.96 would not change any of the reasoning above by anything except a common factor — which is exactly the point made here about correlation.
And “fragile” here means only that a claim is overturnable by measurement error of a stated size. It says nothing about the model whose predictor produced the claim, which may be wrong in ways no error bar reaches. A robust refutation of monotonicity is a robust statement about two numbers; whether those numbers should have been compared at all is a separate question.
The generalisation
The transferable finding is about what an unknown nuisance parameter can and cannot do to a conclusion.
Without this argument, the correlation is an unknown that might change anything. With it, the correlation is an unknown that changes exactly one thing — the common scale of every price — and cannot change the comparison between prices. That is a much better position to be in than having estimated the correlation, because it did not require estimating it.
The move that gets there is to write the nuisance parameter’s effect in closed form and look at what it multiplies. A factor that multiplies every quantity in a comparison is invisible to the comparison; a factor that multiplies them differently is not. It is the same reasoning that lets a parameter that never finds a value be set aside rather than estimated: what matters is not the parameter but where it appears. Here it turned out to be the first, and the way to find out was one line of algebra rather than a sensitivity study.
The negative half generalises too. When a bracket’s far end comes out infinite, the right response is to say so rather than to pick a large finite value and present the pair as a range. An infinite end is informative — it says the quantity is unbounded above under a stated and quite ordinary assumption — and a range from 0.034 to 100 would have hidden that behind a number somebody chose.
Who found it, and when
The propagation of correlated errors through a difference is textbook and dates from the nineteenth century; the ionisation energies, strain energies and magnetic moments compared here are quoted from standard compilations. The prices, the factor and everything above are new arithmetic, done to answer a question about a convention.
The number worth carrying is not a price. It is that the ranking of prices does not depend on the assumption the prices were computed under.
Still open: correlations per predictor, and the slope floor’s denominator
The obvious open question is a correlation that differs between predictors. Everything above assumes one correlation structure across all the claims, which is what makes the factor common and the ordering rigid. Measurements of ring strain and measurements of ionisation energy come from different literatures and there is no reason their internal correlations should match — so the interesting version assigns a correlation per predictor and asks how far apart two of them would have to be before an ordering changes. The algebra above already says the answer for the top two claims is “no distance at all, because the required factor is below one”, and it does not say that for every pair in the ranking.
The nearer question is the slope floor’s denominator, which is left untouched here. A slope floor is a rise over a run, and the run is computed by the predictor rather than measured, so it carries no experimental error and every price above prices only the numerator. The largest floor is a rise of 0.880 electronvolts over a run of 0.0308 — one per cent of that predictor’s range — so a small change in the model that produces the run changes the floor enormously while no measurement error touches it. Pricing the denominator needs a different currency from the one used here, and inventing it is the work.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- Three chains and ninety orderings ruled out — both name error propagation, measurement uncertainty, model selection, monotonicity, predictor, underdetermination
- A floor on models written in one scale — both name model limit, monotonicity, predictor, underdetermination
- Six of fifteen change verdict — both name closed form, measurement uncertainty, model limit, underdetermination
- A bond order between atoms that do not interact — both name closed form, model limit, underdetermination
- A verdict inside its own error bar — both name closed form, model limit, underdetermination
- The lever that was supposed to be smaller — both name error propagation, measurement uncertainty, model limit
Named objects
A dashed tag is an object no other essay names yet.
Closed formError propagationMeasurement uncertaintyModel limitModel selectionMonotonicityPredictorUnderdetermination