The ranking moved and the headline did not
Worth reading first: The other end of the bracket is not a number · The error bar that would be needed.
The other end of the bracket is not a number set out to bracket a quantity and found the upper end of the bracket did not exist. It was pricing refutations — asking, for each refuted claim, what fraction of a measurement’s error would have to be wrong before the refutation collapsed, which is the currency an error bar that would be needed introduced — and the question was what correlated measurement errors do to those prices.
The algebra gave a factor. With correlation ρ and a second measurement whose error is r times the first’s, a difference’s standard error is σ√(1 + r² − 2ρr), so every price is multiplied by
which is 1 at ρ = 0 with equal precision and unbounded at ρ = 1. There is no upper end.
It also gave something sharper than the bracket it was looking for. That factor depends on nothing about the claim — not its size, not its predictor, not what it is about. So one correlation applied to every claim multiplies every price by the same number, and an ordering multiplied throughout by a constant is the same ordering. Correlation slides the ranking rigidly; it cannot scramble it.
And then it named the assumption that result rests on, which is that there is one correlation.
Why one correlation is the wrong model
Ring strains and ionisation energies are not measured by the same people with the same instruments, and the predictors that turn them into claims are as different — a Hückel eigenvalue and a computed angle strain share no calibration and no technique. The correlation between two errors in a difference comes from whatever they share — a calibration, a standard, a systematic in a technique — and two literatures share nothing in particular.
So the honest model gives each predictor its own correlation, and the question becomes how far apart two of them must be before an ordering changes.
That question has an exact answer rather than a sweep, and the reason is the same algebra. Two claims swap when their prices cross. Each price is its own fraction times its own predictor’s factor, so the crossing condition is
— a ratio of factors, which is precisely what a per-predictor correlation supplies and a common one cannot. Holding one predictor at independence and solving for the other’s correlation gives a single number per pair: with equal precision f(ρ) = 1/√(1−ρ), so the required correlation is 1 − 1/need², where the need is the ratio of the two fractions.
No sweeping, no bracketing. Each adjacent pair in the ranking has a price of admission, in correlation, and it can be written down.
One pair cannot be bought at all
Before any number, there is a structural answer, and it turns out to cover the part of the ranking that matters most.
Two of this collection’s claims come from the same predictor — the highest occupied Hückel eigenvalue — one priced by discordance and one by a slope floor. They are the two most fragile claims in the collection, ranked first and second.
Because they share a predictor, they share its correlation, and therefore its factor. Their prices are 0.0342 f and 0.1004 f for whatever f that predictor’s correlation produces, and the ratio between them is 2.933 at every correlation there is. The factor cancels out of a comparison between two claims that carry it.
So the collection’s headline — this claim is the most fragile one it has — is immune to a per-predictor correlation structure, and immune for a reason that has nothing to do with how large its margin is. Any number could be substituted for either fraction and the pair would still be locked.
That is a stronger kind of safety than the common-factor result could offer, and it is worth distinguishing from the kind that comes from a comfortable margin. A margin can be eroded by a better measurement. A cancellation cannot.
And the rest can be bought cheaply
The other three adjacent pairs cross predictors, and their prices are ρ = 0.606, 0.154 and 0.736.
The middle one is the problem. A correlation of 0.154 is weak. It is the sort of correlation two quantities measured in the same laboratory with the same standard have by default, and nothing in this collection would notice it. At that value the ring-strain discordance claim and the spin-only slope-floor claim change places in the ordering.
So the answer to the question is that a per-predictor structure does what a common one could not — it scrambles the ordering — and it does so at a cost that is not exotic.
What a swap is, and what it is not
It is worth being precise about what changes at ρ = 0.154, because “the ordering is not robust” reads more alarmingly than the thing it describes.
Nothing about any refutation changes. Every claim in this collection that was refuted stays refuted at every correlation on this page; a correlated error makes a difference more precisely known, not less, so correlation only ever makes a refutation harder to overturn. The prices all rise together and none of them reaches one.
What changes is which claim is listed as the most vulnerable, and that is a statement about the self-assessment rather than about any chemistry. The ranking is kept to know where to spend effort — which refutation would repay a better measurement — and a reordering means that effort would go somewhere else.
So the finding is about a research-priority list, not about a set of conclusions. That makes it less alarming and not less useful: a priority list nobody can rely on the middle of is worth knowing about before the priorities are acted on, and this one has a locked top and a movable middle.
The three that can be bought, read individually
The three purchasable swaps are not equally interesting, and the prices say which is which.
At ρ = 0.606, the ionisation slope floor falls past the ring-strain discordance. That is a substantial correlation and the two claims differ in price by a factor of 1.59, so it is a genuine question about which of two comparably fragile claims is more so.
At ρ = 0.154, the ring-strain discordance falls past the spin-only slope floor. The two prices differ by a factor of 1.087 — under nine per cent — which is why almost nothing is needed to swap them. These two are effectively tied already, and calling either “more fragile” was reading a difference the numbers barely support. That is the same shape as the pair that is not a tie, arriving from the other side: there the question was whether two predictions were close enough to be indistinguishable, and here it is whether two prices are.
At ρ = 0.736, the spin-only slope floor falls past the ring-strain slope floor, which differ by a factor of 1.95. Expensive, and rightly so.
Read that way the ranking has a shape the list does not show: a locked pair at the top, then a near-tie in the middle that no correlation argument can resolve because it was never resolved by the numbers, then two genuine separations. The 0.154 is not evidence that correlation is dangerous — it is evidence that two entries were never distinguishable and the ordering was reporting a coin flip as a rank.
The top and the middle are different objects
Putting the two results side by side gives the finding its useful form.
Reordering the easiest adjacent pair needs ρ = 0.154. Displacing the collection’s most fragile claim outright — having some claim from a different predictor overtake it — needs ρ = 0.954, because the nearest such challenger is nearly five times less fragile and the factor curve is steep only very near one.
A factor of six separates those. They are answers to different questions, and an ordering quoted as a single list invites them to be read as one.
What was computed, and how
Nothing was measured and nothing was swept. The claims and their fractions are the ones already priced, computed from the same quoted measurements and the same predictors. The factor is the closed form above. Everything here is the solution of one equation per adjacent pair.
That is worth saying because the obvious way to answer this question is a two-dimensional sweep over both correlations, reading off where the ordering changes. It would have produced a grid, a boundary drawn through it, and a number with the grid’s resolution in it. The algebra produces the number exactly, and it also produces the structural result — that same-predictor pairs are locked — which a sweep would have shown as a suspiciously flat region and left to be interpreted.
Seven checks. Three are about the structure: that there is more than one predictor, that some adjacent pair shares one, and that sharing is what locks it rather than anything about the sizes. One says the first pair in the ranking is such a pair, which is the finding about the headline. Two are about the numbers: that the remaining pairs can be swapped, and that the easiest needs only a weak correlation. And the last says displacing the top claim needs a correlation near one, which is the contrast the essay turns on.
The check that the easiest swap is below a quarter is worth pointing at, because the natural expectation runs the other way. The expectation going in is that reordering would need a substantial correlation and the ranking would prove robust; a claim of ρ > 0.3 fails at 0.154, and the conclusion is the opposite of the expected one.
Where the model stops
Equal precision is assumed throughout — the ratio r is held at one — which is what makes the factor 1/√(1−ρ) and the inversion a single line. Sweeping r as well shows that any r ≠ 1 restores a finite ceiling at ρ = 1. With unequal precisions the required correlations here would all change, and the structural result would not: a shared factor cancels whatever it is.
The correlations are treated as free parameters with no evidence behind them. Nothing here knows what the correlation between two ring-strain measurements actually is, and it is not estimated. That is the same posture a parameter that never finds a value had to adopt, and for the same reason: the quantity is real, nothing here measures it, and a threshold is what can honestly be reported instead. What it produces is a threshold — the value at which something changes — which is the useful shape when the quantity is unmeasured, and is the same shape the error bar that would be needed uses.
And “one correlation per predictor” is itself a simplification. A predictor’s correlation with a measurement could differ from claim to claim within it, and then even the locked pair unlocks. That is a more elaborate model with more free parameters than there are claims, which is why it is not the one used — but the locking result is exactly as strong as the assumption that a predictor has one correlation, and no stronger.
The generalisation
The transferable point is about when a common factor is a friend.
A quantity that multiplies every item in a comparison equally is invisible to the comparison, and that is usually reported as a limitation — the ordering cannot recover the scale. Here it is the opposite: because the factor is common within a predictor, the ordering among that predictor’s own claims is exactly the thing that survives when everything else is uncertain.
So the useful question about a nuisance parameter is not “does it cancel” but “over what set does it cancel”, and the answer partitions the comparison into a part that is safe and a part that is not. This collection’s ordering has one locked pair and three purchasable ones, and knowing which is which is more informative than any statement about the ordering as a whole.
The second point is about how such an ordering should be published. Not as a list, which implies every adjacent relation has the same standing. The honest form carries the price of each swap beside it, so a reader can see that the first relation costs infinity and the third costs a sixth — and can decide for themselves whether their own belief about the correlations puts the ranking’s middle in doubt.
Why the threshold and not a distribution
The natural objection to a threshold is that it answers a question nobody asked. A reader wants to know whether the ordering is wrong, and this essay reports the correlation at which it would be — which is a conditional whose antecedent is unmeasured.
The alternative is to put a distribution on the correlations and report a probability that the ordering holds. That would produce a single number and it would be worse, for a reason that has come up before: the distribution would be invented. Nothing here measures inter-laboratory correlation, so any prior over it is a guess, and a probability computed from a guessed prior is a guess wearing a decimal point. A weight that depends on how it is weighed is the standing example of a number whose apparent precision comes entirely from an arbitrary choice upstream of it.
A threshold has the opposite property. It is a fact about the claims — arithmetic on their prices — and it contains no assumption at all about what the correlations are. A reader who believes inter-laboratory correlations are typically under a tenth reads off that the ordering holds; a reader who thinks two per cent is more like it reads off that it holds comfortably; a reader who has actually measured one for their own field substitutes it. All three get an answer, and none of them inherits a belief from this essay.
That is the general argument for reporting the crossing rather than the verdict whenever the deciding quantity is unmeasured, and it applies well beyond this case. The cost is that the result is a sentence rather than a number, and the sentence has an “if” in it. The benefit is that the “if” is visible instead of buried in a prior.
Who found it, and when
The propagation of correlated errors through a difference is textbook. The pricing of refutations by what fraction of a measurement error would overturn them is built on discordant pairs, the rigid-sliding result comes from the common-factor analysis, and the per-predictor algebra is new here.
The common-factor analysis deserves the credit for stating its assumption in a form that could be attacked. It proved a result about one correlation and wrote down, in the same breath, that there was no reason to believe there was one — which is what made this a solvable question rather than a vague worry about robustness.
Still open: the denominator, and the ranking’s block structure
The obvious open question is the denominator, not touched here. A slope floor is a rise over a run, the run is computed by a predictor rather than measured, and every price here prices only the numerator. The largest floor is a rise of 0.880 electronvolts over a run of 0.0308 — one per cent of that predictor’s range — so a small change in the model producing the run changes the floor enormously while no measurement error touches it at all. Pricing that needs a different currency from the one used here, and inventing it is the work.
The nearer question is whether the locking result generalises to the whole ranking rather than to one adjacent pair. Every pair of claims sharing a predictor is locked, not just neighbouring ones, so the ranking decomposes into blocks that are internally rigid and mutually purchasable. Working out that decomposition — which subsets of the ordering are fixed by structure and which are only fixed by the current numbers — would turn one observation into a description of the whole ranking’s shape, and it is arithmetic over a list already in hand.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- A verdict inside its own error bar — both name approximation, closed form, model limit, underdetermination
- One number decides which way it breaks — both name approximation, closed form, model limit, underdetermination
- Six of fifteen change verdict — both name closed form, measurement uncertainty, model limit, underdetermination
- The lever that was supposed to be smaller — both name approximation, error propagation, measurement uncertainty, model limit
- The product a curve measures — both name approximation, closed form, model limit, underdetermination
- The residue is below its own noise — both name approximation, closed form, model limit, underdetermination
Named objects
A dashed tag is an object no other essay names yet.
ApproximationClosed formCorrelationError propagationMeasurement uncertaintyModel limitPredictorUnderdetermination