The error bar that would be needed
Worth reading first: The pair that is not a tie · The curve between two rows.
The pair that is not a tie established two things a pair of measurements can say about a model without fitting anything. A discordant pair — two systems the predictor and the measurement order oppositely — refuses monotonicity outright, not approximately. And a slope floor — the largest ratio of measured difference to predicted difference over all pairs — is a lower bound on how steep the model must be somewhere, surviving any number of fitted parameters.
Both treat a quoted number as exact. The obvious repair is to propagate a stated uncertainty into the pair arithmetic, so that three discordant pairs becomes three discordant pairs, of which two survive their error bars.
The quoted measurements carry no uncertainties, and where molecular orbital theory dissociates is the standing reason to be careful about that: a number that arrives from outside is treated as it arrived, and anything added to it has to be declared. Inventing plausible ones would be worse than having none, because the answer would then be a property of what was invented. So the question is asked the other way round: how large would the error have to be before each claim stopped holding?
Why the uncertainties are not there
It is worth saying plainly why the obvious route is closed, because the closure is a principle rather than an oversight.
Every predictor here divides into a computed half and a quoted half. The predictions come out of a calculation — Hückel eigenvalues, spin-only moments, angle strain from a ring’s geometry — and can be recomputed by anybody at any precision. The measurements are quoted from the literature, and the rule is that a quoted number appears as the source gives it. Strain energies arrive as 115, 110, 26, 0, 26, 40 kilojoules per mole: integers, with no error bar attached, because that is how they are published in the places a reader would find them.
Attaching plausible uncertainties would mean inventing them, and every number downstream would then be a property of the invention. One spectrum, a line of models is the standing warning about exactly that — a quantity that looks measured and is a consequence of a choice made upstream.
The curve between two rows needs the same repair for the same reason, and its closing paragraph said so: a ceiling computed against a quoted number is only as sharp as the number. What follows applies there unchanged.
Pricing a refutation
A discordant pair says the measured difference runs opposite to the predicted difference. It stops saying that when could be zero or the other way. Two measurements each with standard error give a difference with standard error , so the refutation is destroyed at about .
That number is in the measurement’s own units, which makes four predictors in four unit systems incomparable — so it is divided by the range of measurements in that set. The result is a percentage: this refutation needs an error of sixteen per cent of everything the set spans.
Nothing there is quoted and nothing is assumed. It is a property of the claim, and it is the number a reader can hold against whatever they happen to know about how the measurement was made.
The three are the sturdy ones
The three are cyclobutane against cycloheptane, cyclobutane against cyclooctane, and cyclopentane against cyclohexane, on strain energies that span 115 kJ/mol. Their measured differences are 84, 70 and 26 kJ/mol, so they are destroyed by errors of 59, 49 and 18 kJ/mol — 51.6, 43.0 and 16.0 per cent of the range.
These are strain energies quoted to the nearest kilojoule. No uncertainty anybody would defend comes within an order of magnitude of eighteen kilojoules, let alone fifty-nine. All three survive, comfortably, and the expectation that one would fall is wrong in the direction that matters: they are not marginal refutations, they are the most robust claims in the collection.
There is one crumb of information in the quoting itself. Values given as integers are rounded to at worst half a unit, so the rounding alone puts a floor of about half a kilojoule on the uncertainty of each strain energy. That floor is a hundred times smaller than what any of the three refutations would need, which is a weak statement and the only one available without leaving the collection.
The rounding also settles a question the numbers invite. Cyclopentane and cycloheptane are both quoted as 26, and cyclohexane as 0 — round figures that are conventions as much as measurements, since cyclohexane’s strain is defined as the zero everything else is measured from. So one of the three robust refutations rests on a value that is a reference point rather than an independent measurement, and its error is the error of the convention. That does not weaken it — a reference point is exact by construction — but it is worth knowing which of the six numbers is which.
And the fragile one is elsewhere
A different predictor has a single discordant pair, and that one is fragile. Hexatriene against naphthalene, ordered one way by the highest occupied Hückel eigenvalue and the other by the measured ionisation energy, with a measured difference of 0.150 eV. It is destroyed by an error of 0.106 eV on each — 3.4 per cent of that set’s range.
An ionisation energy is not usually known to better than a few hundredths of an electronvolt, and different sources for the same molecule differ by more than that often enough. So this is a refutation that a reader should hold loosely, and it is the only one here that is.
The ordering is the finding, and it is worth saying why it goes that way. A discordance’s price is its measured difference against its set’s range, so a refutation is cheap when the two systems are close in the measured property. The three ring-strain pairs involve rings whose strains differ by tens of kilojoules; the ionisation pair involves two molecules whose ionisation energies differ by a seventh of an electronvolt in a set spanning three. Nothing about the predictors decides it. The original worry was sound and it was pointed at the wrong claims: it counted the three that cannot be moved and did not price the one that can.
The contrast with the ring-strain set is worth holding. There the measurements span 115 kJ/mol and the refutations differ by tens; here they span 3.1 eV and the refutation differs by 0.15. It is not that ionisation energies are measured worse — they are measured far better in absolute terms. It is that the discordance is small relative to the set, and a refutation’s strength is a ratio rather than an absolute difference.
That is the same lesson two systems a model cannot tell apart draws for indistinguishability: whether two things are separable is a question about the separation against the resolution, never about the separation alone.
A floor is a division
The slope floor behaves differently, and the difference is worth more than the discordance result.
A floor is — a measured difference divided by a computed one. Only the numerator carries an experimental error, so an error bar can lower a floor and never raise it. But the interesting sensitivity is in the denominator, which carries no error at all and can be arbitrarily small.
The largest floor of the four is 28.5, and it comes from hexatriene against anthracene: a rise of 0.880 eV over a run of 0.0308. That run is one per cent of the predictor’s own spread — the predictor very nearly ties those two molecules. A floor of 28.5 sounds like a strong statement about how steep a model must be, and it is a modest rise divided by a near-tie.
Compare the ring-strain floor of 3.38, which comes from a run of 32.5 against a range of 115 — a healthy fraction of the set. The two numbers are not comparable, and reporting them side by side without their denominators invites exactly the comparison that should not be made.
A parameter that never finds a value is the case of a quantity whose determination collapses because the data cannot constrain it. A slope floor resting on a run of 0.031 is the mirror image — not a quantity that cannot be determined, but a bound that is enormous precisely because the denominator is nearly nothing. Both are read off the same way: look at what is being divided by.
What no error bar can touch
One claim in the collection is immune to this whole analysis, and it is worth seeing why.
Cyclopentane and cycloheptane have the same measured strain — 26 kJ/mol — and the predictor separates them by 97.14. That is the discordant-pair instrument run backwards: instead of a tie in the predictor bounding what the predictor can be worth, a tie in the measurement shows how far the predictor moves where the property does not move at all.
An error bar on the measurement changes how confidently two numbers can be called different. It cannot make two equal numbers unequal in a way that helps, and it cannot rescue a predictor that varies where the measurement is flat. So this claim has no price: it is not robust, it is unpriceable, and the distinction matters because a table of fragilities with a blank in it should not be read as a claim that survives everything.
What was computed, and how
Four predictors, each with its own systems, its own measured property and its own units. The predictions are computed — Hückel eigenvalues, spin-only moments, angle strain from a ring’s geometry — and the measurements are quoted, which is the standing division and matters here because only the quoted half carries an error.
For each set: every pair, the discordant ones, the steepest ratio, and then the arithmetic above. The whole of it is a few hundred subtractions and runs in under a second, which is the point — the reason this was not done before is that it did not occur to anybody to ask, not that it was expensive.
The refusal is a predictor with nothing to refute. The spin-only moment orders every pair correctly, so it has no monotonicity claim, and its fragility must come back as absent rather than as zero or as infinity — a set with no discordance is not a set whose discordances are indestructible.
What this does not rescue
Pricing a claim says how much error would overturn it. It says nothing about whether the claim was worth making, and the distinction is worth drawing before the habit is recommended.
A refutation that needs half the range to overturn is robust against measurement error and may still be uninteresting — it refutes monotonicity, which is a weak property, and a predictor can fail it while being useful for everything anybody wanted it for. Conversely the fragile ionisation pair, if it survived a proper error analysis, would refute monotonicity for a predictor that is otherwise the best of the four.
So the price is one axis and importance is another. A weight that depends on how it is weighed is the warning against collapsing several axes into one number, and it applies here to the temptation to rank claims by fragility alone.
Where the model stops
The and the one-sigma convention are choices. Two errors might be correlated — measurements from one laboratory, one method, one calibration — in which case a difference is better determined than and every price above is an underestimate. Nothing here can tell, and the numbers should be read as an order of magnitude rather than a threshold.
The measured range is also a crude denominator. It is set by the two extreme systems, so a set with one outlier has an inflated range and its refutations look cheaper than they are. A more careful normalisation would use a spread that is not dominated by the endpoints, and on sets of six systems there is not enough to compute one.
And this prices only what a pair can say. Both claims are pairwise by construction, which is what makes them model-free; a claim resting on all six systems at once — a fitted slope, a correlation — would need a different treatment and would generally be more robust to error, since errors partly cancel across many points.
The generalisation
The habit is one question, asked in place of a different one.
The natural request when a claim rests on measurements is what are the uncertainties? — and when they are not available, the usual outcomes are that the analysis stops or that plausible values are invented. Both are bad: the first discards a claim that may be perfectly robust, and the second produces a result that depends on the invention.
The replacement is: compute the error that would change the answer. It needs nothing but the claim, it is usually a single line of arithmetic, and it converts an unanswerable question into a comparable number. A reader who knows the measurement to five per cent and is shown a claim that needs fifty per cent to overturn has been told everything they need.
It also ranks claims against each other, which quoted uncertainties do not do well across unit systems. Sixteen per cent of a range and three per cent of a range are directly comparable; eighteen kilojoules per mole and a tenth of an electronvolt are not.
A last note on what the prices are for. They are not a test to be passed. A reader who knows a strain energy to two kilojoules learns from “this needs eighteen” that the refutation stands; a reader who knows an ionisation energy to a tenth of an electronvolt learns from “this needs 0.106” that it does not, or only just. The number’s job is to let a reader who has information not given here finish the argument themselves, which is the same job two pictures, one plane gives its two descriptions — put both in front of the reader rather than choosing for them.
Who found it, and when
Asking how large a perturbation would change a conclusion is standard practice under several names — sensitivity analysis, breakdown point, reverse uncertainty — and belongs to statistics rather than to chemistry. The specific arithmetic here is elementary.
What is new here is the measurement on these four sets: which of their model-free claims are cheap to overturn and which are not, and that the ordering is the reverse of what was assumed.
Still open: correlated errors, and the predictor’s own arbitrariness
The obvious open question is the correlated case. Every price above assumes two independent errors, and measurements of one property across a series usually come from one source and share a calibration, which makes differences better determined than the independent case allows. Recomputing every price under perfect correlation gives the other extreme, and the true value lies between two numbers the same arithmetic can produce in one pass — which would turn each price from a point into a bracket.
The nearer question is the denominator. The slope floor’s sensitivity lives in , which is computed and carries no experimental error but does carry the model’s own arbitrariness: a different but equally defensible predictor — a different Hückel parameterisation, a different definition of angle strain — would give a different run and a very different floor. Recomputing the floors under a few such variations would say whether the largest floor is a fact about the molecules or about one choice of predictor, and that is the same question worth asking everywhere else, applied to the one quantity that has so far been treated as exact.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The lever that was supposed to be smaller — both name error propagation, measurement uncertainty
Named objects
A dashed tag is an object no other essay names yet.
Error propagationMeasurement uncertaintyModel selectionMonotonicityPredictor