Bonding models

The pair that is not a tie

An exact tie bounds every model of a form with no fitting, and it is available only for predictors that read integers. A near tie bounds the model's derivative instead, and a pair the predictor orders the wrong way round refuses monotonicity outright. Run on four standard predictors, the new instrument ranks them in a different order from the old one — and its worst case is the control the old one could not see.

Worth reading first: Two systems a model cannot tell apart · The model is what is fitted.

Two systems a model cannot tell apart built an instrument out of ties. Where a predictor assigns two systems the same number and their measurements differ, the spread inside the tie is a floor on the error of every model of that form — whatever the parameters are, however many of them there are, and whatever function sits between the predictor and the measurement.

It is sharp, it needs no fit, and it is available for three of the four predictors. That is the same virtue a control that cannot be a mechanism has and for the same reason: a statement that does not depend on the model cannot be argued with by improving the model. It is available for those three because they read integers: a count of unpaired electrons, a pair of counts of bonds and lone pairs, a graph. The fourth — the computed angle strain of a ring — assigns a different number to every ring, has no ties at all, and served there as the control that the instrument must say nothing about.

Almost every predictor in chemistry is of the fourth kind. So the tie instrument names its own extension:

Every tie here is exact because every predictor reads integers or a graph; a continuous predictor has near-ties instead, and the same argument applies with the spread compared against the separation rather than against zero.

It does, and it gives two statements rather than one.

Two instruments, and they do not agree. Each of this collection's four predictors under both tests. The exact-tie instrument reports the share of the variation a predictor demonstrably cannot account for, and it ranks VSEPR worst and has nothing at all to say about the ring-strain control. The near-tie instrument reports how much steeper any model must be somewhere than the set is overall, and it ranks the control worst and cannot speak about VSEPR, whose predictor is a label rather than a number. Neither instrument is the general one, and a predictor that passes one has not been audited.
Fig. 1 The four predictors under both instruments. The worst case of each is invisible to the other.

Two statements from one pair

The slope floor. Take any model y=f(x)y = f(x) whose derivative is bounded by LL. Then for every pair of systems ΔyLΔx|Δy| ≤ L\,|Δx|, so

L    maxpairsΔyΔxL \;\ge\; \max_{\text{pairs}} \frac{|Δy|}{|Δx|}

is a floor on the steepest slope the model must have somewhere. It is a statement about the derivative rather than about the function, which is exactly why it survives any number of fitted parameters: adding a parameter can change where the steep part is and cannot remove it.

The discordant pair. If the predictor and the measurement order two systems oppositely, no monotone model fits both — not approximately, not with a large error, at all. One such pair refuses monotonicity outright, and monotonicity is how a predictor is almost always claimed: more strain, more energy; a higher orbital, a lower ionisation energy.

Neither statement needs a model, a form, or a fit. Both are of the kind a careful fit keeps needing — a bound established before any parameter is chosen, rather than a residual reported after every parameter has been.

Every pair, and the slope each one demands. Every pair of 6 systems, plotted as how far apart the predictor puts them against how far apart the measurement does. A model with a derivative bounded by L lies under the line of slope L through the origin, so the steepest point sets a floor: 3.382 kJ/mol per unit of a ring's geometry, against 0.241 for the set as a whole. Points in the warned colour are pairs the predictor orders the wrong way round, and each of those refuses a monotone model outright rather than bounding it.
Fig. 2 Every pair of six rings, plotted as the separation the predictor gives against the separation the measurement gives. A model with slope at most L lies below the line of slope L, and the steepest point sets the floor.

What it says about the control

The ring-strain predictor is the tie instrument’s control: no ties, nothing said, and the instrument was asked to be silent about it precisely so that a routine reporting a large fault for everything could be caught.

Under the near-tie instrument it is the worst of the four.

predictor slope floor the set’s own slope ratio discordant pairs
the spin-only moment 1.963 1.017 1.93× 0 of 32
the highest occupied Hückel eigenvalue 28.55 3.551 8.04× 1 of 13
the computed angle strain 3.382 0.241 14.04× 3 of 15

Any model of a ring’s measured strain as a function of its computed angle strain must be, somewhere, fourteen times steeper than the whole set is — and three of its fifteen pairs are ordered the wrong way round, so no monotone model of it exists.

How much steeper each model has to be somewhere than it is overall. For each predictor that assigns a number: the steepest slope any model of it must have somewhere, divided by the slope of the whole set. One would be a model that can be a straight line. The control — the ring strain the tie instrument had nothing to say about — is the worst at 14.0, and the spin-only moment, which the exact instrument condemned, is the best at 1.9.
Fig. 3 The ratio for each numeric predictor. The control is the one that fails hardest, and it is the one the exact instrument had nothing to say about.

Silence was not a clean bill of health, and the reason is worth stating precisely: a predictor with no ties has no ties because of a property of the predictor — that it never repeats a value — and not because of anything about how well it works. The tie instrument’s control was there to catch an instrument that faulted everything. It was doing that job. It was not saying the control was good.

The three pairs that refuse a monotone model

The discordant pairs are worth naming, because a bound is abstract and a refutation is not.

pair angle strain differs by measured strain differs by
cyclobutane and cycloheptane +39.6 −84.0
cyclobutane and cyclooctane +141.2 −70.0
cyclopentane and cyclohexane +25.0 −26.0

Cyclobutane has less computed angle strain than cyclooctane and more measured strain, by seventy kilojoules a mole. No function that rises with the predictor can produce that, whatever its shape and however many constants it has.

The pairs the computed angle strain orders the wrong way round. Each of the 6 systems at its predicted and measured values, with a line drawn between every pair the predictor and the measurement disagree about. There are 3 of 15, and one is enough: no model that rises with the predictor can fit both members of a pair whose measurement falls. That is a refutation of monotonicity rather than a bound on an error, and it needs no fitting and no functional form.
Fig. 4 The six rings at their predicted and measured values, with a line drawn between every pair the two orderings disagree about.

That is a much stronger statement than a poor correlation coefficient, and it is a different kind of statement: a correlation is a summary of how well a particular family of models does, and this is a proof that a whole family contains no member at all.

Why it happens is not mysterious and is the reason the essay is worth writing rather than merely the number. The computed angle strain of a planar equilateral ring rises without limit as the ring grows, because a planar ring’s interior angle is fixed by its size and runs away from tetrahedral. A real large ring is not planar; it puckers, and puckering is what buys the angle back. So the predictor is measuring a geometry the molecule does not adopt, and the sign of its error changes at about six.

That is a diagnosis the instrument did not supply and could not have. What the discordant pairs establish is that no monotone model exists; why none exists is chemistry, and it took one sentence about planarity. The instrument’s value is that it makes the question unavoidable — a correlation of 0.7 invites a better fit, and a discordant pair does not.

The instrument’s own limits, stated by the fourth predictor

VSEPR is not in the table above, and the reason is the sharpest thing to say about what the two instruments are for.

A slope needs a scale. VSEPR does not predict a number; it predicts an arrangement. The tie instrument encodes that arrangement as two counts so that ties can be found in it, which is legitimate because a tie asks only whether two values are equal — a question a label can answer. A ratio of a measured difference to a difference of labels is a number with no meaning, and the instrument is declined rather than computed.

So the two instruments have complementary domains rather than nested ones:

  • exact ties need a predictor that repeats itself, and work on labels;
  • near ties need a predictor on a scale, and work when nothing repeats.

Between them they cover the four predictors and neither covers all four. A predictor that passes one has not been audited, which is a sentence worth having, because the natural reading of the tie instrument is that it is general.

The pair that sets each floor. For each predictor, the two systems whose separation demands the steepest slope. A floor is set by one pair, so it is worth naming it: an outlier in the measurement and a near-tie in the predictor make the same demand, and only the second is a statement about the predictor.
Fig. 5 For each predictor, the pair that sets its floor. A floor is set by one pair, so it is worth naming which.

What the ranking disagreement means

Under the exact instrument, VSEPR is the disaster: 92.4 per cent of the variation in the four angles it predicts is inside its own ties. Under the near-tie instrument it cannot be measured, and the ring-strain control is the disaster.

The spin-only moment is the reverse case. The exact instrument condemns it — 22.5 per cent of the moment variation is inside ties, which is a fifth of the effect being unaccounted for — and the near-tie instrument makes it the best behaved of the three, at 1.93× with no discordant pair at all.

Both readings are correct and they are about different failures. A predictor can be blunt without being wrong-headed: the spin-only moment cannot distinguish two ions with the same unpaired count, and everywhere it does distinguish two ions it orders them correctly and by roughly the right amount. The angle strain of a planar ring is the opposite — it separates every ring and gets three of the orderings backwards.

Those are different defects and no single number reports both, which is the general lesson and the reason to run two instruments over the same table rather than one.

There is a third failure neither instrument sees, and it is worth naming so that the pair is not mistaken for a complete set. A predictor can be blunt, correctly ordered and correctly sloped, and still be a report about the analyst — which is what a fitted exponent that belongs to its window is, and what a repulsion that ranks with a shortfall it does not cause is. Ties and near-ties both take the predictor’s values as given and ask what they permit; neither asks where the values came from.

What the floors are, in their own units

The ratios above are dimensionless on purpose, because a slope floor is in the predictor’s own units and two predictors’ floors cannot be compared. The floors themselves are worth reading once, though, because each is a sentence about a particular pair.

The spin-only moment: at least 1.963 μB per unit of predicted moment, set by V³⁺ and Co²⁺, which the formula puts 1.045 apart and the measurements put 2.05 apart. Twice the slope the set has, on a predictor whose whole claim is that it is the moment.

The Hückel eigenvalue: at least 28.55 eV per unit of eigenvalue, set by hexatriene and anthracene — 0.031 apart in the predictor and 0.88 eV apart in the ionisation energy. That is the near-tie case in its purest form: two molecules the predictor almost cannot distinguish, whose measurements differ by nearly an electronvolt. Eight times the set’s own 3.55, and it is a lower bound.

The angle strain: at least 3.382 kJ/mol per kJ/mol, set by cyclobutane and cyclohexane. A predictor in the same units as its measurement whose model must be more than three times steeper than one-to-one.

The second of these is the one that shows why the instrument earns its name. Hexatriene and anthracene are not a tie — the eigenvalues differ in the second decimal — and treating them as one would be a fudge. Treating them as a pair with a stated separation is exact.

What is quoted, and what is computed

Every measurement is quoted and every one of them is a standard measured value: nine magnetic moments, four bond angles, six ionisation energies, six ring strains. Nothing new is claimed about any of them.

Every predictor is computed by the library that owns it — the spin-only moment from a count, the Hückel eigenvalue from a graph, the angle strain from a ring’s geometry.

The slope floor and the discordant count are then arithmetic over pairs: no fitting, no functional form, and no free parameter anywhere. The one number that involves a fit is the set’s own slope, which appears only as a denominator to make the floor dimensionless, and it is a least-squares line through points nobody claims is a line.

What this cannot say

A floor is not an error. A model must be fourteen times steeper somewhere than the set is overall; it does not follow that the model is bad, only that it cannot be gentle. A genuinely steep physical relation would produce the same number honestly.

The floor is set by one pair, and one pair can be a bad measurement. Cyclobutane appears in two of the three discordant pairs and in the steepest one, so a single wrong strain value would move every number in the ring-strain row. The tie instrument has the same exposure and says so; what neither can do is tell an outlier from a refutation.

A slope floor says nothing about which end is steep. The maximum over pairs is one number and the model it constrains could satisfy it anywhere — at the small rings, at the large ones, or in between. Locating the steep part needs an ordering assumption the instrument deliberately does not make, and a reader who reads fourteen times steeper as fourteen times steeper at the small rings has read something that was not established.

And a small set makes a weak floor. Six rings give fifteen pairs, and the maximum of fifteen ratios is a much weaker lower bound than the maximum of fifteen hundred. The floor can only rise as systems are added, so every number here is a floor on a floor.

What was checked

The control is no longer silent, at more than twice the ratio a straight line would need — checked as an inequality, because the point is that the instrument now says something and not that it says a particular number.

The two instruments rank the predictors differently, checked as a comparison between the control’s ratio and the spin-only moment’s rather than as a pair of numbers, so it fails if the disagreement ever goes away.

A predictor that assigns a label is declined rather than measured. That is checked as an explicit refusal: a routine that quietly encoded the label as a number and reported a ratio would produce a plausible table with a meaningless row in it.

And the refusal is a set built as y=2xy = 2x: a slope floor of exactly two, a ratio of exactly one, and no discordant pair. A set built as y=x2y = x^2 over the same predictor has the same ordering and a ratio well above one, which is what shows the instrument is measuring curvature rather than noise.

What the instrument says about a set it cannot fault. The refusal, run in the open. A set built as y = 2x has a slope floor of exactly two, needs to be no steeper anywhere than it is overall, and has no discordant pair — so the instrument reports nothing, which is what it must do. A set built as y = x² has the same predictor and the same ordering, and its floor is 1.42 times its overall slope, because a curve is steeper at one end than on average.
Fig. 6 The refusal, run in the open: a straight set and a bent one, with the same predictor and the same ordering.

Which uncertainty a discordant pair has to survive

Propagating a stated uncertainty is the obvious repair, and for structural measurements it is not sufficient, because the stated uncertainty is not the largest one.

The reason is already measured. A bond length determined from a rotational spectrum, one from electron diffraction and one from a computed potential minimum are three different quantities, differing by amounts of order ten milliångström for a bond to hydrogen — larger than the precision any of the three is quoted to, and systematic rather than random.

So a discordant pair assembled from two techniques carries an offset that no error bar reports. Two angles differing by two degrees, one measured by microwave spectroscopy and one by diffraction, may differ by less than that once both are converted to the same quantity — and a pair whose discordance is smaller than the systematic difference between its sources is not evidence of anything.

That gives the repair a shape rather than just a formula.

Propagate the stated uncertainties, which catches the pairs that are within noise.

And record the technique, which catches the pairs that are within a convention. A pair from two measurements of the same kind is worth considerably more than a pair from two of different kinds, even when the stated precisions are identical — because the first has only random error between it and the verdict, and the second has a definition as well.

The practical form is a column rather than a calculation. A table of pairs that records where each number came from can be sorted into same-technique and mixed-technique pairs, and a verdict that survives on the first set is a verdict; one that rests on the second is a verdict about a convention.

Still open: other relations, floors at a scale, and measurement errors

The obvious open question is the many other relations of the same shape. Four predictors were audited because four were assembled for the tie instrument, and there are a dozen more relations of exactly this shape — an overlap against a bond order, a computed splitting against a measured band, an electronegativity difference against a dipole. Every one of them has a slope floor and a discordant count already implied by numbers that are already computed, and running the sweep would either confirm several of them or turn up new discordances.

There is also an instrument sitting one step further along, which the arithmetic here already reaches. A slope floor over pairs bounds the derivative; a slope floor over pairs at a stated separation bounds the derivative at a scale, and comparing the two says whether a model has to be steep locally or merely somewhere. That is the same distinction an underdetermined force field draws between a constant and a combination, and the arithmetic is a sort rather than a computation.

The nearer question is what a floor is worth when the measurements have errors. Every statement here treats a quoted number as exact, and a discordant pair whose measurements differ by less than their uncertainties is not a refutation of anything. Propagating a stated uncertainty into the pair arithmetic would turn three discordant pairs into three discordant pairs, of which two survive their error bars — and it is the same repair the ceiling on a rate ratio needs, for the same reason: a verdict computed against a number is only as sharp as the number.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationConventionHückel theoryIonisation energyLeast-squaresMagnetic momentModel limitReference stateRing strainUnderdetermination