Orbitals

A control that outranked the mechanism

Ruling polarisation out left one candidate, and the closed-shell overlap ranks at 0.857 against the additivity shortfall — which looked like the answer until the control was read. The cation's formal charge, which cannot be a mechanism, ranks at 0.9524. Eight pairs split four and four by charge cannot separate anything, and within a charge group two candidates both rank perfectly.

Worth reading first: The surface a neighbour moves · The radius that was tabulated.

Polarisation, proposed to explain the failure of radius additivity, fails in the sharpest available form: the polarisation correction ranks at −0.9524 against the shortfall it was proposed to explain, so the pairs that need most correction get least.

That leaves one candidate. The shortfall is largest for the small hard pairs; a repulsion between two closed shells can be computed from the overlap of their orbitals; and computing that repulsion for the eight pairs, fitting nothing, would say whether the missing distance is the repulsion’s.

It does not say. The interesting part is why not, and the reason is one that appears from the other direction: the radius that was tabulated is a fitted quantity, and a set assembled to fit one thing is rarely a set that separates two others.

A control that ranked better than the mechanism. Rank correlations against the additivity shortfall, over eight ion pairs. The overlap of the two closed shells ranks at 0.8571 — but the cation's formal charge, which cannot be a mechanism, ranks at 0.9524, so the set is confounded: its eight pairs split four and four by charge and everything else rises with it. Held fixed within a charge group the overlap still ranks at 0.80 — and so does the softness, at -1.00. Four pairs cannot separate two candidates.
Fig. 1 Every candidate ranked against the shortfall, over eight pairs. The top bar is a control that cannot be a mechanism, and it is the longest.

The repulsion, and the number it gives

Each ion’s outermost orbital is built at the effective charge Slater’s rules give it — a 2p for the ten-electron ions and a 3p for the eighteen-electron ones — and the two are overlapped at the measured separation. A closed-shell repulsion is second order in that overlap, and no constant of proportionality is used anywhere, because only the ranking is read.

pair additive radii miss by overlap²
NaF −14.71% 4.79 × 10⁻⁵
NaCl −2.67% 3.56 × 10⁻⁴
KF −3.31% 5.84 × 10⁻⁴
MgO +3.15% 4.51 × 10⁻⁴
KCl +6.62% 3.59 × 10⁻³
CaO +13.22% 4.25 × 10⁻³
MgS +17.04% 1.28 × 10⁻³
CaS +26.40% 1.47 × 10⁻²

The rank correlation is 0.8571 — strong, in the right direction, and exactly what a closed-shell repulsion would produce: the pairs whose shells overlap most are the pairs whose additive radii overestimate most.

Against −0.9524 for polarisation, that looks like a decided question.

The shortfall against the overlap of the two closed shells. How far the sum of two equal-fraction radii misses the measured separation, against the square of the overlap between the two ions' outermost orbitals, on a logarithmic axis. The eight pairs rank together at 0.8571 — the pairs whose shells overlap most are the pairs whose radii overestimate most, which is what a closed-shell repulsion would do. Nothing is fitted: the repulsion is a squared overlap with no constant in front of it, and only the ranking is used.
Fig. 2 The same eight pairs drawn: the shortfall against the overlap squared, on a logarithmic axis. The 1+ pairs are at one end and the 2+ pairs at the other, which is the difficulty.

The control

A test that could never have failed proves nothing, so a candidate that cannot be a mechanism is ranked beside the two that can.

The cation’s formal charge. An integer, one or two, the same for four pairs at a time. It has no business predicting an interionic distance to any accuracy at all.

It ranks at 0.9524 — better than the repulsion.

That is not a fluke and it is not a defect of the instrument. It is the instrument working: eight pairs that split four and four by charge, with charge, ionic size, hardness and overlap all rising together, cannot distinguish between quantities that track any of them. A rank correlation of 0.86 over such a set is worth very little, and the way to find that out is to rank something that could not possibly be responsible.

The −0.9524 for polarisation is subject to exactly the same objection, and it survives it for a reason worth stating: a correlation with the wrong sign cannot be produced by a confound that has the right one. A refutation is robust to confounding in a way a confirmation is not.

Holding the confound fixed

The remedy for a confound is to hold it still. Within each charge group the formal charge is constant and predicts nothing, and the ranking is over four pairs rather than eight.

charge 1+ charge 2+
overlap of the two shells 0.80 0.80
softness of the two ions −1.00 −1.00

The overlap still ranks. So does the softness — perfectly, with the opposite sign.

Four points and two candidates that both track ionic size is not a test. A Spearman coefficient of ±1 on four points is what two monotone quantities give whenever they are monotone in the same underlying thing, and both of these are monotone in how large the ions are: a bigger ion is softer and, at these separations, overlaps more.

So the honest answer to the question is that the data cannot say. Not that the repulsion is refuted, and not that it is confirmed: eight pairs of which four differ only in charge cannot separate a repulsion from a softness.

The circularity that had to be removed first

There is a second objection and it had to be dealt with before any of the above meant anything.

The shortfall is (sum of radii − measured distance) / measured distance. The repulsion was computed as the overlap at the measured distance. So the two halves of the comparison shared an input, and a pair that happens to be closer than its radii suggest is a pair whose orbitals were overlapped at a shorter distance and therefore had a larger overlap — a correlation guaranteed by the arithmetic and about nothing.

A test whose two halves share an input is not a test, and this one shared the most important one. It is the same fault a fitted exponent has when its window is the thing being measured, arriving in a correlation instead of in a slope.

The repair is to compute the overlap at a distance every pair shares — 2.8 Å for all eight, which is nobody’s bond length and is therefore nobody’s advantage. That number is 0.8571 as well, identically. So the circularity was not doing the work, and the finding survives the objection it deserved.

It is worth noticing that the fixed-distance version is the better quantity anyway, and not only because it is not circular: it measures how compact the two ions are, which is a property of the pair, where the measured-distance version mixes that with how close they got.

What would settle it

A set in which hardness and overlap vary independently. That needs pairs where a large ion is hard or a small one soft, and the eight here have no such case — every one of them is a hard cation with a soft anion, differing in degree.

The nearest thing available in this collection’s own data is the isoelectronic series: O²⁻, F⁻, Ne, Na⁺, Mg²⁺, Al³⁺ all have ten electrons and the same screening constant, so they differ in effective charge and in nothing else. Their sizes span a factor of three, their softnesses a factor of eighty, and their overlaps with a common partner would too — but in the same direction, so the series is confounded the same way.

What is needed is a pair of ions of similar size and different hardness, which exists in chemistry and does not exist in eight alkali and alkaline-earth halides and chalcogenides. The set was chosen to test additivity, and a set chosen to test additivity is not a set chosen to separate mechanisms.

Radii at a fixed enclosure do not add up. Eight rock-salt separations, against two ways of building a radius. Shannon's reproduce them because they were fitted to them, which is a check on the arithmetic rather than a result. Radii set so that every ion's surface encloses ninety-nine per cent of its own density miss them by between six per cent short and thirty-seven per cent long — so the convention used here for drawing an orbital is not the convention a crystal uses for spacing two ions.
Fig. 3 Where the shortfall itself comes from: radii set so that every ion encloses the same fraction of its own density, added, against the measured separations. The errors run to more than a tenth in both directions.

What a cannot say is worth

A conclusion of the data cannot say is worth having, and it is worth being clear about what it adds.

It closes a candidate honestly rather than by silence, which is what recording shortfalls is for. Polarisation is ruled out and the repulsion was the pointer; the repulsion has now been computed, found it ranks well, and found that ranking well over this set means almost nothing. Leaving the 0.857 in the record without the 0.9524 beside it would have been a stronger claim than the arithmetic supports.

It produces a reusable instrument. Ranking a quantity that cannot be a mechanism, beside the ones that can, is a check any correlation on a small set should be given — and it costs one line. A predictor’s own ties are the other member of the same family: both are ways of asking what a good-looking number is worth without fitting anything.

And it locates the missing measurement. The obstacle is not the model, the arithmetic or the integral; it is that eight pairs vary along one axis. That is a statement about what to compute next rather than a defeat.

The general shape of the mistake

It is worth naming what went wrong here, because it is a mistake of a kind that a figure-first collection is unusually exposed to.

A small set with a strong correlation makes an excellent picture. Eight points rising together on a logarithmic axis is a convincing figure, the caption writes itself, and every number in it is computed rather than fitted. It would have been drawn, published and believed, and nothing in it would have been false — the correlation really is 0.857 and the repulsion really is what a closed-shell repulsion would look like.

What would have been wrong is the sentence underneath it. The picture supports these two quantities order together; the caption would have said this is the mechanism, and the gap between the two is exactly the confound.

The cheap defence is to rank something absurd. A formal charge takes one line to compute and cannot be a mechanism, and it is longer than the bar it was meant to make look short. Anything that passes that test has earned a little of the reader’s confidence; anything that fails it has to be reported with the failure attached, which is what is done here.

The habit generalises past correlations. Wherever a small set gives a good-looking number, the question is not is the number large but what else would give a number this large, and the way to find out is to compute one.

What is quoted, and what is computed

The whole comparison is a ranking, which is what a predictor’s ties also reduce to when no fitting is allowed.

Two things are quoted per ion: a tabulated radius and the measured interionic separations. Both are measurements, and neither is used in the repulsion — the radii enter the shortfall and the separations enter both, which is the circularity the fixed-distance version removes.

The effective charges are computed from Slater’s rules for each ion’s own configuration, which is why every member of an isoelectronic series comes out with the same screening and a charge differing by exactly one.

The overlaps are integrals of hydrogenic functions at those charges, and the repulsion is their square with no constant. Nothing anywhere here is fitted, which is what makes a rank correlation the only statistic used.

What this cannot say

Nothing here is a crystal. Eight interionic separations are measurements in rock-salt structures, where every ion has six neighbours and the shortfall is a property of the lattice as much as of the pair — and a two-ion overlap is not a six-neighbour repulsion.

Hydrogenic functions are not ionic ones. A real F⁻ orbital is not a hydrogenic 2p at an effective charge, and its tail — which is what an overlap at these separations is made of — is exactly where the difference is largest.

The repulsion’s form is assumed. Second order in the overlap is the standard result and it is taken as given; a different power would give a different ranking, though not a very different one, since the ranking of a positive quantity is unchanged by raising it to a positive power.

Eight is a small number. A rank correlation over eight points has a standard error large enough that 0.857 and 0.9524 are not meaningfully different from one another, which is a second reason not to read the ordering of the two bars as a result. What the control establishes is that a mechanism-free quantity can reach this range at all.

And a rank is not a size. Even had the set been clean, a rank correlation says the two quantities order together and nothing about whether the repulsion is large enough to account for a 26 per cent shortfall. That is a separate calculation and would need the constant deliberately left out here.

A ninety per cent surface with a neighbour beside it. The contour enclosing 90 per cent of a one-electron ion's density at an effective charge of 2.2, drawn with no field as a circle and in a field of 0.04 atomic units as the closed curve. The surface moves out by 10.6 millibohr on the side the field pulls the density towards and in by the same amount on the far side — 0.88 per cent of its own radius. What the sphere encloses does not change to first order; only where the surface is does.
Fig. 4 The candidate already ruled out, drawn: an orbital in a neighbour’s field, with the polarisation the correction was built from. It is real, it is computable, and it ranks the wrong way.
A different fraction for every pair. The enclosed fraction at which two ions' own surfaces would just touch at the measured separation, pair by pair. It runs from 91.26 per cent to 99.42 — so there is no single surface that a crystal spaces its ions by, and the one orbitals are drawn at here is not it.
Fig. 5 The enclosed fraction each ion’s tabulated radius corresponds to, which is where the shortfall comes from: a tabulation fitted to distances is not a tabulation of a density’s contour.

Those two are the same failure approached from opposite directions. The first fits a radius to the answer and then reports how well the answer is predicted; the second computes a radius and finds the correction pointing the wrong way. Only the second can be wrong, which is the only reason it is worth drawing.

The correction that goes the wrong way round. Eight rock-salt pairs. The upper bar is how far the two ninety per cent radii fall short of the measured separation; the lower one is how far the neighbour's field moves the two surfaces towards each other. The two rank at -0.952 — almost perfectly opposed — so the pairs that need the most get the least, and the two that need least are over-corrected by a factor of two and of seven. Polarisation is not what the additivity failure is made of.
Fig. 6 The correction that goes the wrong way round, which is the same confound seen from the other end. Every quantity ranked here is built from a radius defined by an enclosed fraction rather than fitted to a measurement, and that is what makes the comparison possible at all — a fitted radius already contains the answer it is being asked to predict.

What was checked

The additivity shortfall ranks strongly with the overlap of the two closed shells — the finding as it first appeared.

The softness of the two ions ranks with it hardly at all, at 0.0476 over the whole set.

But a formal charge, which cannot be a mechanism, ranks better still — checked as an inequality against the repulsion. It would fail if the set ever stopped being confounded, which is the point of it.

Within each charge group the repulsion still ranks, above 0.5 in both.

And the softness ranks at least as well there, which is the second check and the one that decides the question: four pairs cannot separate two candidates that both track ionic size.

Every pair’s overlap repulsion is positive and its two effective charges are positive, which are the arithmetic checks that the quantities being ranked are the quantities they are named for.

What a clean set would look like

The obstacle is stated above and it is worth writing down precisely what would remove it, because the specification is short.

A set that separates a repulsion from a softness needs pairs in which the two do not vary together. That means, for at least some of them, a large ion that is hard or a small one that is soft — and the periodic table supplies both: a transition-metal cation is far more polarisable than a main-group one of the same radius, and a fluoride is much harder than a chloride of nearly the same charge density.

Two rows would do it. Zn²⁺ and Mg²⁺ differ by a hundredth of an ångström in the tables and by a factor of several in polarisability; their oxides’ interionic distances are both measured. A single pair of that kind is worth more than eight of the kind used here, because it varies along the axis the eight do not.

What that needs is data rather than new calculation: two effective charges, two radii and one distance apiece. Everything else — the overlaps, the ranking, the control — runs unchanged. It is worth saying so plainly, because the data cannot say is a conclusion that invites being left there, and this one has a two-line remedy.

The same shape of remedy appears wherever a small set is asked a question it was not assembled for, and two systems a model cannot tell apart is the general version: what a comparison can establish is decided by what varies in the set, before any arithmetic is done.

The test that would not have caught it

One thing about this failure is worth recording for the next person who quotes a correlation over a small set, because the obvious safeguard would have passed it.

A rank correlation of 0.857 over eight points is significant by the usual standard — the threshold at the conventional level for eight pairs is near 0.74 — so a significance test run on the candidate would have reported a real association and stopped. The control ranks at 0.9524, which is more significant still, on a quantity that cannot be a mechanism.

Significance and confounding are different failures, and only one of them has a routine test. A significance test asks whether a pattern could have arisen from noise; nothing here was noisy. What went wrong was that several quantities rise together across the set, so the pattern is real and does not belong to the candidate that was measured.

The instrument that catches that is not a statistic. It is a second quantity, chosen because it cannot possibly be the cause, run through the identical arithmetic — and the only cost of it is remembering to look.

Still open: a pair of the same size and different hardness

The obvious open question is the measurement located here. A pair of ions of similar size and different hardness would separate the two candidates in one comparison, and the place to look is a transition-metal cation against a main-group one of the same radius — Zn²⁺ against Mg²⁺, say, which differ by a hundredth of an ångström in Shannon’s tables and by a great deal in polarisability. That is two more rows of data and no new calculation at all.

The nearer question is the constant deliberately left out. A rank says the two order together; a size would say whether the repulsion is big enough to matter, and the repulsion’s constant is not free — it is fixed by the same overlap integrals through the standard closed-shell expression, the same expression that prices a noble-gas contact well. Evaluating that expression for the eight pairs, in energy rather than in rank, would say whether a repulsion of the computed size displaces two ions by the tenths of an ångström the shortfall amounts to. That is a quantitative test, and unlike a ranking it cannot be passed by a confound.

And there is a third thing to do with the instrument rather than the question. Every correlation quoted over a small set is exposed to what happened here, and most of them have no control beside them. A rank correlation of −0.9524 is safer than a positive one for the reason given above, but safer is not checked — and running a formal-charge control against each of them is an afternoon that would either confirm several findings or retire them.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationClosed-shell configurationsConventionEffective nuclear chargeIonic radiusLeast-squaresModel limitOverlap integralPolarisabilityProbability densityScreening