A gap that was a choice of bonds
Worth reading first: A dispute is a small difference · Four tables and one molecule to disagree about.
Why do four tables of electronegativity, all claiming to measure the same thing, disagree about the direction of some bonds and not others? The natural expectation is that the answer is about awkward elements — hydrogen especially, which is where every disputed sign had turned up, and which the table that could not have mattered had already shown to be the element the scales handle worst.
It was not. Sorting the collection’s thirty-one bonds by the size of the electronegativity difference, normalised to each table’s own range, put every disputed bond at or below 0.0507 and every agreed one at or above 0.0584, with nothing in between. The elements did not enter at all. A dispute was simply a small difference.
Thirty-one bonds is enough to see a gap and not enough to know it is real, and the four tables cover eighteen elements, which is a hundred and fifty-three pairs, of which the thirty-one are a fifth. Running all of them would either populate the interval, giving the boundary a width worth quoting, or leave it empty, which on a hundred and fifty cases would be a much stronger claim.
It populates it, thoroughly.
What survives and what does not
Two things happen at once when the other hundred and twenty-two pairs arrive, and separating them is the whole of this essay.
The gap closes. On all 153 pairs the largest disputed difference is O–Br at 0.1628, and the smallest agreed one is Na–K at 0.0398 — so the disputed group’s top now sits four times above the agreed group’s bottom, and twenty-seven agreed pairs lie below the largest disputed one. There is no number that separates them.
The line does not move. The threshold that misclassifies fewest of the 153 is 0.05065, which is the same line, to five significant figures, that the thirty-one-bond set drew. It was not a fluke of position; it was a fluke of emptiness.
Drawing the misclassification count against the threshold makes the difference visible. The curve has a clear minimum in the right place, and its minimum is six rather than zero. A boundary with a gap would touch the axis; this one has a flat bottom sitting six pairs above it.
Six pairs, and both kinds of failure
Six exceptions out of a hundred and fifty-three is few enough to name, and naming them says what the rule is actually about.
Two agreed pairs sit below the line: Na–K at 0.0398 and Li–K at 0.0487. Both are alkali pairs, and the four tables are nearly identical about the alkalis — the spread across tables for Na–K is 0.0225, the smallest in the whole set. A difference can be tiny and still be definite when nobody disagrees about it.
Four disputed pairs sit above the line: Be–Si at 0.0678, Be–B at 0.0947, O–Cl at 0.1044 and O–Br at 0.1628. Every one involves beryllium or oxygen, and every one has a large spread — 0.12, 0.25, 0.31, 0.27. A difference can be substantial and still be disputed when the tables disagree with each other by more than the difference they are being asked to resolve.
So the mechanism is not the size of the difference alone. It is the size of the difference compared with the disagreement, and the threshold works on thirty-one bonds because there those two things are not distinguishable.
The sixteen disputed pairs make that concrete. Eight of them have differences under 0.03 — C–S at 0.00064, N–Cl at 0.0077, Li–Na at 0.0089, S–I at 0.0114, C–I at 0.0121, B–Si at 0.0270, H–I at 0.0283 — and those are disputed for the reason already given: the difference is small enough that the four tables land on either side of it. The other eight are disputed for the second reason, and they run up to 0.1628. The two populations look alike in the classification and not at all alike in the arithmetic.
There is a further detail worth having, because it constrains any repair. Li–Na is disputed at a difference of 0.0089 with a spread of only 0.0326, and Na–K is agreed at a difference of 0.0398 with a spread of 0.0225. Those two alkali pairs are alike in every respect a rule could see — same group, comparable tiny differences, comparable tiny spreads — and they fall on opposite sides of the classification. Whatever separates disputed from agreed, it is not going to be a smooth function of these two numbers alone.
The relative version does not rescue it
That mechanism suggests an obvious repair: divide. Take the difference and measure it against the spread across the four tables, and the two groups should separate on the ratio even though they do not separate on the difference.
They do not. The steepest disputed pair sits at 0.5950 of the way to the diagonal and the shallowest agreed one at 0.3330, so the groups overlap on that measure as thoroughly as on the other. Placing every pair by its spread and its difference gives a cloud with the disputed points low and to the right and the agreed points everywhere, and no line through the origin cuts them apart.
One thing in that picture is exact, and it is arithmetic rather than chemistry. If four numbers straddle zero then their mean cannot exceed their range, so every disputed pair must lie below the diagonal. All sixteen do, the steepest at 0.595. That is a tripwire rather than a finding — a disputed pair above the diagonal would mean an error in the reading — and it is checked for exactly that reason.
And it is not one awkward element either
The original hypothesis, that the disagreement is about hydrogen, was already refuted on thirty-one bonds. On a hundred and fifty-three it is refuted much more thoroughly.
Fifteen of the eighteen elements appear in at least one disputed pair. Only fluorine, phosphorus and potassium never do, and fluorine’s absence is the least surprising fact in the set — it is at one end of every table, so its differences are large with everything. Hydrogen appears in three of the sixteen disputed pairs, which is the same as carbon, sulphur, iodine and beryllium, and fewer than a hydrogen-centred account would predict.
So the answer to the original question stands: the disagreement is about the arithmetic and not about the chemistry. What has changed is that it is a tendency with named exceptions rather than a law.
What happened to the interhalogens
The other finding was about the series proposed as a test. The six interhalogens — Cl–F, Br–F, I–F, Br–Cl, I–Cl and I–Br — are all agreed, and against the threshold measured here they sit at 6.44, 7.59, 9.86, 1.15, 3.42 and 2.27 times it. Five of the six were never candidates for a dispute; only Br–Cl is anywhere near, and it clears.
That survives without change, and the new threshold does not move it: Br–Cl at 0.0584 sits above 0.05065 and is agreed, so it is one of the 147 the rule gets right. It remains the closest agreed pair to the line in the whole set of 153, which is a slightly stronger version of the earlier statement about it — on thirty-one bonds it was the closest of thirty-one, and it is still the closest of a hundred and fifty-three. What the larger set adds is context. On thirty-one bonds it could be said that the interhalogens were a poor test because they had large differences; it can now be said what a good test would have looked like, which is a set of pairs drawn from the region where the rule is uncertain. There are six such pairs, they are named above, and not one of them is a bond.
That is the same point as the selection effect, arriving from the other side. The series was chosen for having measured dipoles, which is a property of being a real diatomic molecule, and being a real molecule is what keeps a pair out of the interesting region.
Why the small set looked so clean
The obvious worry about a gap that closes when the sample grows is that it was a coincidence, and it is worth saying why it was not.
Six disputed bonds being the six smallest of thirty-one, purely by chance, has a probability of one in seven hundred and thirty-six thousand — which is to say that on its own terms the thirty-one-bond finding was real and would have survived any test of significance anybody put to it. That is not a fluke. It is a selection: the thirty-one bonds are chemically real bonds that other essays had reason to compute. Real bonds are not a random sample of element pairs. They are drawn from a chemistry that puts particular elements together, and the pairs that never occur as bonds — Be–Si, Na–K, Li–K, Be–B — are exactly where the exceptions turn out to live. Every electronegativity table is calibrated on bonded chemistry, so a pair of elements that never meet in a molecule is a pair on which no table was ever checked against another, and the four are free to drift apart there with nothing to object.
That is a more useful lesson than “collect more data”. The small set was not too small in the statistical sense; it was too chosen. Twenty per cent coverage selected on a criterion correlated with the answer will produce clean separations, and it will produce them significantly, and they will not generalise.
What was computed, and how
Each of the four tables — Pauling, Mulliken, Allred–Rochow and Allen — is normalised to its own range across the eighteen elements all four carry, because they are in four different unit systems and only their rankings are comparable. Normalising to the range rather than to a standard deviation is a choice, and it is the one made for the thirty-one bonds; it matters here only in that both calculations make it, so the numbers are comparable across them. For each pair, the four normalised differences are computed; the pair is disputed when they do not all have the same sign, and its size is the mean of them in absolute value.
All 153 unordered pairs are run. The rank statistic is the fraction of disputed-against-agreed comparisons in which the disputed pair has the smaller difference: 0.9745 over 2,192 comparisons. The best threshold is found by trying every observed difference as a cut and counting errors.
Eight things are checked, and two of them are about the calculation being extended rather than about the new one. The thirty-one-bond subset must still reproduce exactly — six disputed, a clean gap — because a calculation that could not reproduce the one it extends would be measuring something else. The rest require that the full set not separate, that many agreed pairs sit below the largest disputed one so that it is not one awkward case, that the best threshold coincide with the small set’s line, that it get a handful wrong rather than none or many, that the rank statistic stay above 0.95, that the relative version fail too, and that every disputed pair satisfy the arithmetic bound.
Where the model stops
The four tables are also not four independent opinions. Mulliken’s is built from ionisation energies and electron affinities, Allred–Rochow’s from an effective nuclear charge and a covalent radius, and none of them is measuring one quantity — but they are all fitted, directly or indirectly, to overlapping bodies of the same data, so the four differences for a pair are correlated in ways this analysis treats as free. A spread computed from four correlated numbers understates the real disagreement, which pushes in the direction of finding fewer disputes rather than more.
Four tables is a small number of opinions, and “disputed” here means those four do not agree — a fifth table would change the classification of some pairs in both directions. Nothing here is a claim about how many electronegativity scales in the literature disagree about O–Br; it is a claim about these four.
Eighteen elements is also the intersection of what all four carry, which is itself a selection: the tables cover more elements individually, and the ones they have in common are the ones everybody had a value for, which correlates with being common and well-studied. The selection effect this essay is about applies one level up as well, and it cannot be removed with these four tables.
And a disputed sign is not the same as a disputed bond dipole. The quantity no scale prints is the actual charge separation, and a table’s difference is a proxy for it that should never be conflated with it. The tables disagree about the direction of the electronegativity difference; whether the physical dipole points that way is a further question involving lone pairs and hybridisation that no table addresses at all. That distinction was made at the start and it has not moved.
The generalisation
Two things travel.
The first is about clean boundaries. A separation with nothing in between is the most persuasive result a sorted list can produce, and it is also the one most easily manufactured by a sample chosen on anything correlated with the answer. The test is not more data of the same kind; it is data selected differently, which here meant every pair rather than every bond.
That is the same discipline a correlation is not an account applies to a fitted line: a statistic computed over a chosen population is a statistic about the choice as much as about the population.
The second is about what to do when it fails. The gap here dissolved and the line did not — the number was right and the claim around it was too strong. That is a common shape and it is worth naming, because the reflex on losing a clean separation is to abandon the finding, and the arithmetic said to keep it and weaken it. Ninety-seven per cent of comparisons in the right order is a good rule about a real thing; it is simply not a boundary, and it should have been quoted as the first from the beginning.
Who found it, and when
The four scales are Pauling’s of 1932, Mulliken’s of 1934, Allred and Rochow’s of 1958 and Allen’s of 1989. Their disagreements are well known in outline. The normalisation, the pair-by-pair sign test, the threshold and everything above are new arithmetic, and this is a check on the thirty-one-bond result rather than a report of anything published.
The number worth carrying out of it is 0.05065, and the sentence worth carrying is that it is a threshold and not a wall.
Still open: a two-variable boundary
The obvious open question is a fifth table. Every claim here is a claim about four opinions, and “disputed” is a property of that panel rather than of the pair; adding one more scale would move pairs across the boundary in both directions, and how many move is a measure of how stable the classification is. It would need one quoted table and no new calculation.
The nearer question is the two-variable rule. The exceptions divide cleanly by mechanism — agreed-but-small pairs have tiny spreads, disputed-but-large pairs have huge ones — so a rule using both the difference and the spread ought to do better than either alone, and the single ratio tried here is only the crudest way to combine them. Fitting the best straight boundary in the plane of the two, and reporting how many of the 153 it still gets wrong, would say whether six is the floor or whether the right two-variable rule separates them completely. It is one two-dimensional search over data already computed, and it decides whether the finding is a rule with exceptions or the wrong projection of a clean one.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The rule is not the lever — both name convention, electronegativity, empirical scale
- The value that only exists in the bond — both name bond dipole, convention, electronegativity
- The worst of the six was the one we asked about — both name convention, electronegativity, empirical scale
- What a dipole cannot tell apart — both name bond dipole, convention, electronegativity
- A capacity that is largest where there is none — both name convention, electronegativity
- A mean that is low rather than right — both name convention, electronegativity
Named objects
A dashed tag is an object no other essay names yet.
Bond dipoleConventionElectronegativityEmpirical scalePeriodic trend