No panel of this kind can find an exception
Worth reading first: Two numbers caught what one could not · A gap that was a choice of bonds.
Two numbers caught what one could not found the best boundary in the plane of electronegativity difference against inter-table dispute, and left three exceptions: agreed pairs on the disputed side of every line. Then it asked whether three meant anything. For a pair whose difference is near zero each table’s sign is close to a coin flip, four coins all land the same way one time in eight, and the rule flags nineteen pairs — so 2.375 of them should look agreed for no reason. Three is not a discrepancy.
It called that figure a lower bound, because the tables are certainly correlated and correlated signs agree more often, and it closed with a question about the instrument rather than the chemistry:
If four tables generate two or three spurious agreements among nineteen flagged pairs, then no panel of four can ever demonstrate an exception to this rule, however the boundary is drawn. Working out how many tables are needed before one exception would be significant, and whether that many independent scales exist, would say whether the rule’s exactness is testable at all or is a claim this method cannot reach.
The answer depends on a number that essay did not measure — the correlation between the tables — and once it is measured the answer is not a count of tables at all.
Ten, if tables were coins
The arithmetic for independent tables is short. With k tables, a flagged pair looks agreed by chance with probability 2·2⁻ᵏ, so nineteen flagged pairs are expected to produce 19·2¹⁻ᵏ spurious agreements. For a single observed exception to be significant at five per cent, the chance of even one spurious agreement must be below one in twenty, which for a rare event means an expectation below −ln 0.95 = 0.0513.
Four tables give 2.375. Eight give 0.148. Ten tables give 0.037, the first panel for which one exception would mean something.
That is already more tables than anyone uses to argue about a bond’s polarity. But it is the optimistic answer, and the reason is in the phrase the chance calculation used for its own assumption.
The tables are not coins
Electronegativity tables are built from overlapping data — bond energies, ionisation energies and electron affinities, effective nuclear charges, spectroscopic valence energies — and they are not one quantity but they are not four unrelated ones either. Where two elements’ true difference is large every table sees it, and they agree. What matters for the chance calculation is how much they share where the difference is nearly zero, which is exactly where their signs are close to coin flips.
Measuring that has a trap in it. The nineteen flagged pairs were selected by a boundary fitted to which pairs are disputed, so estimating the tables’ agreement on those pairs would use the disputes to measure the chance of disputes. The correlation is measured instead on the N pairs with the smallest normalised differences, chosen by size alone, with N from nineteen upward.
Two routes share no arithmetic. One takes the Pearson correlation between each pair of tables’ normalised differences over the window and averages the six. The other counts how often each pair of tables gives the same sign and converts that fraction to a correlation by Sheppard’s formula, which for two normal variables with correlation ρ says their signs agree with probability ½ + arcsin(ρ)/π.
Over the nineteen smallest pairs the first route gives 0.322 and the second 0.325. Widen the window and both rise — 0.49 and 0.54 over twenty-five pairs, 0.66 and 0.79 over forty, 0.97 and 0.98 over all 153 — because a wider window admits pairs whose real difference every table sees. The number that matters for the flagged pairs is the one at the boundary, and it is about a third.
Two of the four are nearly one
The average hides a structure. Near the boundary, Allred–Rochow and Allen correlate at 0.92. Pauling correlates with each of them at about 0.7. Mulliken correlates with Pauling at 0.10 and with the other two at −0.29 and −0.20 — slightly against them.
Allred and Rochow built their scale from an effective nuclear charge felt at the covalent radius, and Allen built his from the average energy of an atom’s valence electrons measured spectroscopically; near zero difference those two readings of how tightly a valence shell is held move together almost perfectly. Mulliken’s scale averages ionisation energy and electron affinity, and near the boundary it goes its own way. A panel of these four is not four opinions.
A correlation matrix says how many independent opinions it holds through its participation ratio, (Σλ)²/Σλ² over its eigenvalues, which is four for independent tables and one for identical ones. Over the nineteen smallest pairs it is 2.02. Over forty it is 1.59, and over all 153 it is 1.05. Near the boundary the four tables are worth about two.
That has a consequence for a single molecule as well as for the rule. When four tables disagree about hydrogen iodide’s dipole by 1.245 debye, the spread looks like the verdict of four witnesses and is closer to the verdict of two, since two of the four are nearly repeating each other near a small difference. A spread across a panel is only as informative as the panel is independent, and a bond-by-bond sum of such predictions — the construction a dipole was never simply — inherits the same limit on every bond whose difference is small.
A model that can be checked
Counting tables at a correlation needs a model of correlated signs, and the simplest is the one Sheppard’s formula already assumes: each table’s difference for a pair near the boundary is a normal variable, and every two tables’ variables correlate at the same ρ. The chance that k such signs all agree is then an integral over the part they share,
which returns one in eight for four independent signs and reduces to Sheppard’s formula for two — both checked, the second by a route with no arcsine in it.
Equicorrelation is a simplification the pairwise structure above visibly breaks. So the model is tested against something it did not use: the fraction of pairs on which all four tables agree, which the pairwise sign agreement predicts through and which can be counted directly.
Over the nineteen smallest pairs, all four agree on 31.6 per cent against a predicted 29.5. Over twenty-five, 44.0 against 42.3; over thirty, 50.0 against 48.8; over forty, 62.5 against 61.4. The prediction is a little low and never by more than two points. A model built only on pairs of tables predicts what the whole panel does, which is what licenses using it to ask about panels of other sizes.
What correlation does to the count
For independent tables each new table halves the expected number of spurious agreements, and the curve drops through the significance line at ten. For correlated tables it does not. The tables share part of their luck, and a shared lucky draw is not diluted by adding another table that shares it; the expectation falls only as a power of the panel size rather than exponentially.
At a correlation of 0.1 the requirement is fifteen tables. At 0.2 it is twenty-seven, at 0.3 fifty-eight, at 0.4 a hundred and sixty-five. At a correlation of a half no panel of four hundred tables reaches it — four hundred tables would still be expected to produce 0.095 spurious agreements among nineteen pairs, nearly twice the threshold.
At the four tables’ measured correlation near the boundary, 0.325, the requirement is seventy-three tables. That number has an honest uncertainty attached, because it rests on nineteen pairs. Resampling those pairs four hundred times and recomputing the sign correlation puts its tenth and ninetieth percentiles at 0.11 and 0.55 — a range over which the requirement runs from about sixteen tables to more than any panel could hold.
The exceptions were below chance
The same model reads the three exceptions differently. Independent tables expected 2.375, and three sat slightly above that. At the measured correlation four tables all agree on a flagged pair 29.5 per cent of the time, so nineteen flagged pairs are expected to show 5.6 spurious agreements. Three is below that.
So the earlier essay’s reading was right in direction and too cautious in size. Its expectation was a lower bound, as it said, and the true expectation is more than twice it. The three exceptions are not borderline evidence of anything; they are fewer agreements by chance than this panel would usually produce.
Whether seventy-three tables could exist
They could not, in the sense that matters. Electronegativity scales in the literature number in the tens rather than the hundreds, and the question is not how many tables are printed but how many independent readings they amount to near the boundary. Four of them amount to two. Adding a fifth that correlates with the existing ones at a third adds less than a table’s worth of independence, and a fifth built from the same kind of data as Allred–Rochow or Allen would add almost nothing.
The kinds of evidence are few and can be listed. Thermochemistry gives the bond-energy scales that descend from Pauling’s. Atomic spectroscopy gives ionisation energies and electron affinities, from which Mulliken’s and Allen’s are built, and they share that source though they use it differently. Electrostatics gives the scales that use an effective nuclear charge and a size, Allred and Rochow’s among them. Finer distinctions exist within each — which bonds, which radii, which valence state — but they are variations on a kind rather than new kinds, and a variation inherits most of its parent’s errors near a small difference.
A panel that could test the rule would need tables built from genuinely unrelated kinds of evidence, and even then the requirement for truly independent tables is ten, which is more than the number of distinct kinds of evidence electronegativity has been built from. So the rule’s exactness is not a claim a panel of this kind can confirm or refute. The instrument’s noise floor sits above the effect, and a residue below its own noise is not something more of the same instrument resolves.
What was counted, and how
The four tables — Pauling, Mulliken, Allred–Rochow and Allen — each cover the same eighteen elements, and each is normalised by its own range, the step that makes four incommensurable scales comparable. For every one of the 153 pairs the four normalised differences, their mean size and their spread come from the whole-panel comparison, and the nineteen flagged pairs and the three exceptions from the best boundary in the plane.
The orthant probability is integrated on a grid of 12,000 points over the shared normal part. The requirement at each correlation is the smallest panel for which nineteen flagged pairs are expected to produce fewer than 0.0513 spurious agreements, found by search up to four hundred tables. The bootstrap resamples the nineteen smallest pairs with replacement, four hundred times, from a fixed seed.
The checks are these. Four independent signs agree one time in eight, and two signs at a correlation of a half agree at Sheppard’s rate, by the integral. Independent tables need exactly ten. Near the boundary the correlation is above 0.2 by both routes and the routes agree within a tenth. The correlation rises through every wider window to above 0.9 over all pairs. At the measured correlation the requirement is more than three times ten, and at a half no panel of four hundred reaches it. The observed exceptions are fewer than the measured correlation predicts. And the refusal: in the three narrowest windows the fraction of pairs on which all four signs agree must match what the pairwise agreement predicts to within five points, which a model that did not describe these tables would fail.
Where the panel model stops
Equicorrelation is not the tables’ structure. Two tables nearly duplicate each other and one runs against them. A model with a separate correlation for each pair would give a different requirement, and it would not give ten: the independent case is the most favourable one and every positive correlation raises the count.
Nineteen pairs is few. The bootstrap range is wide because the estimate is, and the conclusion is stated over that range rather than at its centre — from roughly sixteen tables to more than can be assembled.
A five per cent threshold is a convention, and so is asking for one exception rather than two. A stricter test needs more tables; a laxer one fewer. Neither changes the shape of the curve, which is the finding.
And the flagged set is itself a product of four tables. A different panel would draw a different boundary and flag a different number of pairs, and fewer flagged pairs need fewer tables. The requirement here is for this rule on this panel, which is the question as it was asked.
A consensus has a noise floor
The transferable point is about classification by panel. Whenever a label is assigned by asking several instruments and recording whether they agree — whether tables agree on a sign, raters on a category, methods on a ranking — the label’s reliability near its boundary is set by how correlated the instruments are there, not by how many there are.
Independence makes a panel powerful: every added member halves the chance of an accidental consensus. Correlation takes that away, and it does so most where the instruments share their construction. The check is cheap and should come before any exception is counted: measure the members’ agreement near the boundary on cases chosen without the label, and convert it to an effective panel size. Here four tables were two, and the rule’s exceptions were fewer than two tables’ worth of luck produces. The same caution applies wherever a ranking is read as a difference or a table could not have mattered — the arithmetic of agreement has to be done on the instruments before it is done on the subject.
Who counted it first
Sheppard’s result on the signs of correlated normal variables is from 1899, and the orthant probability for equicorrelated normals is a standard reduction. The four scales are Pauling’s of 1932, Mulliken’s of 1934, Allred and Rochow’s of 1958 and Allen’s of 1989.
The measurement of the four tables’ correlation near the dispute boundary, the check of the correlated-sign model against the whole panel, and the panel size one exception would need are computed here. Two numbers caught what one could not deserves the credit for turning three exceptions into a question about the instrument, and for calling its own expectation a lower bound — which it was, by more than a factor of two.
Still open: which fifth table would help most, and a rule that flags fewer
The obvious open question is the fifth table, and the correlation structure now says what kind to want. A table that correlates with Allred–Rochow and Allen near the boundary adds almost nothing; one that behaves like Mulliken there adds most. The effective number of tables a panel gains from a fifth member is computable from its correlations with the four, so a candidate scale could be scored for its value to the panel before its values for these eighteen elements are even compared.
The nearer question is the rule itself. The requirement grows with the number of pairs flagged, and nineteen is what the best straight boundary flags. A boundary that flags fewer pairs at the same miss rate would need fewer tables for the same test — so whether a slightly different rule, trading an error for a narrower flagged set, would bring the requirement within a panel that could exist is one search over boundaries the plane already holds.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The quantity no scale prints — both name convention, electronegativity, model limit, polarity, rank correlation
- A mean that is low rather than right — both name convention, electronegativity, model limit, rank correlation
- A size the fit was not made from — both name electronegativity, empirical scale, model limit, rank correlation
- The worst of the six was the one we asked about — both name convention, electronegativity, empirical scale, model limit
- A capacity that is largest where there is none — both name convention, electronegativity, model limit
- The lone pair is not the missing term — both name electronegativity, model limit, polarity
Named objects
A dashed tag is an object no other essay names yet.
ConventionElectronegativityEmpirical scaleModel limitPolarityRank correlation