What is taught wrongly

The rule is not the lever

One exception among nineteen flagged bond pairs would need seventy-three electronegativity tables to mean anything, and a rule that flagged fewer pairs looked like the way to need fewer. Searched over sixty-one thousand boundaries, nineteen is already the fewest that misses no disputed pair; each missed pair buys a table or five; and a rule flagging a single pair would still need fourteen. The requirement lives in the correlation between the tables, which moves it nine times as far as any rule can.

Worth reading first: No panel of this kind can find an exception · Two numbers caught what one could not.

Four electronegativity tables disagree about which atom is the negative end of some bonds. A boundary in the plane of difference and dispute separates the bonds whose polarity the tables dispute from those they agree on, with three exceptions: agreed pairs on the disputed side. Whether those three mean anything is a question about chance, and the chance turned out to depend on how correlated the tables are near the boundary. At the measured correlation of 0.325, four correlated tables produce 5.6 spurious agreements among the nineteen flagged pairs — more than the three observed — and a single exception would only be significant with seventy-three tables. Nobody has seventy-three.

That essay left two routes open. One was a fifth table of the right kind. The other was the rule itself. The requirement grows with the number of pairs the rule flags, because each flagged pair is another chance for tables to agree by luck, and nineteen is what the best boundary happened to flag. A rule that flags fewer pairs, even at the cost of letting a disputed pair or two through, would need fewer tables for the same test. The question is how many fewer, and it has a definite answer, because every boundary the plane can hold can be tried.

Letting disputed pairs through narrows the rule and barely moves the panel. Over 61,353 straight boundaries in the plane of difference and dispute, the fewest pairs any boundary flags for each number of disputed pairs it lets through (bars), and the number of tables that flagged count needs at the measured correlation of 0.325 (dots). Missing none, nineteen is the least — the published rule. Each miss saves one or two flags and two to five tables; at eight misses, half the disputed pairs, the rule flags 8 and still needs 46 tables.
Fig. 1 For each number of disputed pairs a boundary lets through, the fewest pairs any boundary flags, and the tables that count needs.

Nineteen is already the narrowest

A boundary here is a straight line in the plane of two variables for each bond: the mean normalised electronegativity difference across the four tables, and the spread of the four. Pairs below the line are flagged as disputed. Every line is a slope and a threshold, and for each of the slopes the original search tried, every threshold that changes which pairs are flagged was tried too — 61,353 boundaries in all. For each, two numbers were kept: how many pairs it flags, and how many of the sixteen genuinely disputed pairs it fails to flag.

The slopes run from −2 to +2 in steps of a hundredth, 401 of them, which covers boundaries that rise with the spread, fall with it and ignore it; a slope of zero is a rule on the difference alone. The original search found 41 slopes tying for its fewest errors, from 0.25 to 0.72, and chose the one with the widest margin. The search here asks a different question of the same lines — not which has fewest errors, but which flags fewest pairs for a given number of disputed pairs missed — so it can return a different slope even where it returns the same count.

Among boundaries that miss no disputed pair, none flags fewer than nineteen. The published rule is already on the edge of what the plane allows: it cannot be narrowed without letting a disputed pair through. The three exceptions are agreed pairs that sit among the disputed ones in both variables, and no straight line removes them without removing a disputed pair with them.

It helps to be clear about what an exception would have had to be for the question to arise. The disputed pairs are those where the four tables do not all put the negative end of a bond on the same atom — pairs where the tables are measuring different things and the difference between the atoms is small enough for that to show. A rule that predicts dispute from the size of the difference and the tables’ spread is a statement that dispute is a small-difference phenomenon. An exception is a bond the rule says should be disputed on which all four tables happen to agree. If it were real, it would be a bond whose polarity is settled for a reason the rule does not contain; if it is chance, four correlated tables landed the same way on a near coin flip. Telling those apart is what the panel of tables is for.

So narrowing the rule means trading. The frontier of that trade is the least number of flags for each number of misses. One miss allows seventeen flags; two, fifteen; three, fourteen; four, twelve; and at eight misses — half of the sixteen disputed pairs called clear — the narrowest boundary flags eight.

Each miss buys a table or five

At the measured correlation, nineteen flagged pairs need 73 tables. Seventeen need 68, fifteen need 64, twelve need 57, and the eight flagged at eight misses need 46. Each missed disputed pair saves between two and five tables.

That is a poor exchange on any accounting. A rule whose purpose is to say where a sign cannot be trusted is made worse by every disputed pair it calls clear, and after giving away half of them the panel needed has fallen by little more than a third. It remains more than ten times the four tables that exist.

It is also worth setting the numbers against what could be assembled. Published electronegativity scales number in the low tens, and many are built on the same few inputs — ionisation energies and electron affinities, bond energies, effective nuclear charges, atomic radii — so a panel of them would be more correlated, not less, than the four used here. Forty-six tables is not a panel that exists; seventy-three even less so. The frontier does not move the requirement from impossible to difficult. It moves it from one impossible number to a smaller impossible number, and it does so by damaging the rule it was meant to serve.

Three boundaries: missing nothing, missing four, missing eight. Every pair of bonds with a mean difference under 0.35, placed by the four tables' mean difference and their spread, disputed pairs filled and agreed pairs open. Three boundaries from the frontier are drawn: the narrowest that misses no disputed pair, flagging 19; the narrowest missing four, flagging 12 with no agreed pair among them; and the narrowest missing eight, flagging 8. The agreed pairs the first rule flags are the three exceptions; the price of losing them is four disputed pairs called clear.
Fig. 2 The pairs in the plane of dispute and difference, with the narrowest boundaries that miss none, four and eight disputed pairs.

The picture shows why the trade is so poor. The disputed pairs are not a tight cluster with the agreed pairs around them; they are spread along the bottom of the plane, at small differences, over the whole range of spreads. A boundary that excludes the three agreed pairs among them has to flatten, and flattening it cuts off disputed pairs at large spread. At four misses the narrowest boundary has almost no slope and flags twelve pairs, none of them agreed — the price of losing the three exceptions is four disputed pairs called clear. At eight misses the slope is zero, and the rule has become a plain threshold on the difference — a one-variable rule of exactly the kind the second variable was added to improve.

The exceptions themselves also disappear along the frontier, which is worth noticing for what it does not mean. A rule with no exceptions is not thereby a better rule; it is a rule that has stopped flagging the part of the plane where agreed and disputed pairs mix.

Even one flagged pair needs more than four

The limit of narrowing is a rule that flags a single pair. It cannot be a useful rule, but it bounds what any rule could achieve.

A rule that flags a single pair still needs more than four tables. The tables needed before agreement on one flagged pair is significant at five per cent, against the correlation between tables. Four independent tables agree on a sign one time in eight, which is above one in twenty, so even at zero correlation one pair needs 6 tables. At the measured correlation it needs 14, and at the top of the bootstrap range 58. No boundary, however narrow, makes four tables enough.
Fig. 3 Tables needed before agreement on one flagged pair is significant, against the correlation between tables.

Four independent tables all agree on a sign one time in eight, and a significant exception needs the chance of spurious agreement below one in twenty. So even a single flagged pair, with tables as independent as coins, needs six. At the measured correlation it needs fourteen. At the top of the correlation’s bootstrap range, 0.547, it needs fifty-eight.

The one-in-twenty standard deserves a word, because it is the conventional threshold rather than a demanding one. It asks only that an observed agreement be something chance produces less than five times in a hundred. A single flagged pair with four independent tables fails it by a factor of two and a half before correlation is considered at all.

No boundary, however it is drawn, makes four tables enough. The four tables that exist cannot establish a single exception to any rule of this kind, at any correlation, including none. That is a stronger statement than the one the panel calculation made, because it does not depend on the rule at all.

The correlation is the lever

The requirement has two inputs: how many pairs are flagged and how correlated the tables are. Their effects can be put side by side.

The panel needed grows with the log of the flagged pairs and far faster with the correlation. Tables needed before one agreement among the flagged pairs is significant at five per cent, against the number of pairs flagged, for six correlations between tables, both axes logarithmic. At zero correlation the requirement runs from 6 for a single pair to 13 for all 153. At the measured 0.325 it runs from 14 to 215. At the top of the bootstrap range, 0.547, a single flagged pair needs 58 and nineteen need 1,903; beyond thirty pairs no panel of two thousand is enough.
Fig. 4 Tables needed against pairs flagged, for six correlations between tables, both axes logarithmic.

Along the flagged axis the growth is slow — roughly a fixed number of extra tables for each doubling of the flagged count, because each flagged pair adds a chance and chances add logarithmically in the required panel. At zero correlation, going from one flagged pair to all 153 pairs in the data raises the requirement from 6 to 13. At the measured correlation the same range runs from 14 to 215.

Along the correlation axis the growth is steep. Correlated tables share part of whatever pushes their signs, so adding a table adds less independent evidence the more correlated it is, and at high correlation additional tables barely reduce the chance that all agree. At nineteen flagged pairs, a correlation of 0.110 needs 16 tables and one of 0.547 needs 1,903; above thirty flagged pairs at that correlation, no panel of two thousand is enough.

The correlation moves the requirement nine times as far as any choice of rule. The range of tables needed as three things are varied, on a logarithmic axis. Every rule from one flagged pair to all 153 spans 14 to 215 at the measured correlation. Every trade of misses for narrower flagging spans 46 to 73. The correlation's own bootstrap range, at the published nineteen, spans 16 to 1,903.
Fig. 5 The range of tables needed across every choice of rule, across the trade of misses for narrower flagging, and across the correlation’s bootstrap range.

A worked comparison makes the asymmetry concrete. Halving the flagged count from nineteen to nine at the measured correlation lowers the requirement from 73 tables to 49. Lowering the correlation from 0.325 to 0.2 at the same nineteen lowers it to 27, and to 0.11 lowers it to 16. A modest change in how much the tables share is worth more than any rule the data allow, and it does not cost a single disputed pair.

Every choice of rule, from a single flagged pair to all 153, moves the requirement by 201 tables at the measured correlation; the correlation’s bootstrap range moves it by 1,887 at the published rule’s nineteen. Trading misses for narrowness moves it by 27. The rule was never the lever. The correlation was, and it is not yet even known to better than a factor of five.

How the frontier and the requirement were computed

Each bond pair carries its four tables’ normalised electronegativity differences, their mean and their spread, and a label saying whether the four agree on the sign. A boundary with slope s and threshold t flags every pair whose mean minus s times spread is at most t. The slopes are the ones the original boundary search swept, and for each slope every distinct value of the projection is used as a threshold, which covers every distinct partition a line of that slope can make. For each number of disputed pairs left unflagged, the boundary flagging fewest pairs is kept; where several tie, the first found is kept, which changes the slope reported and not the count.

The number of tables needed for F flagged pairs at correlation ρ is the smallest k for which F times the probability that k equicorrelated signs all agree falls below −ln 0.95. That probability is the orthant probability of equicorrelated normals, computed as a one-dimensional integral over their shared part, exactly as the panel calculation did; this is the same requirement applied to other flagged counts.

The narrowest boundary for each number of misses. For each number of disputed pairs let through, the boundary that flags fewest pairs: its slope, the pairs it flags, how many of those the four tables agree on, and the tables needed at zero correlation and at the measured correlation.
Fig. 6 The narrowest boundary for each number of missed disputed pairs: slope, pairs flagged, exceptions among them, and tables needed.

The checks, run wherever these figures are drawn. No boundary that misses nothing flags fewer than the published nineteen. Each step along the frontier saves at most two flags and seven tables per miss, and at seven misses the requirement is still above forty. A single flagged pair needs more than ten tables at the measured correlation and more than fifty at the top of its range. The correlation’s bootstrap range moves the requirement further than the whole range of flagged counts. The refusal is the published calculation itself: nineteen flagged pairs must need seventy-three tables at the measured correlation and ten at zero here too, or the requirement being narrowed would not be the one that was measured.

Where the calculation stops

Straight boundaries only. A curved boundary, or one drawn in more variables, could flag a different set, and the frontier here is the frontier of lines. The mixing of agreed and disputed pairs along the bottom of the plane suggests a curve would not do much better, but that is a reading of the picture and not a search.

Equicorrelated tables. The requirement treats every pair of tables as correlated by the same amount. Two of the four are nearly one table and one runs slightly against the others, so the real correlation structure is not uniform, and a fifth table’s value depends on which tables it resembles. That is the other open route and it is not taken here.

A five per cent test. A looser standard needs fewer tables; a stricter one, more. The comparison between rule and correlation does not depend on the level, because both enter through the same expectation.

One sign per table per bond. Each table contributes only which end of a bond it makes negative. A test using the size of each table’s difference, rather than its sign, would extract more from the same four tables, and the requirement computed here is the requirement for sign agreement specifically. Whether a magnitude-based test could establish an exception with fewer tables is a different question, and a natural one, but it would be testing a different claim: not that the tables agree on a polarity, but that they agree on how polar.

And the flagged pairs are the data’s. Other bonds, or other elements, would move every count. The conclusion that the correlation dominates rests on the shape of the requirement, which is logarithmic in one input and steep in the other, and would survive a different data set; the numbers would not.

Where a test’s power comes from

The instinct when a test cannot reach significance is to sharpen the question — to ask about fewer cases, so that fewer things can go wrong by chance. Here that instinct buys almost nothing, and the reason is general. When the evidence comes from several instruments, how independent the instruments are sets the power of the test far more than how many cases are asked about, because cases add chances logarithmically and correlation removes information from every instrument at once.

Electronegativity tables are a clean example because they are few and openly correlated, but the same arithmetic governs any consensus of methods built from overlapping inputs. Asking fewer questions of a correlated panel does not make it a better panel; the residue a consensus leaves is set by what the members share.

The same tables have shown the pattern before in other forms. They agree about the order of the elements almost perfectly and disagree about sizes by large factors; they give a dipole moment that runs over a factor of eleven for one molecule. In each case what the four tables share decides what their agreement is worth, and in each case the agreement looks stronger than it is if the tables are counted as four. Here the sharing has a number, and the number turns out to be the only input to the test that matters.

Still open: what a fifth table is worth

The obvious open question is the one this left aside. A fifth table’s value to the panel depends on its correlation with each of the four, and the four are far from equicorrelated near the boundary. Replacing the single correlation with the measured matrix, and computing the orthant probability for a panel of five under a candidate’s correlations with the others, would score a table for its usefulness before its values were compared — and would say whether any table that could exist brings seventy-three down to something a chemist could assemble.

The nearer question is the correlation’s own uncertainty. The bootstrap puts it anywhere from 0.11 to 0.55, and that range is worth a factor of a hundred in the panel needed. Measuring the correlation on more pairs near the boundary — more elements, or bonds beyond the eighteen elements used — would narrow the one input that decides the answer, and that is a question of data rather than method.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ConventionElectronegativityEmpirical scaleModel limitPolarityRank correlation