Where the atoms go

Two numbers caught what one could not

No single threshold on the electronegativity difference separates the bond pairs whose polarity the four tables dispute from those they agree on — the best misses six of a hundred and fifty-three. Adding how much the tables disagree halves that to three and catches every disputed pair. And three is what four coin-flips produce: the rule flags nineteen pairs, an eighth of which should look agreed for no reason at all, which is 2.375 against the three observed.

Worth reading first: A gap that was a choice of bonds · A dispute is a small difference.

Widening the panel took a clean result and broke it on purpose. The thirty-one-bond comparison found that on thirty-one common bonds, the pairs whose polarity the four electronegativity tables disagree about are exactly the pairs whose difference is small — a clean gap, with every disputed pair below a threshold and every agreed one above it.

Extending that to all one hundred and fifty-three pairs the four tables all cover — a set that is five times the size and includes many pairs no common molecule contains — the gap disappeared. The same threshold is still the best one available, and it now misclassifies six pairs instead of none. A boundary with a width is a different claim from a boundary with a gap, and that is worth saying.

It also noticed how the six failed, and made a prediction from it. The agreed pairs that fall below the line have tiny disputes; the disputed pairs that fall above it have enormous ones. So a rule using both the difference and how much the tables disagree ought to do better than either alone. The single combination it tried — the ratio of one to the other — did not.

A ratio is one line through the origin. The search here covers all the lines.

Why the search is easy

The problem looks two-dimensional and is not, which is worth saying because it is the reason the search costs nothing.

A straight boundary in the plane is size = c + m × spread, and a pair is called disputed when its difference falls below the line. Rearranged, that is a threshold on the single quantity size − m × spread. So for any fixed slope the question collapses to the one-variable problem already solved: given one number per pair, find the threshold that misclassifies fewest.

That inner problem has an exact answer rather than a fitted one. The best threshold is always one of the projected values themselves — nothing between two adjacent points classifies differently from the lower of them — so it is found by trying each and counting, not by optimising anything.

So the whole search is 401 slopes, each with its exact optimum. There is no fitting anywhere in it, and the comparison against the one-variable rule is a comparison between two optima rather than between an optimum and something plausible.

Two quantities, and the best straight line between the two kinds of pair. Every pair the four electronegativity tables all cover, by how much the two elements differ and by how much the tables disagree about them. Filled points are pairs whose sign the tables dispute. The line is the boundary that misclassifies fewest — 3 of 153, against 6 for the best rule using the difference alone. It slopes upward, which is the mechanism: more dispute buys a larger difference and still leaves the sign in doubt.
Fig. 1 Every pair the four tables cover, by difference and by dispute, with the disputed ones filled and the best boundary drawn.

What it finds

The best boundary misclassifies three of the hundred and fifty-three. The best rule on the difference alone misclassifies six.

A second variable halves the mistakes and changes what kind they are. The best rule on one variable against the best on two, by how many of the 153 pairs each gets wrong and in which direction. The one-variable rule makes both kinds of mistake; the two-variable rule makes only one, calling three agreed pairs disputed and no disputed pair agreed. As a screen for where a sign cannot be trusted, that is the useful way round.
Fig. 2 The two rules by how many pairs each gets wrong, and in which direction.

Halving the errors is the headline, and it is the less interesting half. The more interesting one is that the kind of mistake changes.

The one-variable rule makes both kinds: it misses disputed pairs, calling them clear, and it over-flags agreed pairs, calling them disputed. The two-variable rule makes only the second kind. It catches every one of the sixteen disputed pairs and wrongly flags three agreed ones.

For a rule whose job is to say where a computed bond polarity cannot be trusted — and the polarity is what a dipole is read from — those two directions are not equivalent. Over-flagging costs a reader some caution about a pair that turns out to be fine. Missing costs them a sign they had no reason to doubt and which the tables do not support. A rule that never misses and sometimes over-flags is a usable screen; one that does both is a description.

The disputed pairs, and where each falls against both rules. Every pair whose sign the four tables do not agree on, with its electronegativity difference, how much the tables disagree, and whether each rule catches it. The one-variable rule misses the pairs with large differences and large disputes; the two-variable rule catches all of them.
Fig. 3 The sixteen disputed pairs against both rules. The one-variable rule misses the ones with large differences and large disputes.

And it still does not separate them

The pairs no line in this plane gets right. Every pair the best two-variable boundary misclassifies, by how far it sits on the wrong side. All three are agreed pairs called disputed, and all three sit close to the line — which is what makes them exceptions rather than counterexamples: they are pairs the rule is unsure about, not pairs it is confidently wrong about.
Fig. 4 The three pairs no line in this plane gets right, by how far each sits on the wrong side.

Three is not zero, and none of the 401 slopes reaches zero. So the answer to the question — is six the floor, or does the right two-variable rule separate them completely — is neither. Six was not the floor, and the exceptions survive.

That matters because it rules out a particular hope. If the two groups had been separable in this plane, the six failures of the one-variable rule would have been a projection artefact: the right quantity existed and the one-variable rule had been looking at a shadow of it. They are not — so within this panel, any rule of this shape has exceptions.

The three survivors are all close to the line, within five hundredths of it in the projected coordinate against a range spanning a full unit. So they are pairs the rule is unsure about rather than pairs it is confidently wrong about. The next section asks whether they are pairs at all, or arithmetic.

Or are the three exceptions chance

There is a calculation that has to be made before the exceptions can be called exceptions, and it is two lines long.

“Agreed” means four tables gave the same sign. For a pair whose difference is genuinely near zero, each table’s sign is close to a coin flip — the tables are built from different data and a hair’s-breadth difference falls whichever way that data pushes it. Four independent coins all landing the same way happens one time in eight.

The rule flags nineteen pairs as disputed. If an eighth of those show spurious agreement, that is 2.375 pairs which the tables happen to concur about and which the rule therefore counts as mistakes.

Three are observed.

Three exceptions, against the two and a half that four coins produce. The rule calls 19 of the 153 pairs disputed. For a pair whose difference is near zero each table's sign is close to a coin flip, and four independent signs all fall the same way one time in eight — so 2.38 of those 19 are expected to look agreed for no reason. Three do. The exceptions are not evidence that the boundary is wrong; they are evidence that four tables is a small panel.
Fig. 5 The three exceptions against the number four coin-flips produce among the flagged pairs. The expectation is a lower bound.

A ratio of 1.26 between observed and expected is not a discrepancy, and the direction of the bias makes it weaker still. The four tables are not independent — they are built from overlapping measurements and their signs are certainly correlated — and correlated signs agree more often than independent ones. So 2.375 is a lower bound on what should be seen, and three sits above a lower bound, which is where it belongs.

That changes the conclusion. The boundary is not a rule with three exceptions. It is a rule whose exceptions are indistinguishable from the noise a four-table panel generates, and the honest statement is that the classification may be exact and this panel cannot show it.

It also relocates the open question. The interesting number is no longer how many pairs the rule gets wrong; it is how many tables would be needed before a real exception could be told from a lucky agreement. With four, chance produces two or three. With eight, it would produce one in a hundred and twenty-eight — about a seventh of a pair — and a single exception would mean something.

The answer is a band, not a line

How many pairs each slope gets wrong, and the flat bottom it has. The number of misclassified pairs against the slope of the boundary, with the best threshold taken exactly at each slope. At slope zero the rule is the one-variable one — the difference alone — and it misses 6. The minimum is 3 and it is reached by 41 different slopes, from 0.25 to 0.72, so the answer is a band of lines rather than one.
Fig. 6 Misclassifications against the slope of the boundary. The minimum has a flat bottom forty-one slopes wide.

Forty-one of the four hundred and one slopes tie at three mistakes, spanning 0.25 to 0.72 — nearly a factor of three in slope. Every one of them classifies all one hundred and fifty-three pairs identically.

Every line that does equally well, drawn together. The 41 boundaries that tie at 3 mistakes, from slope 0.25 to 0.72, with the disputed pairs marked. They fan across a wide range and classify identically, because the pairs are sparse where the lines differ. Quoting one of them as "the rule" would be reporting the sweep's first entry rather than a property of the data.
Fig. 7 The tied boundaries drawn together. They fan widely and separate the same points, because the pairs are sparse where they differ.

This is the honest limit on what the rule can be said to be. The data rule out most slopes and cannot choose among these. Reporting “the boundary has slope 0.61” would be reporting which tied line the sweep happened to prefer on a secondary criterion, dressed as a measurement.

What the data do support is the sign and the rough size. Every tied slope is positive and between a quarter and three quarters, so more dispute genuinely does buy a larger difference — the mechanism argued for — at a rate somewhere in that band and not more precisely than that.

What was computed, and how

Four quoted electronegativity tables, every element all four cover, and every pair of them: 153. For each pair, the difference between the two elements on each table normalised to that table’s own range, giving four normalised differences per pair. The size is their mean and the dispute is their range. A pair is contested when the four differences do not all have the same sign — that is, when the tables do not agree which of the two elements is the more electronegative.

None of that is new; it is the one-variable construction, reused unchanged so that this extends that measurement rather than a neighbouring one. The normalisation in particular is the step that made four incommensurable scales comparable at all, and every quantity here inherits it.

What is new is the sweep, and it is 401 slopes with an exact threshold at each. The tie-breaking deserves a note: among the slopes that tie for fewest errors, the one reported is the one with the widest margin — the largest gap between the lowest agreed pair and the highest disputed one in the projected coordinate. That is a secondary criterion and it is stated rather than silently applied, because without one the reported slope would be whichever tied value the sweep reached first, which is a property of the array of slopes.

Seven checks. One reproduces the one-variable rule at slope zero, which is the check that this is the same problem widened. One says two variables beat one. One says no line separates them, which is the refusal that keeps the result from overclaiming. One says the slope is positive, in the direction the mechanism predicts — and it is worth knowing that the sign is easy to get wrong, because the projection subtracts the slope and the boundary adds it. One says every remaining mistake is in the safe direction. One says a band of slopes ties. And one says the two kinds of mistake sum to the total, so nothing is being dropped from the count.

Why the improvement is smaller than it looks

It is worth resisting one reading of the halving, because it is an easy mistake to make.

Going from six errors to three on a fixed set of a hundred and fifty-three points, by adding a parameter, is not by itself evidence that the second parameter belongs. A boundary with two free numbers has more freedom than one with a single threshold, and more freedom fits a finite set better whether or not the extra freedom corresponds to anything. Exactly the same caution applies here as applies to a fitted residual gaining a second term: the improvement is what the extra degree of freedom buys, and the question is whether the improvement is the kind the mechanism predicted.

Three things say it is, and none of them is the error count.

The direction was predicted before the search. The argument came from the way the six failures divided — small differences with tiny disputes, large differences with huge ones — that the boundary should rise with the dispute. Every one of the forty-one tied slopes is positive, so the search found the sign the argument called for rather than whichever sign fitted better.

The kind of error changed, which a mere extra degree of freedom has no reason to do. A rule with one more parameter would be expected to trim errors of both kinds roughly in proportion. This one removed every error of one kind and left the other.

And the magnitude is right. The slope band runs from a quarter to three quarters, so the rule trades roughly half a unit of difference for a unit of dispute — the same order, which is what one expects if the two quantities are measuring comparable contributions to how doubtful a sign is. A slope of a hundredth or of fifty would have fitted equally well in principle and would have meant that one variable was doing all the work.

So the case for the second variable rests on the shape of what changed, and the error count is the least of it. That is the same discipline needed when a two-term fit halves a misfit and leaves its sign pattern intact.

Where the model stops

Four tables is a panel and “disputed” is a property of that panel rather than of the pair. A fifth table is the obvious next step and that remains true; adding one would move pairs across the boundary in both directions, and the three exceptions here might become two or five.

A straight boundary is a choice of shape, and the finding that no line separates the groups is not a finding that no curve does. A curve would almost certainly reach zero errors on 153 points, and it would mean nothing: with enough freedom any two finite sets separate. That is the same trap a table that could not have mattered set from the other direction, where a rule looked strong because the set it was tested on was small. The line is the right shape here because the mechanism argues for it — two effects, each adding to how doubtful a sign is — and because it has two parameters against a hundred and fifty-three points.

And the normalisation is doing work. Each table’s differences are scaled by that table’s own range, which is what makes four incommensurable scales comparable, and it is the normalisation used whenever the tables are put side by side. A different one would change both quantities and could change the slope; it would not change whether a line separates the groups, since that is invariant under scaling each axis.

The generalisation

The transferable point is about what “the rule has exceptions” means before and after a search like this.

Before, it is ambiguous between two very different situations. Either the rule is nearly right and the exceptions are noise — in which case a better projection of the same data removes them — or the rule is genuinely approximate and the exceptions are part of the phenomenon. A one-variable failure cannot distinguish those, and that ambiguity had stood since the gap first closed.

After, it is settled: the exceptions survive the best line in the best plane the data offer, so they are the second kind. That is a more useful thing to know than the error count, and it is what a two-dimensional search buys that a cleverer one-dimensional rule could not.

The second point is about flat minima. A search that returns a best value should always be asked how much better than the runners-up it is, and here the answer is that forty-one values tie exactly. Reporting the winner without the band would have produced a slope to quote, and quoting it would be the familiar mistake with a threshold: treating the output of an optimiser as a measurement when the data cannot distinguish it from its neighbours.

Who found it, and when

The four electronegativity scales are quoted and old; the normalisation, the dispute measure and the classification are constructions rather than quotations. The two-dimensional search is arithmetic over numbers already computed — 401 slopes times 153 pairs.

That cheapness is worth noticing rather than passing over. A question that is one search over data already in hand should be asked as soon as it is posed, because a cheap test put off for later tends to stay unrun.

Still open: a fifth table, and how many tables are enough

The obvious open question is the fifth table. It is the only way to find out whether “disputed” is a property of the pair or of the panel, and everything else here is conditional on the answer. It needs one quoted table and no new method, and the same search would run on it unchanged.

The nearer question follows from the chance calculation and is a question about design rather than about chemistry. If four tables generate two or three spurious agreements among nineteen flagged pairs, then no panel of four can ever demonstrate an exception to this rule, however the boundary is drawn — the noise floor of the classification sits above the effect being looked for. Working out how many tables are needed before one exception would be significant, and whether that many independent scales exist, would say whether the rule’s exactness is testable at all or is a claim this method cannot reach. That is a calculation about the instrument rather than the subject, and it is the kind of calculation any classification by panel eventually needs.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationDipole momentElectronegativityModel limitPolarisability