The worst of the six was the one we asked about
Worth reading first: A correlation is not an account · A capacity that is largest where there is none.
Take a constructed quantity — the capacity, the largest amount of electron a cubic energy curve says an atom can accept before its chemical potential turns round — and ask what it is. It correlates with the second ionisation energy, and a quantity that is ninety per cent of one input has a simpler name than the one it carries — which is the test a correlation is not an account applied and failed to pass.
It is not ninety per cent of it. A straight line explains 2.9 per cent, the best of five two-parameter forms explains 43.8, and the counterexample is a pair: nitrogen and potassium sit 2.02 electronvolts apart in the second ionisation energy and their capacities differ by a factor of 193. No monotone function of one input can do that.
Then it asked for the same search against everything else — the first ionisation energy, the electron affinity, the hardness, the atomic radius — on the reasoning that an input for which no close pair has a large ratio would be an input the capacity might actually be a function of.
Running that produced two things: an answer, and a defect in the question.
The test had to be repaired first
The question says “the closest pair”, and the obvious reading is by value: two atoms within some small fraction of the input’s own range. That is what the first test did for the second ionisation energy, with a window of three electronvolts.
Run that way against every input, and one of the results is impossible. The control — the capacity tested against itself — returns a ratio of 43.9, for lithium against carbon. The capacity is not a function of the capacity.
The reason is not subtle once seen. The capacities run from 0.073 for lithium to 31.54 for nitrogen — a factor of four hundred and thirty — and a tenth of that linear range is 3.15, which is a window containing eight of the fourteen atoms at once. A tenth of a linear range is a tiny window at the top of a heavy-tailed distribution and an enormous one at the bottom, so a proximity test on a raw scale is not a proximity test on the quantity. It is a proximity test on wherever the quantity happens to be large.
Nor is this an artefact of choosing a tenth. Any fixed fraction of a linear range has the same shape, and the only fractions small enough to separate lithium from sodium are far too small to admit any pair at the top at all. The measure is wrong rather than badly tuned.
Adjacency in rank has no such problem: the closest pair in an input is the two atoms next to each other when the atoms are sorted by it, whatever the units and whatever the spacing. Run that way the control returns 4.95 — and that number is the second thing the control gives.
One more thing follows from the repair, and it is the reason the raw-scale version is worth keeping beside the repaired one. The original search used a window of three electronvolts on the second ionisation energy, which is a raw-scale window — and it found nitrogen and potassium, which the rank version also finds. So the defect did not change that answer. It would have changed the answers for the inputs not yet tested, which is exactly the situation in which a silently repaired test is most dangerous: the one case anybody had checked was the one case the defect did not touch.
The floor
Fourteen atoms spread over a factor of four hundred cannot be packed tightly. Even sorted by the capacity itself, some adjacent pair differs by a factor of five — bromine at 6.368 and nitrogen at 31.54 are neighbours in the ordering because nothing lies between them. Nitrogen is the outlier that sets the floor, and it is the same nitrogen that appears in the original counterexample pair; the two facts are related, and neither is an accident of the test.
So 4.95 is a floor, not a pass mark. No input can be shown to do better than that on this data, and every number in what follows is read against it. An input returning 5 would be indistinguishable from a perfect one; an input returning 200 is saying something. The distance between those two is the whole resolving power of this collection on this question, and it is about a factor of forty.
That is worth stating because the alternative reading of a test where everything fails is that the test is too strict. It is not too strict — the control shows exactly how strict it is, and the spread between the best input and the worst is a factor of twenty-four, which is far more than the floor could account for.
The ranking
Six inputs, ordered by their worst adjacent pair:
The cubic coefficient, ×7.9 — oxygen against potassium. The hardness, ×10.4 — lithium against aluminium. The electronegativity, ×24.2 — nitrogen against oxygen. The first ionisation energy, ×24.2 — oxygen against nitrogen, the same two atoms, because they are adjacent in both. The electron affinity, ×43.9 — lithium against carbon. And the second ionisation energy, ×192.6 — nitrogen against potassium, which is the original pair, recovered by a test that was not built to look for it.
Every one of them is several-fold, so the answer to the question as put is that there is no input the capacity is a function of. What the ranking adds is that the input under suspicion is not one candidate among several. It is last, by a factor of twenty-four over the best and by thirty-nine times the floor.
There is no external input to find
The ranking has a structure, and reading it honestly requires saying what the six inputs are.
The capacity comes from a cubic fitted through four points — the dication, the cation, the neutral atom and the anion. Its three coefficients are the electronegativity, the hardness and the cubic term; the three energies it is fitted through are the first ionisation energy, the electron affinity and the second ionisation energy. Every candidate the model can offer is one or the other.
And the ranking splits exactly on that line: the three coefficients occupy the first three places and the three data the last three, with no interleaving at all. That is not a discovery. It is what a working test ought to produce, since the capacity is computed from the coefficients and the coefficients are what the data were compressed into — so it functions as a second control, on the test’s ability to see structure it should see.
Which means the proposed search cannot succeed within the model. There is no input here that is not already part of the cubic, so the best possible outcome was always “the capacity is nearly a function of one of its own coefficients”, which says nothing about what the capacity is. To do better one would need a quantity from outside — a measured polarisability, an atomic radius from a different construction — and none is available here in a form all fourteen atoms share.
It is worth being clear about what kind of failure that is. It is not that the test was run and returned nothing; it is that the test could not have returned anything, and knowing so is a result about the construction rather than about the atoms. A quantity built by compressing three numbers into three coefficients has exactly three views of itself available, and asking which of them it is most nearly a function of is asking which compression lost least.
A correlation of a half, and neighbours differing two hundredfold
There is one more thing in the table and it sharpens the original finding.
The second ionisation energy’s rank correlation with the capacity is −0.5033. That is a moderate correlation, and it is negative — so the relation that a rising straight line was fitted to, explaining 2.9 per cent, is in fact a falling monotone-ish relation that a straight line was the wrong shape for. Both facts are consistent: a rank correlation says something about ordering and an R² says something about a particular functional form, and the two can disagree completely. The same distinction appears in a ranking is not a difference — there about four scales that agree on order and not on magnitude, here about one relation that has an order and no usable magnitude.
But the correlation is not rescued by being restated in ranks, because a rank correlation of one half sits alongside two neighbours whose capacities differ by a factor of 193. Plotting the two against each other for all six inputs shows they do not simply trade off: the electron affinity correlates weakly and has a bad pair, the second ionisation energy correlates moderately and has by far the worst.
The electronegativity and the first ionisation energy are also worth a glance here, because they return the identical ratio of 24.2 on the identical pair. Nitrogen and oxygen are neighbours in both orderings, which is not a coincidence — Mulliken’s electronegativity is built from the ionisation energy and the affinity, so the two inputs are close to being one input twice. Two of the six candidates are not independent, and the test says so without being asked.
That is the sharpest form of the original point. A correlation is a statement about the average behaviour of an ordering. An account is a statement that constrains individual cases. Nitrogen and potassium are the demonstration that the second gives out long before the first does.
What was computed, and how
Fourteen atoms have all six inputs and a finite positive capacity. For each input the atoms are sorted by it, every adjacent pair is taken, and the largest ratio of the two capacities is reported along with the median across all thirteen adjacent pairs. The control is the same procedure with the capacity as the input; the scale-dependent version is retained and reported precisely because it fails.
The test requires eight things. That there are enough atoms to run the test. That the second ionisation energy returns nitrogen and potassium at a ratio above a hundred, so this is the same test as the original. That every input has a several-fold pair, which is the answer. That the questioned input ranks last, and by more than a factor of five over the best. That the best input is one of the cubic’s own coefficients, and that the three coefficients rank ahead of the three data with no interleaving. And the refusal: that the raw-scale version of the test returns a ratio many times the rank version’s for the capacity against itself, so the defect in the question is exhibited rather than quietly fixed.
That last one is the check worth having. A repair made silently looks like the test having always been right.
Where the model stops
Fourteen atoms is few, and the floor of 4.95 is the price of that. On forty atoms the floor would fall and the differences between the inputs would be more finely resolved; nothing here says how the ordering would change, only that the spread between best and worst is currently twenty-four times and the floor is five.
The capacity is also a property of a cubic fitted to four points, and that construction gives lithium a negative electronegativity and two alkali metals no solution at all. Those three atoms are still in this sweep — lithium, sodium and potassium hold the three smallest capacities — so the ordering that the whole test is run against has its bottom end supplied by the part of the model that is least defensible. The capacity of an atom is a real idea; this model’s version of it carries every defect that fit carries, and the cubic coefficient — which is the highest derivative of a fit to four points — carries the most of all. That the capacity is closest to being a function of that coefficient is therefore not reassuring.
And “adjacent in rank” is a coarse notion of closeness. It ignores how far apart the neighbours actually are, so an input with one large gap in the middle of its ordering is treated the same as one that is evenly spaced. The gap is reported alongside each worst pair for exactly that reason, and for the second ionisation energy it is 2.02 electronvolts out of a range of 16.35 to 75.64 — three and a half per cent of the range, so genuinely close and not an artefact of ranking. The median adjacent pair in that input differs in capacity by 3.61, which is below the floor, so the worst pair is a genuine outlier within its own input rather than the typical case.
The generalisation
The transferable part is the control, and specifically that it was run at all.
A test of the form “does anything in this list explain the quantity” is only as good as its behaviour on the quantity itself. This is the same discipline as insisting that a difference does not make a transfer — checking that an instrument returns the trivial answer on the trivial case before believing it on a hard one. Running the capacity against the capacity costs one line and it caught a defect that would otherwise have propagated into every number in the table — and the defect was not in the arithmetic, which was correct, but in the reading of the word closest. That is the kind of error no check of correctness catches, because everything is working exactly as written.
The second transferable part is the floor. A test where everything fails is uninformative unless what passing would look like is known, and here the same control that found the defect also supplies the best achievable score. Without it, six inputs all failing reads as a test too strict to use. With it, the ranking is legible: two inputs are within a factor of a few of the best possible, and one is thirty-nine times worse.
Who found it, and when
The cubic-through-four-points construction, the capacity it implies, and the pair test are constructions rather than quotations. The ionisation energies and electron affinities are quoted; everything derived from them here is computed. The idea that a chemical species has a limit to the charge it will accept is a real one and much older than this model of it, and that caution has not moved. Nothing here is evidence for or against that idea; it is evidence about one construction’s version of it, and specifically about which of its own inputs that version most resembles.
The second ionisation energy’s place at the bottom of the ranking does have a reading, though, and it is worth stating because it is the only positive thing here. The capacity is a property of the anion end of the curve — how much more electron the atom will take — and the second ionisation energy is a property of the far cation end, two electrons in the other direction. That they should be the least alike of the six is what a chemist would have guessed, and the original correlation was the surprise rather than this.
The habit of asking what a constructed quantity is a function of, rather than what it correlates with, is one worth applying to one’s own inventions as readily as to anybody else’s, and it is the reason the capacity is best kept at arm’s length.
What this establishes is narrower than the question that prompted it: not what the capacity is, but that nothing available inside the model is what it is, and that the suspected input is the least like it of the six.
Still open: a fifth point, and an input from outside
The obvious open question is the third derivative, which is untouched here. The capacity is finite or infinite according to the sign of the cubic coefficient, and that coefficient is the highest derivative of a fit to four points, so it carries whatever the fit’s error carries. Recomputing the capacity from five electron counts rather than four, and seeing which atoms change sign, would say how much of the ordering here is the quantity and how much is the fit’s resolution — and given that the cubic coefficient came first in this ranking, that is now the most load-bearing number in the construction.
The nearer question is an input from outside. Everything tested here is inside the cubic, so the search was constrained to fail from the start; the interesting version needs a quantity the cubic was not fitted to. The atomic radius is the obvious candidate, it is uncontroversial, and radii are tabulated for other purposes — but not on a scale all fourteen of these atoms share. Bringing one in, from a stated source, and running the same rank-adjacent test against it would be the first time this question was asked of something the answer could not have been built from.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The value that only exists in the bond — both name closed form, convention, electron affinity, electronegativity, hardness, ionisation energy
- Four quantities go and one question stays — both name electron affinity, electronegativity, ionisation energy, model limit
- No panel of this kind can find an exception — both name convention, electronegativity, empirical scale, model limit
- The rule is not the lever — both name convention, electronegativity, empirical scale, model limit
- A band gap is not a bond energy — both name electron affinity, ionisation energy, model limit
- A bond order between atoms that do not interact — both name closed form, convention, model limit
Named objects
A dashed tag is an object no other essay names yet.
Closed formConventionElectron affinityElectronegativityEmpirical scaleHardnessIonisation energyModel limit