What is taught wrongly

A correlation is not an account

A quantity called the chemical capacity is ordered against the thing it was supposed to predict and correlates instead with the second ionisation energy, which invites asking how much of it that accounts for — on the reasoning that a quantity which is ninety per cent of one input has a simpler name than the one it carries. It is three per cent of it. Nitrogen and potassium sit two electronvolts apart in the second ionisation energy and differ two hundredfold in capacity.

Worth reading first: A capacity that is largest where there is none · The table that could not have mattered.

A quantity defined from a cubic fit to an atom’s energy against its electron count — the chemical capacity, how many electrons an atom can take before its chemical potential turns round — is ordered against the thing it is supposed to be about. The three atoms that cannot hold an extra electron are among the four largest capacities; the three smallest capacities all bind their anions.

Having established what it is not, that essay noticed what it does track: the second ionisation energy, by rank correlation. Its closing paragraph proposed the natural next step and gave the reason for it — a quantity that turns out to be ninety per cent of one input is a quantity with a simpler name than the one it has.

It is three per cent of that input, and two atoms two electronvolts apart in it differ two hundredfold in capacity — which no function of one variable can do, whatever form it takes.

The capacity against the quantity it was supposed to be. The chemical capacity of 14 atoms against their second ionisation energy, with the capacity on a logarithmic axis because it spans nearly three orders of magnitude. The suspected relation is not there: N and K are 2.02 electronvolts apart in the second ionisation energy and differ by a factor of 193 in capacity, which no function of one variable can produce.
Fig. 1 The capacity against the second ionisation energy, on a logarithmic axis because the capacity spans nearly three orders of magnitude. The circled pair is the argument.

Why the question was worth asking

The suspicion behind it is a good one and it has paid off before. Electronegativity is not one quantity makes the same move: a word that names several different constructions is a word doing less work than it appears to, and the repair is to find what each construction actually measures. If the capacity had turned out to be the second ionisation energy rescaled, the honest response would have been to say so and use the shorter name.

The quantity no scale prints is the other precedent — a quantity that exists in the definitions and in no published table, which had to be computed before it could be examined. The capacity is the same kind of object, and the same discipline applies: it has to be tested against the simpler things it might secretly be.

So the closing paragraph was not being sceptical for its own sake. It named a specific alternative, gave a specific threshold, and made the test arithmetic. That is what makes the negative answer worth having.

What the correlation was, and was not

The rank correlation is real. Across the fourteen atoms with both quantities it is 0.503 in magnitude, and negative — a bigger second ionisation energy goes with a smaller capacity, which is the direction physical sense would predict, since an atom that holds its second electron tightly should be less willing to take more.

A rank correlation of a half is a visible relation. What it is not is an account. Rank correlation asks whether the orderings agree; it is insensitive to how far apart the values are, and it can be substantial when one quantity varies over three orders of magnitude and the other over a factor of five, as here. The question the closing paragraph asked — how much of the capacity is the second ionisation energy — is a question about variance, and the two statistics are not comparable.

A correlation is not an account. The capacity correlates with the second ionisation energy — by rank, 0.503 in magnitude, which is a real and visible relation. What a rank correlation does not say is how much of the quantity the other one accounts for, and the answer here is almost none of it. The two bars measure different things, and only the lower ones bear on the question the closing paragraph asked.
Fig. 2 The rank correlation beside the fractions of variance, with the level a quantity would have to reach before it deserved a simpler name.

Fitted directly, the second ionisation energy accounts for 2.9 per cent of the capacity’s variation. That is the form the proposal implies — a two-parameter fit of one to the other — and it explains essentially nothing.

This is not a subtle statistical point, but it is an easy one to walk past. A ranking is not a difference is the same distinction applied to the tables themselves: four scales can agree on every ordering and disagree about every difference computed from them. Here two quantities agree on much of their ordering and one accounts for almost none of the other’s variation, which is the same gap between an ordinal statement and a numerical one.

Five forms, and the best of them

The proposal did not name a functional form, and it would be unfair to test only the one that fails hardest. So five were tried: the direct fit, the reciprocal, and three with the capacity’s logarithm taken, since a quantity spanning three orders of magnitude is a natural candidate for one.

Five ways of asking, and none of them says most of it. The fraction of the capacity's variation each two-parameter form accounts for. The form the question implies — capacity against the second ionisation energy directly — explains 2.9 per cent. The best of the five explains 43.8, and needs a logarithm the question did not offer. The line at nine tenths is where a quantity would have to sit before it deserved a simpler name.
Fig. 3 The fraction of the capacity’s variation each two-parameter form accounts for, against the level the proposal’s own reasoning names.

The best is the capacity’s logarithm against the second ionisation energy, at 43.8 per cent. That is a real relation and it is less than half. It also costs something the proposal did not offer: the logarithm was chosen after looking at the data, and a transformation selected that way is not free — with five forms tried on fourteen points, the best of them is expected to look better than it is.

Even taken at face value, 43.8 per cent is not what “essentially is” means. A quantity that is essentially another quantity does not leave the majority of itself unexplained.

What the best fit still gets wrong, atom by atom. The residual of the best of the five forms — log capacity ~ a + b·I₂ — for each atom, in the units it is fitted in. A fit that had captured the quantity would leave a scatter without structure and without large members. These span 5.11 natural logarithms, which is a factor of 166 in the capacity itself — on a quantity the fit was supposed to account for.
Fig. 4 What the best fit still gets wrong, atom by atom, in the units it is fitted in.

The residuals of that best fit span 5.11 natural logarithms — a factor of 166 in the capacity itself. Whatever the second ionisation energy is doing here, the part it does not do is larger than the part it does.

What forty-four per cent would have to mean

It is worth pausing on the best form rather than dismissing it, because 43.8 per cent is not nothing and a reader is entitled to ask what it represents.

The relation it captures is the physically sensible one: atoms that hold their second electron tightly tend to have small capacities. Lithium, sodium and potassium have the three largest second ionisation energies among the alkali metals here and three of the smallest capacities. That much is real and the fit finds it.

What the fit cannot do is carry that tendency into a prediction. A relation accounting for 44 per cent of a quantity spanning three orders of magnitude leaves an uncertainty of more than an order of magnitude on any individual atom, which is not a prediction anybody would use. The tendency is worth knowing and the function is not worth having, and those are different conclusions that the single number 0.438 does not distinguish between.

The pair that settles it

A fit can always be argued about. Two atoms cannot.

Two atoms, one second ionisation energy, two hundred times the capacity. N and K have second ionisation energies of 29.60 and 31.63 electronvolts — 2.02 apart, closer than almost any other pair here. Their capacities are 31.542 and 0.164, a factor of 193. Any function of the second ionisation energy alone must give them nearly the same answer, so no such function can be what the capacity is.
Fig. 5 Nitrogen and potassium: their second ionisation energies, and their capacities.

Nitrogen’s second ionisation energy is 29.60 eV and potassium’s is 31.63 — 2.02 apart, closer than almost any other pair in the set. Their capacities are 31.54 and 0.164, a factor of 193.

Any function of the second ionisation energy alone, of any form whatever, must give two nearly equal inputs two nearly equal outputs. So no such function is what the capacity is, and this does not depend on which fit was tried or how many. One pair is enough, and the pair is not a marginal case: nitrogen has the largest finite capacity in the collection and potassium one of the three smallest.

There is one more objection worth answering, because it would make the result less interesting than it is: that the capacity’s three orders of magnitude are dominated by a few extreme atoms, and the fit is being judged unfairly on outliers.

It does not rescue the fit. Removing nitrogen, much the largest capacity, raises the direct form from 2.9 per cent to 28.5 — better, and still nowhere near an account. And the largest misses of the best fit are not all at the top of the range: they are nitrogen at +2.89, then potassium at −2.22, aluminium at −1.64 and sodium at −1.34, whose capacities are 0.16, 0.76 and 0.12. The fit misses the small ones as badly as the big one, in the other direction.

What the capacity is instead

Nothing here says what the capacity is a statement about, and it is worth being explicit that no answer has been found.

Every atom, both quantities. The second ionisation energy and the chemical capacity for each atom that has both, ordered by the first. Reading down the second column is the argument: it rises and falls without reference to the first, spans nearly three orders of magnitude where the first spans a factor of five, and puts its largest and one of its smallest values within two electronvolts of each other.
Fig. 6 Every atom with both quantities, ordered by the second ionisation energy.

Reading down that table, the capacity rises and falls without reference to the column beside it. Nitrogen at 31.5 sits between chlorine at 6.3 and potassium at 0.16. The quantity is built from three fitted coefficients of a cubic — a chemical potential, a hardness, and a third derivative — and the third of those decides where the potential turns round and hence whether the capacity is finite at all. A quantity whose value is set by the third derivative of a fit to four points is not obviously a statement about anything, and that possibility is not excluded by anything here.

What has been excluded is the specific simplification the closing paragraph proposed. The capacity keeps the name it has, for now, on the grounds that no shorter one has been shown to fit.

That is the shape where a closed form stops being one warned about from the other direction: a construction can be perfectly well defined and still be sensitive to something nobody intended it to depend on. The capacity is well defined; whether it is about anything is the open question, and it has been narrowed by one candidate rather than answered.

What was computed, and how

The capacity comes from the standard cubic construction: an atom’s energy at four electron counts, fitted, differentiated, and solved for where the chemical potential turns. Fourteen atoms have a finite capacity and a tabulated second ionisation energy, and those are the fourteen fitted.

Ionisation energies and affinities are quoted values; the capacities are computed from them. That division is the collection’s standing rule and it matters here more than usual — the whole argument is about what a computed quantity is a function of, so a computed quantity fitted against a second computed quantity would be measuring the two constructions against each other rather than against anything. The second ionisation energy is the one input in this essay that nobody here derived, which is exactly why it was a good candidate for the capacity to secretly be. A difference does not make a transfer establishes how carefully these quantities have to be kept apart. The fits are ordinary least squares, and the fraction of variance is the squared correlation between fitted and actual in whatever units the form is fitted in — which is why the logarithmic forms are not directly comparable with the linear ones, and why the direct fit is quoted separately as the one the proposal actually implies.

The refusal is the fitting arithmetic itself. A quantity regressed on its own logarithm must return unit slope and account for all of the variance; if that ever comes back as anything else, every other number here is worthless. It is the cheapest possible check and it is the one that would catch a transcription error in the fitter.

Where the model stops

A mean that is low rather than right is the methodological neighbour here — a quantity computed by a rule that is defensible in isolation and produces a systematically wrong answer in use. The difference is that there the rule’s failure was demonstrable against a reference; here there is no reference for the capacity to be wrong against, which is precisely why it is hard to say what it is.

Fourteen atoms is a small set and they are all main-group. The three with infinite capacity — beryllium, magnesium and phosphorus — are excluded because a fit cannot take them, and excluding them is not neutral: they are excluded for having the extreme behaviour, which is exactly the behaviour a relation would have to explain. A version of this question that could handle them would be a better test and would need a different formulation of the capacity.

The value that only exists in the bond is the other warning about atomic quantities: some of them are not properties of an atom at all, and a number computed per atom can be a number about the procedure. The capacity may yet be one of those, and ruling out the second ionisation energy does not rule that out.

The five forms are also five, not all. There is no theorem here that no two-parameter function of the second ionisation energy fits; there is a demonstration that five reasonable ones do not, and a pair of atoms that rules out every monotone one. The second is the load-bearing argument and the first is supporting.

And the capacity itself is a construction of this collection’s, built to make a definition precise enough to test. Nothing here says the construction is the one a reader would find elsewhere under the same words.

The generalisation

The habit this argues for is one line long: when a correlation is offered as evidence that one quantity is another, ask what fraction of the variance it accounts for, and check a close pair.

The two checks catch different failures. The variance fraction catches a relation that is real but small — the case here, where a genuine rank correlation of a half corresponds to almost no explanatory power because the two quantities live on wildly different scales. The close pair catches a relation that is real but not a function — two inputs that agree and two outputs that do not, which no fitted form can repair and which one line of arithmetic finds.

The close-pair check has a further virtue: it is immune to the choice of functional form, which is the part of a fitting argument that can always be disputed. A variance fraction depends on what was fitted and how the data were transformed. Two atoms with equal inputs and unequal outputs rule out every function at once, and they can be found by sorting.

Neither is expensive, and both are more informative than the correlation coefficient that prompted them. The reason they get skipped is that a correlation is usually reported as a single number with a familiar interpretation, and the familiar interpretation of “these correlate” quietly becomes “this is basically that” somewhere between the calculation and the sentence.

That slippage is what the ninety-per-cent proposal made explicit and testable, which is why the question was worth asking even though the answer is negative. A proposal that names its own threshold — ninety per cent — can be refuted by arithmetic. One that says two things are related cannot.

Who found it, and when

Chemical hardness and the conceptual density-functional quantities around it are Parr and Pearson’s, from the 1960s onward, and the practice of fitting an atom’s energy against a continuous electron count to obtain them is standard. The specific quantity here is a construction made for these essays rather than a standard one.

The distinction between a rank correlation and a fraction of variance is elementary statistics and belongs to nobody. It is restated because the step from one to the other is where the proposal would have gone wrong, and because the arithmetic that separates them took less time than writing this paragraph.

Still open: how much of the ordering is the fit

The obvious open question is the third derivative. The capacity is finite or infinite according to the sign of the cubic coefficient, and that coefficient is the least determined of the three — it is the highest derivative of a fit to four points, so it carries whatever the fit’s error carries. Recomputing the capacity from five electron counts rather than four, and seeing which atoms change sign, would say how much of the reported ordering is the quantity and how much is the fit’s resolution. That is the same construction with one more point and it would settle whether nitrogen’s 31.5 is a number at all.

The nearer question is the pair test, run over everything. Nitrogen and potassium were found by looking for the closest pair in the second ionisation energy with the largest ratio in capacity, and the same search can be run against every other candidate input — the first ionisation energy, the electron affinity, the hardness, the atomic radius. Each would return its own worst pair, and a quantity for which no input produces a close pair with a large ratio would be a quantity worth naming after that input. It is one loop over pairs on fourteen atoms, and it turns “the capacity is not the dication energy” into a statement about what it is not a function of at all.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Chemical hardnessCorrelationElectron affinityElectronegativityEmpirical scaleIonisation energy