Beyond the octet

Six disagreements and three calculations

Two localisation criteria classified six of the family's forty-eight cage-and-filling pairs differently, and three of the six were consecutive fillings of one icosahedron with spreads identical to three figures. They are not three coincidences and not one degeneracy being filled: they are one calculation, because the survey's forty-eight rows are thirty-three distinct questions.

Worth reading first: The cage is on both sides · A second criterion left a gap too.

An earlier essay ended with six cages on both sides of a line — six cage-and-filling pairs whose alternative localised descriptions are identical under one criterion and different under the other — and with a small observation it could not follow up. Three of the six were consecutive fillings of the same icosahedron, at fourteen, sixteen and eighteen skeletal electrons, and their relative spreads agreed to three significant figures under both criteria. That is not what three separate coincidences look like, and the reading offered was that one degeneracy was being filled across them.

The spreads agree because the three rows are the same calculation. Nothing is being filled across them at all.

Forty-eight rows, thirty-three calculations. Every cage-and-filling pair the family survey covers, one cell per pair, with cells that hand the localisation the same set of orbitals joined. A localisation here selects its occupied orbitals by asking which have any occupation at all, and Hund's rule puts one electron into each member of a degenerate shell before pairing any of them — so adding two electrons to a half-filled shell pairs a spin and changes nothing the search can see. Fifteen of the forty-eight rows repeat an input already in the table.
Fig. 1 Every cage-and-filling pair the family survey covers, with the pairs that hand the localisation the same set of orbitals joined.

What a row of the survey is

A localisation takes an occupied space and finds the unitary mixing of it that makes the orbitals as local as some functional says. The occupied space is the input and everything else is apparatus, so the question is how each row of the survey chooses one.

It asks which orbitals have any occupation at all. That is the natural implementation and it is what both criteria here do: gather every orbital whose occupation number is greater than zero, mix them, maximise.

Then the electron count is fed in and the orbitals are filled by Hund’s rule — one electron into each member of a degenerate shell before any of them is paired. So a five-fold shell reached by a cage with nine electron pairs’ worth of room receives one electron, then two, then three, four, five, and only then does it start doubling up.

Which makes the selected set a staircase. An orbital holding one electron is selected exactly as an orbital holding two is, so every electron that goes into a shell for the first time widens the selected set and every electron that pairs one leaves it alone.

Two more electrons, and the search sees nothing. For the twelve-vertex cage: the number of electrons against the number of orbitals the localisation treats as occupied. The line is a staircase because Hund's rule fills a degenerate shell singly before pairing, and an orbital holding one electron is selected exactly as one holding two. The flat treads are where several electron counts become one calculation — and the widest of them is three rows long.
Fig. 2 For the twelve-vertex cage, the number of electrons against the number of orbitals the localisation treats as occupied.

On the icosahedron the shell structure is one orbital, three, five, three. Nine orbitals are selected at fourteen electrons — one, three, and all five of the third shell singly or doubly occupied. Sixteen electrons pairs two of those five and selects nine. Eighteen pairs the rest and selects nine.

Three rows, one input, nine orbitals each time. The spreads are identical because the arithmetic is identical.

Where the treads fall, cage by cage

The staircase is a property of each cage’s shell structure, so it is worth reading the five spectra rather than generalising from one.

The six-vertex cage — an octahedron — has shells of one, three and two. The nine-vertex cage has eight shells, six of them singletons. The ten-vertex cage has six, the eleven-vertex cage eight, seven of which are singletons or pairs. The icosahedron has four shells and they are the widest in the family: one, three, five, three.

A tread is two electron counts wide wherever a two-fold shell is entered, and wider only where the shell is wider. So the six-, nine-, ten- and eleven-vertex cages contribute treads of exactly two rows each — the nine-vertex cage’s three of them at four and six, ten and twelve, sixteen and eighteen electrons — and the icosahedron contributes one tread of three, at fourteen, sixteen and eighteen, because that is the only five-fold shell in the family.

Which is why the group that was noticed was the icosahedral one. It was the only group of three, and a group of three is what looks like a pattern; the seven groups of two sitting beside it in the same table read as adjacent rows agreeing, which is not surprising enough to write down. The defect was visible everywhere and legible only in one place.

There is one more consequence and it runs the other way. A tread of three costs the icosahedron four of its twelve rows, against two of six on the octahedron and three of nine, ten and eleven on the others — so the collapse is not uniform, and the cage whose shells are widest loses the most. Every count the family has produced over rows is therefore weighted towards the cages with narrow shells, which are the ones with least degeneracy, which is the quantity the whole survey is about.

How much of the table repeats

Once the mechanism is named the count is arithmetic on five spectra, and it is not a small correction.

What each cage contributes, counted twice. Per cage: the electron counts the survey enumerates, how many distinct orbital sets those counts actually produce, how many fillings leave a degenerate shell partly occupied, and how many inputs the two criteria classify differently. The second column is the one every count should have been taken over, and it is smaller than the first on every cage.
Fig. 3 Per cage: the electron counts enumerated, the distinct orbital sets those produce, the open-shell rows, and the criterion disagreements counted both ways.

The six-vertex cage’s six fillings are four inputs. The nine-vertex cage’s nine are six, the ten-vertex cage’s ten are seven, the eleven-vertex cage’s eleven are eight, and the twelve-vertex cage’s twelve are eight.

Forty-eight rows, thirty-three calculations. Fifteen rows are a question already asked, and every count quoted over that table — the fraction degenerate, the ranking by count against the ranking by spread, the eleven of thirteen that had anything to count, the six disagreements — is a count with repeats in it.

That does not make any of those findings wrong. A proportion computed over a table with duplicates is still a proportion of something; it is a proportion over rows rather than over questions, and the two differ by however unevenly the duplicates fall. They fall unevenly here: the twelve-vertex cage loses four of its twelve rows and the six-vertex cage loses two of six, so the larger cages are over-weighted in every count the family has produced.

Two of those counts are worth recomputing, because they are the ones later essays leaned on.

The fraction of the family whose descriptions are all one answer was 37 of 48, or 77 per cent. Over inputs it is 27 of 33, or 82 per cent. The finding — that degeneracy is common rather than exceptional — survives and gets slightly stronger, which is what the uneven collapse predicts: the treads sit inside wide shells, wide shells are where degeneracy lives, and removing duplicates removes them mostly from the degenerate side.

The two readings of ambiguity still disagree about which case is worst. By number of descriptions it is the nine-vertex cage at eight electrons; by spread in the functional it is the eleven-vertex cage at twenty. Neither of those sits in a tread, so the disagreement between the two rankings is not an artefact of the repeats, and the essay that established it is untouched.

So the recount changes one finding and leaves two standing. That asymmetry is the useful part: a duplicate hurts a count of cases and barely touches a count of proportions, because a proportion is distorted only by however unevenly the duplicates fall, and here they fall in a direction that happens to be conservative.

The check nobody designed

There is a second reading of fifteen duplicated rows and it is the better one.

Fifteen rows that had to agree, and do. Every repeated input, with the two criteria's relative spreads for each of the rows that map onto it. Two rows that are the same calculation must return the same numbers to every digit, so a disagreement anywhere in this table would be the search reporting a different answer to an identical question. There is none — which makes the defect its own test of the survey's determinism, and the only one the survey has.
Fig. 4 Every repeated input, with the two criteria’s relative spreads for each of the rows that map onto it.

Two rows that are the same calculation must return the same numbers. Not close numbers — the same numbers, to every digit, since the search is deterministic given its input and both rows hand it the same one. A disagreement anywhere in that table would be the search returning different answers to an identical question, which is a defect of a different and much worse kind.

There is none. Fourteen repeated inputs covering twenty-nine rows, and every group agrees exactly under both criteria.

So the survey has a test of its own determinism and did not know it had one. That matters because the basin search is a stochastic procedure — it starts from pseudo-random rotations of the occupied space and reports what it lands on — and nothing else here pins down that its answers are reproducible rather than merely stable-looking. The count of descriptions was chased to four thousand starts to find out when it stopped rising, which is a test of sufficiency and not of determinism.

A duplicate is a defect in a count and a control in a calculation. This table is both, and the second is worth more than the first cost.

Six disagreements, three of them

With the inputs identified the earlier headline can be recounted.

Six disagreements are three cases. Each case where the two localisation criteria classify an input differently, with the electron counts that map onto it. The earlier survey counted six because it counted rows; the three icosahedral rows at fourteen, sixteen and eighteen electrons are one set of nine orbitals and the two ten-vertex rows at six and eight are one set of four. The asymmetry that reading found — four one way and two the other — is two and one.
Fig. 5 Each case where the two localisation criteria classify an input differently, with the electron counts that map onto it.

The three icosahedral rows are one input of nine orbitals. The two ten-vertex rows, at six and eight electrons, are one input of four. The eleven-vertex row at twenty electrons is alone.

Three cases, not six. And the direction of the asymmetry survives: two of the three are degenerate under Boys alone and one under Pipek–Mezey alone, where the earlier survey read four and two.

That reading was used for something. It proposed a trend — the Boys-degenerate cases are the larger cages and the Pipek–Mezey ones the smaller, so a bigger cage has more room for populations to differ at fixed centroids than the reverse — and closed by saying it “may be a real trend, or it may be four cases and two”. It is two cases and one. The hedge was the right one to write and the number it was hedging against is now three times smaller, which puts the question past what an extension of the survey to one more cage size could settle.

What the explanation was, and why it fails

The mechanism proposed for the icosahedral group was a degenerate shell being partly occupied, and it is worth testing directly rather than displacing. It is a reasonable idea: the members of a degenerate shell are interchangeable, a half-filled one is where a wavefunction has the most freedom, and freedom is what a localisation’s ambiguity is made of.

A partly filled shell neither causes the disagreement nor follows from it. The forty-eight rows sorted two ways at once: whether the filling leaves a degenerate shell partly occupied, and whether the two criteria classify it differently. The explanation the earlier survey proposed predicts that the two lower-left and upper-right boxes are empty. Neither is: most open-shell fillings agree, and half the disagreeing rows have a closed shell. The grouping it saw is real; the reason it gave is not.
Fig. 6 The forty-eight rows sorted by whether the filling leaves a degenerate shell partly occupied and by whether the two criteria classify it differently.

Twenty-one of the forty-eight fillings leave a shell partly occupied. Three of those twenty-one are among the disagreeing rows, so the condition holds for eighteen rows where the criteria agree perfectly.

And three of the six disagreeing rows are closed-shell. The ten-vertex case at eight electrons has its HOMO at the top of a filled two-fold shell and its LUMO on a level of its own; the eleven-vertex case at twenty electrons has a non-degenerate HOMO; the icosahedron at eighteen electrons closes its five-fold shell exactly.

So the condition neither implies the disagreement nor follows from it. Half the disagreements do not satisfy it and six sevenths of the rows that satisfy it do not disagree. The grouping that was seen was real and the reason offered for it was wrong, which is a distinction worth keeping: what it observed was that three numbers were identical, and what it inferred was a mechanism. Only the observation had evidence behind it.

The right cause was one level lower down than either — not in the physics of a filled shell but in how a row of the table becomes an input.

What was computed, and how

Every occupied set is read from a Hückel calculation on a deltahedron built by minimising repulsion on a sphere, the same construction the neighbouring essays use. The shells are runs of eigenvalues equal to a part in a billion. The occupied set is the list of orbitals whose occupation is greater than zero, which is what both localisation criteria select on, so it is read from the same quantity rather than from a parallel rule.

Nine things are checked. That the survey has forty-eight rows, so the recount is over the table already built. That it has fewer inputs than rows, with the count stated. That every group of rows sharing an input agrees about how many orbitals that input has, which is the arithmetic the grouping rests on. That no repeated input disagrees with itself under either criterion — the determinism check above, checked rather than noticed. That the disagreements are fewer inputs than rows, and that there are three. That they still run in both directions, so the earlier qualitative finding survives. That the asymmetry is two against one rather than four against two. And that the open-shell condition fails as an explanation in both directions, with each direction stated as its own test so that either could fail alone.

Where this stops

None of this touches the criteria themselves. Both functionals leave a gap in their spreads, which is the finding that made the classification worth trusting, and a gap in a set of numbers is unaffected by some of those numbers being repeats — a duplicated value does not fill an empty stretch. So degenerate remains a property of these cages and not of the functional used to look at them.

And the physics of the icosahedral case is not resolved by finding out that it is one case. Nine orbitals on a twelve-vertex cage genuinely do have descriptions that Pipek–Mezey separates and Boys does not, and why remains open. What has changed is that it is one fact needing one explanation rather than three facts needing a common cause.

The step function is a modelling choice and it may be the wrong one. Selecting by non-zero occupation treats a singly-occupied orbital as fully occupied, which is what makes three electron counts one input. A localisation whose functional read the density would separate them: the three icosahedral fillings have different density matrices — their traces are fourteen, sixteen and eighteen by construction, and their site populations differ, uniform at one and a half per site only in the closed-shell case. Whether a density-weighted localisation of an open shell is the right object is a question about what a localised orbital means when some orbitals hold one electron, and nobody here has asked it.

So the defect admits two repairs with different meanings. Counting the table over inputs is bookkeeping and is done here. Making the open-shell rows into distinct inputs would be a change to what is being localised.

The generalisation

The habit is to ask what the independent variable of a survey actually is, rather than what it was indexed by.

This table was built by two nested loops, over cages and over electron counts, and every count taken from it has been a count over the cells those loops produced. The loops are not wrong; the assumption that distinct loop variables give distinct calculations is, and it fails here through a step function nobody put there on purpose — Hund’s rule on one side and a selection rule of any occupation at all on the other, meeting in the middle.

The tell was available and was written down. Three numbers identical to three significant figures is not a coincidence and is not usually a mechanism either; it is most often the same arithmetic twice. That essay saw the identity, recorded it, and reached for a physical cause, which is the natural move and the wrong first one. Identical outputs are evidence about inputs before they are evidence about anything else.

The corollary is the cheerful half. A survey with duplicates in it is a survey carrying a free control, and the control tests something no deliberate check here does. Finding the duplicates is what turns them from a silent error in a denominator into the only evidence that a stochastic search is reproducible at all.

Who found it, and when

Hund’s rule is 1925 and Pipek and Mezey’s criterion is 1989; Boys’s is 1960. Nothing about the filling rule or either functional is new here. What is computed is the map from electron count to selected orbital set on five deltahedra, the grouping of the survey’s rows by that map, the consistency of every repeated group, and the recount of the disagreements reported earlier.

The number worth carrying is not thirty-three. It is that fourteen repeated inputs agreed with themselves to every digit, in a table where nobody had asked whether they would.

Still open: whether the answer was determined in the first place

The obvious open question is the one the step function makes unavoidable. Selecting by non-zero occupation is what collapses three fillings onto one input, and it does something else at the same time: where the selected set takes only some members of a degenerate shell, the members it takes are distinguishable only by an orientation the eigenvalue routine chose. A rotation inside that shell is not the occupied–occupied mixing a localisation is invariant to; it changes which combinations are occupied, and so it changes the density being localised. How many of the thirty-three inputs have a boundary falling inside a shell is arithmetic already done above, and whether the localisation’s answer moves when the orientation is re-chosen is a handful of extra searches.

The nearer question is the trend that has just lost two thirds of its evidence. Two cases against one is not a direction anybody should defend, and the survey stops at twelve vertices because that is where this collection’s cage builder was last exercised. Thirteen and fourteen vertices are available, each contributes one deltahedron and a dozen fillings, and — counted properly this time — each would add about eight inputs rather than a dozen rows. Whether that is enough to turn two against one into a trend is a question worth pricing before running, because the answer may be that no number of cages settles it while the disagreements stay this rare.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ClusterDegeneracyElectron-deficient bondingFillingHückel theoryLocalisationModel limitMulticentre bondingUnderdetermination