Beyond the octet

A second criterion left a gap too

The localisation spreads are bimodal — an empty factor of two hundred around the degeneracy threshold — and the gap might belong to the criterion rather than to the cages. Boys localisation leaves a gap of seven decades on the same forty-eight pairs, so it belongs to the cages. But the two criteria disagree about six of them, in both directions, and the family's most ambiguous cage under one is exactly degenerate under the other.

Worth reading first: Counting was right except where it mattered · Fifty descriptions of one molecule.

A check on whether the number of distinct bonding descriptions measures a cage’s ambiguity found something better than it went looking for: the number is not a good measure — a population of forty-four descriptions whose functional values agree to two parts in a hundred thousand is one answer found forty-four times. That is fifty descriptions of one molecule read the other way round, and the survey it was measured over is how many descriptions a cage has.

Along the way it measured the distribution of those spreads across all forty-eight cage-and-filling pairs it could compute, and the distribution turned out to be bimodal. Thirty-five pairs have a spread of exactly zero. Of the thirteen with a non-zero spread, two fall below 2.2 × 10⁻⁵ and eleven above 4.5 × 10⁻³, and between those there is nothing — an empty factor of two hundred and five with the degeneracy threshold sitting inside it.

That made the threshold robust: any value in two decades classifies all forty-eight identically, so a line inherited from a single molecule turned out not to matter. And the essay closed by naming what that argument rests on. The spreads are differences in one criterion’s values, and the bimodality is a fact about that criterion on these cages. A criterion that separated descriptions more finely could fill the gap in.

This runs a second criterion.

What a second criterion is

The functional used throughout is Pipek–Mezey: maximise the sum over orbitals and sites of the fourth power of the coefficient, which is a sum of squared atomic populations. It scores a description by how concentrated each orbital is on few atoms, and it knows nothing at all about where the atoms are.

Boys is the other classical choice and it is genuinely different. It maximises the sum of the squared distances of the orbital centroids from the origin — a statement about geometry rather than about populations. On a Hückel wavefunction the centroid of an orbital is the population-weighted average of the vertex positions, so the functional is cheap, and the cage’s coordinates enter the calculation for the first time.

Two criteria that use different information about the same wavefunction is what the question needs. The wavefunctions themselves are the ones a cage needs one pair more than it has corners counts electrons for, unchanged. If both are bimodal, the bimodality is the cages’. If Boys fills in the gap, the robustness was Pipek–Mezey’s coarseness.

Boys leaves a chasm

Both criteria leave a gap, and Boys leaves a chasm. Every non-zero relative spread under each criterion, on one logarithmic axis, with the largest empty stretch shaded. Pipek–Mezey's runs a factor of 206; Boys's runs 2.8e+7 — seven decades, from numerical zero to a real spread with nothing between. So the bimodality belongs to the cages rather than to the functional, and under the second criterion the threshold matters even less.
Fig. 1 Every non-zero spread under each criterion on one logarithmic axis, with the largest empty stretch shaded. One is a factor of two hundred; the other is seven decades.

Pipek–Mezey’s gap is a factor of 210, from 2.1 × 10⁻⁵ to 4.5 × 10⁻³, reproduced here at forty starts rather than the earlier two hundred — which makes it a check rather than a restatement.

Boys’s gap is a factor of 3.2 × 10⁷, from 9.7 × 10⁻¹⁴ to 3.1 × 10⁻⁶.

That lower figure is machine precision, so under Boys the split is not merely bimodal but categorical: a cage’s descriptions are either identical to the last bit the arithmetic holds, or they differ by parts in a million or more. There is no intermediate case anywhere in the family.

So the answer is that the bimodality belongs to the cages. Two criteria using different information about the same wavefunctions both find the family divided into degenerate and genuinely-different populations with nothing between, and the second finds the division more sharply than the first.

The threshold question, which was the reason for asking, is settled a fortiori. If a line anywhere in two decades classifies identically under Pipek–Mezey, a line anywhere in seven does under Boys.

And they disagree about which cages

The same forty-eight cages, measured two ways. Each cage-and-filling pair's relative spread under Pipek–Mezey against its spread under Boys, both logarithmic, with each criterion's own degeneracy line drawn through its own gap. Points in the two shaded quadrants are pairs the criteria classify differently — there are 6 of them, going both ways. Everything on the floor or the left edge is a spread of exactly zero, plotted at the axis limit.
Fig. 2 Each pair’s spread under one criterion against the other, with each criterion’s own line drawn through its own gap. The shaded quadrants are the disagreements.

That would be a tidy result and it is not the whole one.

Placing each criterion’s degeneracy line in its own gap — which is the only defensible way to draw them, since the gaps are seven decades apart — and comparing the classifications pair by pair, the two agree on forty-two of the forty-eight and disagree on six.

Six pairs the two criteria classify differently, and they go both ways. Every cage-and-filling pair the two criteria disagree about, with each one's relative spread and verdict. Two are degenerate under Pipek–Mezey and not under Boys; four are the reverse. A disagreement running one way would mean one criterion is simply coarser; running both ways it means they are looking at different things.
Fig. 3 The six pairs the criteria classify differently, with both spreads and both verdicts.

The disagreements go both ways. Two pairs are degenerate under Pipek–Mezey and not under Boys; four are the reverse. That direction matters: if one criterion were simply the other with more resolution, every disagreement would run the same way — the finer criterion would split populations the coarser one merged, and never merge ones it split. Both directions means the two are looking at different things and each sees structure the other misses.

Which is what one should expect from what they measure. A set of descriptions can be identical in how concentrated its orbitals are and differ in where those orbitals sit, and it can be the reverse.

Nearly the same count, and not quite the same cages. How many of the 48 pairs each criterion calls degenerate, and how many the two classify identically. The counts are close — 37 and 39 — which would be easy to read as agreement. They agree on 42 pairs, so the near-equality of the totals is partly two disagreements cancelling.
Fig. 4 How many pairs each criterion calls degenerate, and how many they classify identically. The totals are closer than the classifications.

The totals are 37 and 39 of 48, which look like near-agreement and are partly two disagreements cancelling — the same reading error as taking a count of descriptions as a measure of ambiguity. Comparing classifications by their totals is the same mistake one level up.

The extreme case flips

The family's most ambiguous cage is unambiguous under the other criterion. The pair with the largest Pipek–Mezey spread in the whole family — 11 vertices, 20 electrons — measured both ways. Under Pipek–Mezey its descriptions differ by 5.3 per cent, the widest anywhere. Under Boys they are identical. Whatever the headline case was showing, it was not a property of the cage independent of how the descriptions were scored.
Fig. 5 The pair with the largest Pipek–Mezey spread in the family, measured both ways.

One of the six disagreements is worth more than the others.

The eleven-vertex cage at twenty electrons has the largest Pipek–Mezey spread anywhere in the family — 5.3 × 10⁻², the case singled out when asking what a spread of five per cent means.

Under Boys its descriptions are identical, to machine precision.

So the family’s most ambiguous cage by one criterion sits on the floor of the other. Whatever that cage was showing, it was not a property of the cage independent of how its descriptions are scored. A reader who took away “the eleven-vertex cage at twenty electrons is where the bonding is genuinely ambiguous” would be taking away a statement about Pipek–Mezey. That is the same hazard as reading a count for a spread, which is the correction to reading a count, arriving one level further up.

Every pair with a spread under either criterion. The relative spreads and verdicts side by side for every cage-and-filling pair that is not exactly degenerate under both. The two columns of verdicts are mostly the same and not entirely, and the rows where they differ run in both directions.
Fig. 6 Every pair not exactly degenerate under both, with the two spreads and the two verdicts.

Why the two criteria can disagree at all

It is worth spelling out what a disagreement means physically, because the two functionals are not merely two scoring conventions and the difference has content.

Pipek–Mezey asks how concentrated each orbital is: an orbital spread over three vertices scores worse than one on two, whatever the geometry. Boys asks how far from the centre each orbital’s charge sits: an orbital displaced towards the surface scores better than one near the middle, whatever its populations.

On a cage those come apart in a specific way. Two descriptions can put the same amount of charge on the same numbers of vertices — identical populations, identical Pipek–Mezey value — and place those vertices differently around the cage, so that one description’s orbitals sit further out than the other’s. Boys sees that and Pipek–Mezey cannot. Equally, two descriptions can have orbitals whose centroids coincide by symmetry while differing in how many vertices each spreads over, and then the reverse holds.

So the four pairs that are degenerate under Boys and not under Pipek–Mezey are cages whose alternative descriptions differ in concentration but not in position, and the two that go the other way differ in position but not in concentration. That is a statement about the cages, and it is the reason a disagreement running in both directions was the expected outcome rather than a surprise once the question is asked properly.

It also says why “which criterion is right” is the wrong question. Neither is an approximation to the other, and neither is measuring ambiguity as such; each is measuring one aspect of how two descriptions differ. A cage that both call degenerate has descriptions that are the same in both respects, which is a stronger claim than either criterion makes on its own — and there are forty-two of those.

What was computed, and how

Forty-eight cage-and-filling pairs, from six to twelve vertices at every even electron count, each localised from forty random starting rotations of the occupied space under each of two criteria, and the descriptions grouped by the ones they land on. The Pipek–Mezey half is the earlier survey with fewer starts, unchanged so that this extends that survey rather than a neighbouring one — the same discipline as three shapes from one search. The Boys half is new.

The Boys sweep finds each pair rotation’s angle by a bracketed line search over the quarter period rather than in closed form. Boys has a closed-form Jacobi angle and it is easy to derive subtly wrong — this collection’s Pipek–Mezey implementation carries a comment recording what a mis-derived pair rotation looked like when it happened, which was a charge density that moved. A one-dimensional maximisation over an angle, using the functional itself to score, cannot fail in that way, and at these sizes it costs nothing.

One thing the Boys survey produces is not usable and is worth saying so rather than quietly not using. Its count of distinct descriptions is meaningless: the descriptions are keyed on participation numbers to three digits, and a line search reaches its maximum with enough numerical scatter that two starts finding the same description can disagree in the third digit. It routinely reports as many descriptions as there were starts. The functional values are unaffected, because a spread is a difference of maxima, so every comparison here reads spreads and none reads that count.

Six results are checked. Two are that both criteria run on essentially every pair and that enough have non-zero spreads for a distribution to exist. Two are that each distribution is bimodal, with Pipek–Mezey’s stated separately so that the earlier finding has to be reproduced before the new one counts. One is that Boys’s gap is much the wider. And the last is that the criteria disagree on some pairs and agree on most — both halves, because agreement everywhere and agreement nowhere would each pass a check of only one direction.

Where the model stops

Hückel on a cage: one orbital per vertex, one hopping integral, no repulsion. Both criteria are applied to the same wavefunctions, so nothing here compares the criteria as approximations to anything real — only to each other.

Forty starts rather than two hundred. That biases the count of descriptions downward and does not bias a spread much, since a missed basin is usually a rare one near an already-found one. The Pipek–Mezey gap comes out at 210 against 205 at two hundred starts, which is the check that the reduction has not moved anything material.

Two criteria is two. Edmiston–Ruedenberg is the third classical choice and needs two-electron integrals a Hückel model does not have, so it is out of reach here rather than merely unrun — the same boundary where two-centre bonding stops runs into when it needs more than adjacency. What can be said is that two criteria using disjoint information both find the family bimodal; a third might not.

And the degeneracy lines are placed in each criterion’s own gap, which is the right thing to do and is also a choice. It means the two classifications being compared are each defined by their own distribution rather than by a common absolute threshold — the alternative, which would compare them at the same numerical value, is meaningless when the two functionals have different units and different scales.

The generalisation

The transferable point is that reproducing a shape across methods is stronger evidence than reproducing a number.

The first finding was a gap. This finds a gap under a criterion that shares no information with the first — populations against positions — and the two gaps are in different places, of very different widths, and separate different sets of cages. None of the numbers agrees with any other. What agrees is that both distributions are bimodal, and that is what the finding was about.

Had the two criteria produced similar numbers, the natural worry would have been that they are not independent. That they produce wildly different numbers with the same structure is the more convincing outcome, and it is worth knowing that in advance, because a replication that agrees too closely is evidence of a shared assumption rather than of a shared truth.

The second point is about where the agreement stops. Establishing that the classification is real does not establish that any particular case belongs to it, and here both results come at once: the family divides, and six of the forty-eight fall on a side that depends on the measure. A finding stated as “these cages are ambiguous and those are not” is stronger than the evidence; stated as “the family divides sharply, and which side a given cage falls is criterion-dependent for an eighth of them” it is exactly as strong as the evidence.

What can now be said about ambiguity

Three calculations have now been spent on one question — how ambiguous is a cage’s bonding description — and it is worth setting out what survives all three, because each corrected the one before it.

The first counted descriptions and found a cage with fifty. The second showed the count is the wrong measure, since most of those fifty are one answer reached repeatedly, and replaced it with the spread in the functional. This one shows the spread is a real division of the family under two independent criteria and that six of the forty-eight fall on a criterion-dependent side of it.

What is left standing is narrower than any of the three essays’ headlines and is not weaker for it. A cage’s alternative descriptions are either the same or clearly different, with nothing in between, under either way of scoring them. That is the finding, it is about a distribution rather than about any cage, and it needed all three to state.

What is not left standing is any claim of the form “this cage is the ambiguous one”. Every such claim so far has been overturned by the next calculation — the fifty-description cage by the spread, the largest-spread cage by the second criterion. That is three for three, and the honest reading is that the method is good at measuring the family’s structure and has not yet found a way to say anything reliable about an individual member of it.

That is worth carrying as a warning rather than as a defeat. A distributional result and a claim about one case are different kinds of statement, and an argument that keeps producing the first while its headlines report the second will keep having to correct itself.

Who found it, and when

Boys localisation is Foster and Boys’s, Pipek–Mezey’s is theirs, and the ambiguity of bonding descriptions in electron-deficient cages is older than either. The survey, both distributions, and the comparison are new here.

The question came from a clean, robust result — a classification insensitive to its own threshold — whose robustness might have been an artefact of the one criterion used. Nothing forced that doubt; the argument was complete without it.

Still open: what the six disagreements are made of

The obvious open question is the six disagreements. Each is a cage whose descriptions are identical by one measure and different by the other, which means the descriptions differ in exactly the information one criterion uses and not the other — same populations, different centroids, or the reverse. Looking at one such pair of descriptions directly, rather than at their functional values, would say what that difference is in bonding terms, and it is the sort of thing a figure can show.

The nearer question is whether the eleven-vertex case’s flip has a symmetry explanation. A cage whose descriptions have identical centroids but different populations is a cage where the descriptions are related by an operation that permutes the vertices while fixing the centre of each orbital — which is what a symmetry operation of the cage does. Those operations are computed here, and checking whether the flipped cases are precisely the ones with a large enough symmetry group would connect these disagreements to the symmetry question raised about basin populations, which has not yet been run.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationDegeneracyLocalisationModel limitMulticentre bonding