Three shapes from one search
Worth reading first: Where the count stops being an effort · The answer a search is most likely to give.
A search on a twelve-vertex cage ran to four thousand random starts, found fourteen distinct descriptions with the last arriving at start 124, and reported one more thing about them. The fourteen fall into two groups: ten holding six to fifteen per cent of the starts each, four holding under one and a half, with a factor of five between the groups and nothing in the gap. A population that continued downward would show a tail; this showed a gap, and that is what made the count trustworthy.
It closed by asking whether the gap is general. If it is, then count the big basins and check for a gap is a stopping rule with a stated confidence, which is what every search of this kind in the collection lacks.
It is not general. Two more cases from the same family, the same search and the same code give two completely different shapes, and one of them is the shape that would mean the count was only a lower bound.
The family had almost nothing to compare
The first thing to check is the obvious one: run the same search on a cage of a different size.
At the filling this collection’s cage builder uses, a six-vertex cage has one localised description, a nine-vertex cage has one, a ten-vertex cage has one, an eleven-vertex cage has one, and only the twelve-vertex one has fourteen. So the proposed experiment returns a single description and no population at all — there is nothing to group and no gap to look for.
Nor is that an accident of these particular cages. A localised description is a way of dividing the occupied orbitals, and how many ways there are depends on how many occupied orbitals there are and how symmetric they are — the same reason four-centre bonding has more descriptions than three-centre bonding does. Changing the cage at a fixed rule for the filling changes both at once and does so unhelpfully.
The multiplicity is a property of the filling, not of the cage. Sweeping the electron count instead turns a null result into a survey, and the largest entry in it is not the twelve-vertex cage at all: it is a nine-vertex cage at eight electrons, which four hundred starts find fifty descriptions of and four thousand find sixty-three.
A smooth tail, and a search that has not converged
The nine-vertex cage’s sixty-three descriptions run from 6.70 per cent of the starts down to 0.025 per cent — one start in four thousand — and the largest step anywhere in that sequence is a factor of 2.00, at rank fifty-six, where the shares are 0.10 and 0.05 per cent and the counts are four and two. That is not a gap; it is two small integers.
So there are no groups to count, and the gap rule has nothing to say. What it would have said, if applied anyway, is the wrong thing: the top six basins are separated from the seventh by a factor of 1.20, so anybody drawing a line there would report six descriptions where there are at least sixty-three.
Worse, the search itself has not finished. Four thousand starts on the twelve-vertex cage leaves a 0.0 per cent chance of a miss; four thousand on the nine-vertex cage leaves 90.0 per cent. The same number of starts on the same code, and one answer is converged while the other is not, with nothing in the output to distinguish them except the miss probability.
One basin and a dust
The third case is a third shape and it is the one that looks safest and is not.
The twelve-vertex cage at twenty electrons has fifteen descriptions, and the first of them takes 97.63 per cent of the starts. The largest step is a factor of 79.7 and it falls at rank one. The other fourteen share the remaining 2.37 per cent, with the rarest reached once in four thousand.
Applied here, count the big basins and check for a gap returns one, with an enormous gap behind it and every appearance of certainty. The right answer is fifteen, and the miss probability at four thousand starts is 90.3 per cent — so a search that returned one description and looked completely converged has in fact almost certainly missed some of what it was looking for.
That case is also the one where two independent calculations agree 95.32 per cent of the time. On the nine-vertex cage they agree 3.51 per cent of the time. Consistency between two published localisations is evidence about the system, not about either calculation — and here the system that produces the most consistency is the one whose description count is most wrong.
What the three cases have that differs
The three differ in what the occupied set is. The twelve-vertex cage at twelve electrons stops in the middle of a five-fold degenerate group; the nine-vertex cage at eight electrons stops cleanly at a non-degenerate level; the twelve-vertex cage at twenty electrons stops in the middle of the same five-fold group from the other side.
So the shapes do not sort by whether the shell is closed. The case with the smooth tail is the closed-shell one, and the two with a degenerate frontier give the gap and the dominant basin. Whatever decides the shape, it is not the obvious candidate, and three cases cannot find it.
That is worth stating rather than dressing up. The question was whether a feature of one population generalises; the answer is that it does not, and the reason it does not is not established here.
The one thing all three agree about
Set against those three shapes, one fact holds in every case and is worth putting beside them: the best description by the functional is not always the one most often reached. That was the twelve-vertex cage’s other finding and it is confirmed here rather than tested — a search that returns its most common answer is returning the largest basin, and a basin’s size is a property of the landscape’s shape rather than of the maximum at its bottom.
The three shapes make that sharper in opposite directions. Where one basin holds ninety-eight per cent of the starts, a practitioner will find it every time and never see the alternatives; where sixty-three basins share the starts almost evenly, a practitioner will find a different answer on nearly every run and has no way to know which is best without keeping the functional values. Neither situation is visible from a single calculation, and both are visible from a hundred starts and a histogram, which nothing costs.
What does survive
One statement holds on all three, and it is the one the twelve-vertex search also gave: the probability that a search of a stated length has missed a description, computed from the basin sizes it observed.
It needs the populations and nothing about how they are grouped, so it is indifferent to which of the three shapes a case has. It says four thousand starts is enough on one of these and not enough on the other two, it says two hundred starts on the twelve-vertex cage has a 62.6 per cent chance of a miss, and it is available from any search that recorded which start landed where.
The practical rule that replaces the gap is therefore not a rule about shape at all. Record the basin populations, compute the miss probability, and quote it with the count. A count without it is a count whose confidence is unknown, and the three cases here show that the confidence ranges from complete to almost none across systems that a reader would put in one class.
It is also worth saying what this does not undermine. The count of fourteen on the twelve-vertex cage is right, its discovery curve is right, and its central claim — that a search which stops finding new answers has found them all only with a probability its own basin sizes give — is exactly what survives. What does not survive is the extra rule about gaps, which was always tentative.
One more thing the three cases settle, because it comes free with them. The twelve-vertex search’s rarest description held 0.60 per cent of four thousand starts — twenty-four of them — which reads as evidence that nothing rarer was hiding, since a population continuing downward would show a tail. Here two populations do continue downward: the nine-vertex cage’s rarest holds one start in four thousand and the twenty-electron case’s holds one as well. So a rarest basin of a few tenths of a per cent is not a floor, and the smallest basin a search finds is mostly a statement about how many starts it ran.
That last observation has a consequence for how the miss probability should be read, and it is the one place the formula can mislead. It is computed from the shares the search observed, so a description that was never reached contributes nothing — the estimate is a lower bound on the risk and not a confidence interval. On a case with sixty-three basins sharing the starts almost evenly, the shares are well measured and the bound is close to tight; on a case where one basin holds ninety-eight per cent, the remaining two per cent is measured by a few dozen starts and the tail below it by none at all. The formula is most trustworthy exactly where the answer it gives is least alarming, and least trustworthy where a miss is most likely.
The remedy is not a better formula. It is to say what the search was for. A practitioner who wants the best description by the functional needs the largest functional value and can stop when the value stops improving; a practitioner who wants to know how many descriptions exist needs coverage and cannot stop until the miss probability is small; and a practitioner who wants to know whether a molecule has a unique localised picture at all — which is the chemical question — needs neither, because the answer is the shape of the histogram and is legible from a hundred starts. Three different questions, three different stopping rules, and one number quoted for all three is how a gap rule comes to be offered.
What was computed, and how
Each search runs a localisation from a random unitary rotation of the occupied orbitals and records the resulting description by its orbitals’ centre counts, rounded to three places. Two starts landing on descriptions with the same sorted list of centre counts are counted as the same basin, a convention kept unchanged so that the twelve-vertex numbers are the same numbers.
The miss probability is over the observed shares — the chance that every description is reached at least once in starts, subtracted from one. It uses the shares the search itself found, so it is a self-consistent statement rather than an external estimate, and it can only be an underestimate: a description never reached contributes nothing to it.
The searches are cached between builds and verified on read by running one fresh start and requiring its description to be one the stored population contains. A localiser that had changed would produce a description the stored population does not have, and the cache would refuse.
The refusal is a system with one description. Two isolated double bonds have exactly one, and the classification must report a single basin rather than two groups — a shape rule that found structure there would be finding it in noise.
Where the model stops
This is a Hückel cage with one orbital a vertex, not a borane. The electron counts swept here run from two to twice the vertex count and most of them correspond to nothing; what they are for is to generate occupied sets of different sizes and degeneracies on a fixed graph, which is the only way to get more than one case out of a family this small.
Fractional occupations are allowed. Where the filling stops inside a degenerate group the occupied set is partly filled, and localising a partly filled set is a different object from localising a closed shell — the localisation transformation is defined over occupied orbitals and does not ask whether they are equally occupied. Two of the three cases here are in that situation and one is not, and the comparison treats them alike.
The cage graphs are also unweighted, so every edge is the same. A real cage’s bonding is not a graph and the descriptions counted here are descriptions of a graph’s occupied space, which is what makes them countable and also what makes them not chemistry.
And the descriptions are keyed on centre counts rather than on the orbitals themselves. Two genuinely different descriptions that happen to share a sorted list of centre counts would be counted once, so every count here is a lower bound for a second reason as well.
The generalisation
A feature of one measured distribution is not a rule until it has been looked for in a second one, and the cost of looking is usually far less than the cost of being wrong.
The twelve-vertex gap is real and the reading of it was reasonable: a population with a clean break in it does look like a population whose small members are a separate kind. What the second and third cases show is that the break is not what makes the count right — the count was right because four thousand starts was enough, and the gap was a coincidence of that system that happened to point the same way.
The same shape turns up in three other places. A constant that cancelled a coupling belonged to one lattice; a local search’s strength belonged to one size; a diagnostic’s calibration belonged to one axis. In every case the demonstration was sound and the range it was demonstrated over was silent about a variable, and in every case the second example was cheap to run.
Who found it, and when
Multiple local maxima of a localisation functional have been known since Edmiston and Ruedenberg’s work of 1963, and the standard practice — start from several places and take the best — is an acknowledgement of it. That the number of maxima found depends on how many starts were used is obvious and rarely reported; that the distribution of basin sizes has a shape worth classifying does not appear to be a standard remark, and the contribution here is mostly that it does not have one shape.
The miss probability is elementary and is a coupon-collector’s calculation with unequal probabilities. Its application to basin populations is straightforward.
Still open: what decides the shape, and the rare descriptions
The obvious open question is what decides the shape. Three cases give three shapes and no variable that sorts them: the degeneracy of the frontier does not, since the two degenerate cases give opposite extremes. The survey already computed contains forty-odd (cage, filling) pairs at four hundred starts each, and running the ten with the most descriptions to four thousand would give ten populations to look for a pattern in — which is the smallest set that could find one.
The nearer question is the one the third case makes urgent. A search whose largest basin takes ninety-eight per cent of the starts will report one description however long it is run in practice, because nobody runs four thousand starts when the first fifty agree. Whether the fourteen rare descriptions there are worse descriptions by the functional, or merely rarer, is one number per description and is already computed — and if they are equally good, then a localisation on that system is reporting one of fifteen equally valid pictures and calling it the answer.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- One scale, from two centres to a cage — both name canonical orbitals, local minimum, localisation, model limit, three-centre bonding, underdetermination, unitary transformation
- A bond order between atoms that do not interact — both name closed form, degeneracy, model limit, three-centre bonding, underdetermination
- A contrast with a closed form — both name closed form, convergence, degeneracy, filling, model limit
- A count rather than an average — both name approximation, closed form, convergence, localisation, model limit
- An interior maximum a third orbital allows — both name canonical orbitals, closed form, localisation, model limit, unitary transformation
- How many descriptions a cage has — both name canonical orbitals, local minimum, localisation, underdetermination, unitary transformation
Named objects
A dashed tag is an object no other essay names yet.
ApproximationCanonical orbitalsClosed formConvergenceDegeneracyFillingLocal minimumLocalisationModel limitThree-centre bondingUnderdeterminationUnitary transformation