Beyond the octet

Where the count stops being an effort

A cage's localised descriptions were counted by a search that stopped when it stopped finding new ones, which makes the count a property of the stopping rule. Run to four thousand starts the count is fourteen and the last new description appears at start 124 — after which three thousand eight hundred and seventy-six starts add nothing. The four rarest are found six times in a thousand, which is what says nothing rarer is hiding.

Worth reading first: The answer a search is most likely to give · The localisation transformation, demonstrated.

A localisation search on a twelve-vertex cage finds fourteen distinct localised descriptions, and a single run of the search returns the best one 7.7 per cent of the time. The count depends on one thing:

A search that stops after three batches add nothing finds ten; four hundred and eighty starts find fourteen. The number of descriptions is therefore not a number at all but a curve against effort, and the useful quantity is where it flattens — or whether it does.

It flattens, and it flattens early.

The count is a curve, and the curve flattens. How many distinct localised descriptions of a twelve-vertex cage had been found after each batch, out of 4,000 random starts. The last new one appears at start 124; the remaining 3,876 add nothing. The dashed curve is what the basin sizes measured here predict — the chance of having hit each description at least once — and it is a closed form rather than a fit to the points.
Fig. 1 The discovery curve out to four thousand starts, with the closed form the measured basin sizes predict. The last new description arrives at 124.

Where it flattens

Four thousand random starts, each a fresh rotation of the occupied set, each localised to convergence:

starts descriptions found
50 10
100 13
200 14
1,000 14
4,000 14

The last new description appears at start 124. The remaining 3,876 starts find nothing that was not already there.

Where each description first showed up. The start at which each of the fourteen descriptions was first found. Ten appear within the first sixteen starts, the remaining four between 65 and 124, and the rest of the run — three thousand eight hundred and seventy-six starts — finds nothing new. That gap is the evidence: a population with more members would show them arriving, and nothing arrives.
Fig. 2 The start at which each description first appeared. Ten within sixteen starts, four more by 124, and then nothing.

The agreement between the measured curve and the closed form is worth a sentence, because it is the check that the two halves of the argument are about the same thing. The expected number found after nn starts is the sum over descriptions of 1(1pi)n1 - (1-p_i)^n, with pip_i the fraction of starts landing on description ii — a formula in quantities the same run measures. The measured curve sits within one description of it throughout, which it need not have: a run whose starts were correlated, or whose basins were mis-identified, would show the measurement running above or below the prediction systematically.

So the count is a count. That is what 480 starts cannot say, and it is not a stronger version of what they do say — it is a different claim. Four hundred and eighty starts establish that there are at least fourteen; four thousand with a flat curve establish that there are probably exactly fourteen, and the probability can be computed.

Why nothing rarer is hiding

The evidence is not the flat curve, which a population with one very rare member would also produce. It is the basin sizes.

description starts landing there share
the ten largest 242 – 582 6.1% – 14.6%
the four smallest 24 – 47 0.60% – 1.18%
How often each description is found. Each of the 14 descriptions with the number of starts that landed on it, out of 4,000. There are two groups and the gap between them is a factor of five: ten descriptions taking between 242 and 582 starts each, and four taking between 24 and 47. The four small ones are why a shorter search finds ten — and their size is what says nothing smaller is hiding.
Fig. 3 How often each of the fourteen is found, out of four thousand starts. Two groups, and a factor of five between them.

The rarest holds 0.60 per cent of the starts and none was found only once. A fifteenth description with a basin half the size of the smallest would appear about twelve times in four thousand, and one a tenth the size about two and a half times. Neither happened.

And the shape of the distribution is the second piece of evidence. The fourteen fall into two groups with a factor of five between them and nothing in the gap — a population that continued downward in size would show a tail, and a tail is what a set with more members hiding at the bottom looks like.

Two groups, with a factor of five between them. The fourteen basins in order of size, logarithmically. Ten of them hold between 242 and 582 starts and four hold between 24 and 47, with nothing in between. A population that continued downward would show a tail rather than a gap, and there is a gap.
Fig. 4 The fourteen basins in order, logarithmically. A gap rather than a tail.

It is worth naming the assumption that turns the basin sizes into evidence, because it is doing real work and it is not free.

The starts are uniform. Each is a random rotation of the occupied set drawn from a fixed sequence, so the fraction of starts landing on a description is an estimate of the fraction of the space of starting points that flows to it. If the sequence favoured some regions — if it were, say, biased towards near-identity rotations — the basins would be mis-measured and a description reachable only from a region the sequence avoids would look rare or absent.

Nothing here tests that directly. What can be said is that the ten large basins take between 6 and 15 per cent each, summing to 87 per cent, and a badly biased sequence would be unlikely to produce a distribution that flat. It is a consistency check rather than a proof, and it is the weakest link in the argument.

What a shorter search was risking

The basin sizes turn the stopping rule from a judgement into an arithmetic. The chance that a search of nn starts misses at least one description is one minus the product over descriptions of the chance of hitting each:

starts chance of missing at least one
100 92%
200 63%
480 12%
1,000 0.4%
2,000 under 0.1%
How likely a search of a given size is to miss something. The chance that a search of a stated number of starts misses at least one of the fourteen descriptions, computed from the basin sizes. At two hundred starts it is 63 per cent; at the four hundred and eighty it is 12; and it is under one per cent only past a thousand. The count at 480 starts was right and the search was not safe — the two are different claims and only the first could be seen at that effort.
Fig. 5 The chance of an incomplete answer against how hard the search is run, with the two earlier efforts marked.

Four hundred and eighty starts had a one-in-eight chance of returning thirteen, and returned fourteen. The three-dry-batch rule stopped at about two hundred, had a nearly two-in-three chance of missing one, and returned ten — which is what the risk looks like when it lands.

There is a sharper way to put the second row. A search of two hundred starts on this cage is more likely than not to report a wrong count — and it reports it with no indication of trouble, because the thing it reports is the number it found and the number it found is all it has. The failure has no symptom.

Neither search was wrong and neither could say what it was risking, because the risk is computed from the basin sizes and the basin sizes are what a longer search measures. The stopping rule and the confidence in it are the same measurement, and a search that stops when it stops finding things has no way to price its own stopping. The arrangement a count cannot pick meets the same wall from the other direction — there the search is exhaustive on a small net and cannot be on a larger one, and the question is how much is lost by giving up exhaustiveness.

The control

None of this means anything unless a system with one answer gives one answer however hard it is looked for.

The control, run just as hard. Two isolated double bonds have exactly one localised description, and 1000 random starts find exactly one. That is what makes the cage's fourteen a count of the cage rather than a count of what a search does to a random starting point — a routine that manufactured maxima out of its own starts would manufacture them here too.
Fig. 6 Two isolated double bonds, a thousand starts, one description throughout.

Two isolated double bonds have exactly one localised description — the two bonds — and a thousand starts find exactly one. So the cage’s fourteen are a property of the cage rather than of what a numerical search does to a random starting point, which is the failure mode a count of maxima always has to rule out.

Two things a stopping rule can mean

The distinction this essay has been circling is worth stating plainly, because it applies to every search of this kind.

A search can stop because it has finished, which is a statement about the population: there is nothing left to find. Or it can stop because it has run out of luck, which is a statement about the effort: what is left is rarer than the sampling reaches. The two look identical from inside the search — several batches adding nothing — and they are distinguished only by a quantity the search does not compute, which is how big the smallest thing it did find was.

That is a cheap addition to any search of this kind: keep the hit counts, and the stopping rule acquires a probability. The projector is unique and the basis is not is a statement about what is being searched for; this is a statement about the searching.

What fourteen means

It is worth restating what is being counted, because fourteen descriptions invites a reading it does not support.

All fourteen give the same density, exactly and not approximately. That is the whole of what a localisation is: a unitary transformation of the occupied orbitals, which leaves the density matrix exactly where it was. Fourteen descriptions is fourteen ways of drawing one thing.

They are genuinely different drawings, and that is established rather than assumed. Each has a different multiset of participation numbers, which no symmetry of the cage can change, so they are not one another seen from another side. And their functionals are separated by far more than the tolerance the maximisation converges to.

And the search returns the best one 7.7 per cent of the time, which is what makes the count matter: a single run of a localisation program on this molecule reports one of fourteen pictures, and which one is decided by a random number.

The basin sizes measured here sharpen it into something 480 starts could not show. The best description is not the most likely one. Its basin holds 8.65 per cent of four thousand starts — close to the 7.7 measured on far fewer — and the largest basin, at 14.6 per cent, belongs to a different description that is not the best by the functional.

So a single run of a localisation program on this cage most often returns an answer that is neither the best nor a random one: it returns the one with the widest catchment, which is a property of the optimisation landscape and not of the chemistry. Nothing in the program says which of the two a reader is looking at.

Four centres and the pair that will not localise is the case where the difficulty is that no localised description exists; this is the opposite difficulty, where too many do, and the two are the same fact about a unitary freedom seen from either end.

What this costs, which is nothing

The practical point is worth making because it is the reason nothing like this was done before.

Four thousand starts on this cage take under two seconds. Four hundred and eighty take a fifth of that, and the three-dry-batch rule stopped at two hundred. The stopping rule was not saving anything.

What it was saving was a decision. A search that runs until it stops finding things needs no judgement about how long to run, and a search that runs a fixed effort needs one — so the dry-batch rule is the one that can be written down without knowing anything about the problem. That is a real virtue and it is why the rule exists.

The repair is not to run longer; it is to keep the hit counts. A search that records how many starts landed on each answer can price its own stopping rule for nothing, whatever rule it uses — and the price is exactly what a stopping rule needs and usually lacks.

That applies beyond localisation. The arrangements a local search can reach and the conformers a ring search finds are both searches with stopping rules and no basin counts, and both would acquire a confidence for one extra array.

What is quoted, and what is computed

The localisation transformation sets out the calculation run four thousand times here.

Nothing is quoted. A twelve-vertex deltahedron, a Hückel Hamiltonian on its edges, and a Boys-style localisation of the occupied set.

Each start is a fresh random rotation of the occupied orbitals, generated from a deterministic sequence so that the run is reproducible; each is maximised to convergence and its answer identified by the sorted list of participation numbers to three decimals. Two starts that converge to the same list are the same description.

The expected discovery curve and the miss probabilities are closed forms in the measured basin sizes rather than fits to the curve, which is what makes the agreement between them a check. The curve is measured and the prediction comes from a different property of the same run.

What this cannot say

Four thousand is not infinity. A description with a basin of one in ten thousand would be missed by this run with probability two thirds, and nothing here excludes one. What the run excludes is a description as common as the ones found, and what the gap in the distribution suggests — rather than proves — is that the population does not continue downward.

The identification is by participation numbers to three decimals. Two genuinely different descriptions with the same multiset would be counted once, and the other check for that is a different test: whether the functionals are separated. Both are necessary and neither is sufficient, which is stated here rather than assumed.

A basin size is a fraction of a space nobody has characterised. The number 0.60 per cent is a fraction of the starting rotations this sequence produces, not of anything with an intrinsic measure — there is no natural volume on the space of unitary transformations that the localisation dynamics respects. So rare here means rarely reached by this procedure, which is what a chemist running the same program cares about and is not a statement about the description itself.

And a cage is one molecule. Fourteen is this cage’s number. Whether the shape of the result — a flat curve, two groups of basins, a gap — is general is a question about a family of cages, and this essay has one.

What the search requires

The last new description appears in the first quarter of the run. Not the curve is flat at the end, which a run that had stopped finding things for its own reasons would also satisfy: the check is where the last arrival is, relative to the effort spent after it.

The rarest description is found often enough that a missing one would be surprising — a share above one in four hundred, which is what makes the miss probabilities small rather than merely uncomputed.

The curve has flattened by the end of the run, which is the arithmetic check that the last batch’s count is the total.

And the refusal is the control. Two isolated double bonds must give one description at every effort, and a search that manufactured maxima from its starting points would fail there before anywhere else.

The stopping rule has a standard form

Stop when new answers stop appearing is a rule with no confidence attached, and the problem it is solving is one another field has already solved — with an estimator that turns the observed frequencies into a prediction of how many answers remain unseen.

The setting is identical. A search returns one description per start, drawn from a fixed population with unknown sizes; the question is how many members of that population have never been drawn. That is the problem of estimating how many species are present from a sample of individuals, and the standard answer uses the rarest observations rather than the total.

The reasoning is that the number of species seen exactly once is what carries the information about those seen zero times. If several descriptions have been found only once in four thousand starts, there is probably at least one more that has been found none; if every description has been found hundreds of times, there is probably nothing left. The estimator makes that quantitative: the number missing is approximately the square of the singleton count divided by twice the doubleton count.

The numbers here are the ones the estimator wants. Fourteen descriptions, four of them found about six times in a thousand starts, none found only once at that effort — which is exactly the configuration that says the population has been exhausted, and says so with a number rather than with a run of empty batches.

That converts the stopping rule from a habit into a statement. Stop when the singleton count reaches zero and report the estimator’s bound, rather than stopping when three batches in a row add nothing — because three empty batches is a statement about the batch size and a singleton count is a statement about the population.

Still open: the average, and the family

The obvious open question is the average over descriptions. If a single run returns the best answer 7.7 per cent of the time, the honest object is not any one description but the average over them — and averaging localised sets is not defined, because they are related by unitary transformations rather than by addition. What is defined is the average of the densities each contributes to each atom, and that average is the canonical density, which is where the argument started. The basin sizes measured here are the weights that average needs, so the circle can now be closed with a figure rather than a paragraph.

The nearer question is the family. This cage’s basins fall into two groups with a gap; whether a thirteen-vertex or a ten-vertex cage does the same is a run of the same search at a different size, and it is cheap — four thousand starts on this cage is under two seconds. If the gap is general, then count the big basins and check for a gap is a stopping rule with a stated confidence, which is what most searches of this kind currently lack.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationCanonical orbitalsClosed formConvergenceDegeneracyLocal minimumLocalisationModel limitThree-centre bondingUnitary transformation