What is taught wrongly

The worst system in the square was the solver's

Every score in the transferability square rests on an unrestricted mean field, and every one was computed from a single conventional start. On a band of nine systems, each with a site energy one or two below the repulsion, that start lands above the lowest solution the method has — by 1.12 at a repulsion of six and a site energy of four, which was the largest error in the whole square. Recomputed on the lowest field, the factor of forty-six that first showed the polarisation to be no rule falls to 2.61 and passes; the grid still fails every candidate, by a factor of fifteen rather than forty-eight; and the typical comparison, which no worst case could show, is far better than chance for every candidate but one.

Worth reading first: A blank is not a pass · The arms were the kindest part of the square.

The transferability square is a single question asked many times. A composite method computes a cheap mean field on the system it wants and adds a correlation correction computed somewhere else — here on a reference ring of four at a repulsion of eight and no site energy — and the question is whether anything cheap warns in advance when the carried correction will be wrong. The mean field’s own symmetry breaking was the first candidate, and since then seven quantities a mean field produces for nothing have been scored on a cross of two axes, censused for where they fail, scored again on a full grid, and scored where a band could not reach.

Every one of those numbers has a mean field in it, twice: once in the composite’s energy and once in the diagnostic. And every one of them took the mean field from a single start — densities twenty per cent off uniform, alternating in spin — iterated until it stopped changing. The answer that iteration reaches was taken to be the mean field’s answer.

It is not always. The unrestricted mean field is variational: its answer is the lowest energy any single determinant reaches, and a self-consistent solution above that is a fixed point the iteration happened to find. A better energy is not always a better answer, but here the method defines its answer as the lower one, and on a band of the square the single start did not find it.

A fixed point is not a minimum

The self-consistent field is an iteration. Each spin’s electrons fill the lowest orbitals of a one-electron problem whose site energies depend on where the other spin’s electrons are, and the densities that come out are fed back in until they reproduce themselves. Any self-reproducing density is a solution — a stationary point of the mean-field energy — and a problem with competing orders has several.

This one has competing orders by construction. The repulsion favours spin alternating round the ring, one up electron on each of two opposite sites and one down on each of the other two. The staggered site energy favours charge sitting on the two low sites, both spins together. Near the site energy at which the second wins, both kinds of solution exist, and which one an iteration reaches depends on where it started.

So the lowest mean field was found by brute enumeration rather than by one start: the conventional start, the symmetric one, every placement of each spin’s two electrons on two of the four sites (thirty-six of them), and four random densities with a fixed seed. Every start that converges is kept and the lowest is the answer.

The conventional start misses the lowest mean field on one band, and nowhere else. A four-site ring at every combination of repulsion and staggered site energy on a 12 by 12 grid. A small dot is a system where the conventional broken-symmetry start reaches the lowest mean-field solution; a circle is one where it lands above it, with area proportional to how far. There are 9, every one with a site energy one or two below the repulsion, and the largest miss is 1.124 at a repulsion of six and a site energy of four.
Fig. 1 Where the conventional start lands above the lowest mean-field solution, over a grid of repulsion and site energy.

On a grid of 144 systems — twelve repulsions from one to twenty by twelve site energies from zero to twelve — the conventional start misses the lowest solution on nine. Every one of them has a site energy one or two below its repulsion, from a repulsion of three to eight. The largest miss is 1.124 at a repulsion of six and a site energy of four; the next are 0.524 and 0.505. Everywhere else, including every system at zero site energy and every system far from the band, the two starts agree to the iteration’s tolerance.

The band sits between two collapse fields

The band has a location that explains it. The warning was that the broken-symmetry solution collapses at a definite site energy — above it the mean field finds no polarised solution — and that the transfer fails where it collapses. That collapse field was itself measured with the single start.

The broken solution survives further than the conventional start reportedThe site energy at which the mean field stops finding a spin-polarised solution, against the repulsion, from the conventional start and from the lowest mean field, with the line where the site energy equals the repulsion. At a repulsion of 6 the conventional start stops at 3.95 and the lowest field at 5.25; the start missed the lowest solution at a site energy of 4 and a site energy of 5.04812160481216repulsion Usite energy at the collapseU 6: lowest 5.25conventional 3.95ringed: 2 missed9 repulsionsa four-site ring at half filling · unrestricted mean field · conventional start against the lowest of many
Fig. 2 The site energy at which the mean field stops finding a polarised solution, against the repulsion, from the conventional start and from the lowest field; the dial picks one repulsion and rings the systems the start missed there.

The lowest field stays polarised further, at every repulsion from two to twelve: 1.05 becomes 1.35 at a repulsion of two, 2.35 becomes 3.15 at four, 3.95 becomes 5.25 at six, 6.25 becomes 7.35 at eight. At twelve the two nearly meet, 11.25 against 11.45, and at sixteen the lowest field’s collapse is slightly the earlier, 15.75 against 16.0.

Seven of the nine missed systems lie between the two collapse fields. In that strip the conventional start reports no polarised solution, a polarised one exists, and it is lower. The missed systems are not scattered failures of an iteration; they are mostly the region where the single start’s collapse field was wrong. The other two sit just below both collapses — a site energy of five at a repulsion of seven, where the two fields collapse at 5.10 and 6.25, and six at eight — and there the start reached a polarised solution that is not the lowest polarised one: at a repulsion of eight and a site energy of six, a polarisation of 0.802 where the lowest solution’s is 0.838.

The warning’s own claim survives this. Across its thirty-two systems its rank agreement with the error was 0.902 and is 0.913 on the lowest field, and the error still rises with the diagnostic at every repulsion before the collapse. The broken solution still collapses, and the transfer still fails where it goes. What moved is where it goes.

The square’s largest error was the start’s

The composite’s error is its energy less the exact one, and its energy is the mean field’s plus the carried correction. A mean field 1.124 too high puts 1.124 into the error, and nothing downstream can tell that from a failure of the transfer.

The largest error in the square belonged to the start, not to the transfer. The composite's error at the grid systems where the two starts disagree, from the conventional start (pale) and from the lowest mean field (solid), and the largest error among the systems where they agree. The conventional grid's largest error, 1.146 at U 6, ε 4, is 0.022 on the lowest field: almost all of it was the mean field landing above its own minimum. The largest error left in the square is 0.351.
Fig. 3 The composite’s error at the grid systems where the two starts disagree, and the largest error where they agree.

Four of the grid’s eighty systems are on the missed band. At a repulsion of six and a site energy of four the error was 1.146 — the largest anywhere in the square, and the point the grid’s figure drew as its biggest mark — and on the lowest field it is 0.022. At a repulsion of eight and a site energy of six it was 0.545 and is 0.040. So nearly all of the square’s two largest errors were the iteration, not the transfer.

Not every change is a reduction. At a repulsion of three and a site energy of two the error grows, from 0.043 to 0.108: the lowest mean field is lower, and the correction carried from the reference was fitted to a different kind of solution, so the two no longer happen to cancel. A mean field that is more right is not a composite that is more right. The largest error left in the square is 0.351, at a repulsion of one and a site energy of ten, and nothing moved it.

The factor of forty-six was one of the missed systems

The polarisation’s failure has been the centre of the whole argument. The first cross found two systems at nearly the same polarisation difference whose errors differed by a factor of forty-six; the census counted that pair as the polarisation’s one failure; the next score quoted 45.7 as its worst ratio, and the gap rule set it beside the candidate’s median gap. One of the pair’s two systems was the site energy of six at a repulsion of eight — the ninth missed system.

On the cross, the polarisation passes once the start is fixed. The worst ratio between the errors of two systems at nearly the same diagnostic value, one on each arm of the cross, for each cheap candidate that has a comparable pair, from the conventional start (pale) and the lowest mean field (solid), on a logarithmic scale, with the census threshold of three. The polarisation falls from 45.7 to 2.61, under the threshold; the others barely move, and the control does not move at all.
Fig. 4 Each cheap candidate’s worst error ratio on the cross, from the conventional start and from the lowest field, against the census threshold of three.

On the lowest field the polarisation’s worst pair on the cross is 2.61, between a repulsion of twelve and a site energy of four, under the threshold of three that the census used to call a failure. The other candidates barely move: double occupancy stays at 4.66, the mean-field energy goes from 3.67 to 4.06, and the two products with the level shift do not change at all. The control — the change in the correlation energy, which is the error rearranged and must score near one — stays at 1.16.

So on the cross, the polarisation passes. The single most-quoted number in the argument was a mean field that had not found its own minimum.

The census loses a failure and keeps its conclusion

The census asked whether the candidates fail together, on one bad corner of the square, or each somewhere else. Its answer was five failing pairs on ten systems with none shared.

The census redrawn: one candidate fewer fails, and still no pair is shared. Each cheap candidate's failing pairs on the cross — two systems, one on each arm, at nearly the same diagnostic value and with errors more than three times apart — from the conventional start and from the lowest mean field. The polarisation's one failing pair disappears. 5 failures on 10 systems become 6 on 11, and no failing pair is shared by two candidates under either field.
Fig. 5 Each candidate’s failing pairs on the cross, from the conventional start and from the lowest field.

On the lowest field the polarisation’s failure disappears, and two others appear: double occupancy now also fails between a repulsion of four and a site energy of six, at 3.38, and the mean-field energy between a repulsion of twenty and a site energy of six, at 4.06. That is six failing pairs on eleven systems, four candidates failing instead of five, and still no pair shared between two of them. The census’s conclusion — the failures are not one corner — is untouched, and is if anything stronger, since the corner it would most have suspected was the one that turned out to be the solver’s.

On the grid every candidate still fails

The grid was where the diagonal made everything worse, and it is the reading that matters most, because it compares any two systems rather than one from each arm.

On the grid, every candidate still fails, by less. The worst ratio between the errors of two grid systems at nearly the same diagnostic value, for each cheap candidate, from the conventional start (pale) and the lowest mean field (solid), on a logarithmic scale, with the census threshold of three. The polarisation falls from 48.0 to 14.7 and still fails by a factor of five; the highest occupied level, which never involved the missed systems, does not move.
Fig. 6 Each cheap candidate’s worst error ratio on the grid, from the conventional start and from the lowest field.

The polarisation’s worst pair on the grid falls from 48.0 to 14.7; double occupancy and the product of the two from 48.0 to 14.6; the mean-field energy from 43.4 to 12.4. All of those worst pairs had the missed system at a repulsion of six as their large-error member. They are now between a system at a repulsion of one and the system at a repulsion of six and a site energy of six — two systems neither of which the correction touched. The highest occupied level stays at 627: its worst pair never involved the band.

So the grid’s verdict stands and its size does not. Every cheap candidate still fails, by a factor of twelve to fifteen rather than forty-three to forty-eight. And the interior is still worse than the arms for every candidate, though not by the same factors the grid first reported: the mean-field energy’s interior is now three times its arms rather than twelve, and the polarisation’s five and a half times rather than barely at all — 48.0 against 45.7 before, because the missed system had inflated both readings at once. The rank agreements barely move: the polarisation’s with the error goes from 0.783 to 0.793, the mean-field energy’s from 0.751 to 0.830, and the highest occupied level’s from 0.063 to −0.018, which is to say still none.

The typical comparison, which the worst case hid

The essay before this one left an open question: every score here is a worst case over pairs, which is the right shape for a warning and makes each score depend on one pair — and this essay has just shown what one pair can be. A median over the grid’s several hundred comparable pairs does not depend on any one of them, and it can be set against chance by shuffling a candidate’s values among the systems and recomputing.

A typical comparison is far better than chance, except for the highest occupied level. For each quantity, the median ratio between the errors of two grid systems at nearly the same value of it, on the lowest mean field (solid mark; the conventional start's is the hollow ring), against the range of the same median when the quantity's values are shuffled among the systems (the bar, from the 5th to the 95th of a hundred shuffles). Every candidate's typical pair agrees to about a fifth, where chance gives a factor of three — except the highest occupied level, whose median of 3.05 sits inside its shuffled range.
Fig. 7 The median error ratio over comparable pairs for each quantity, against the range the same median takes when the quantity’s values are shuffled.

For every cheap candidate but one, two systems at nearly the same diagnostic value have errors that agree to about a fifth: medians of 1.21 for the polarisation over 538 pairs, 1.21 for double occupancy, 1.20 for the mean-field energy, 1.14 for their product. Shuffled, the same medians sit between 2.3 and 4.2. The control reads 1.07. The exception is the highest occupied level, whose median of 3.05 is inside its shuffled range of 2.55 to 3.45: it is no better than a column of random numbers, typical case and worst case alike.

On the conventional start the medians were nearly the same — 1.30 for the polarisation, 3.27 for the level — which is the point: a median does not care about nine systems of a hundred and forty-four. The candidates are usually right and occasionally badly wrong, with about four per cent of the polarisation’s pairs more than ten times apart. That is a different thing to tell a user of a composite method from “no better than chance”, and the worst-case scores, which were the only scores reported, could not tell the two apart.

What stands, and what does not

What each published number was, and is on the lowest mean field. The published statistics of the transferability square that depend on the mean field, computed from the conventional start and from the lowest mean field, with whether the published reading stands. Three were the start's — the first cross's worst pair, the polarisation's failure on the cross, and the square's largest error — and the rest stand or move by a fraction.
Fig. 8 The published statistics that depend on the mean field, on the conventional start and on the lowest field.

Three published numbers were the solver’s: the factor of forty-six on the first cross, the polarisation’s failure in the census, and the square’s largest error. Two moved and stand in substance: the collapse fields, which moved outward by up to a third, and the grid’s worst ratios, which fell by a factor of three and still fail. The rest stand as published — the warning’s rank agreement, the census’s disjointness, the control, the highest occupied level, and every rank correlation to within a tenth.

The check that would have caught all three is cheap and has a standard name. Quantum chemistry codes offer a stability analysis: after convergence, the second derivative of the mean-field energy with respect to rotating occupied orbitals into empty ones is computed, and a negative eigenvalue says the solution is a saddle point with a lower solution downhill. That test would catch a converged solution sitting on a ridge between two orders. It would not catch a solution that is a genuine local minimum above a lower one — a stable fixed point with a deeper basin elsewhere — and for that nothing short of more starts will do. On four sites, forty-two starts cost milliseconds. On a molecule the configurations cannot be enumerated, and the defensible practice is the same idea at a coarser grain: start from each chemically distinct order the system could plausibly adopt, and report which one won.

The published figures in the essays before this one are left as they were computed, each with a note pointing here, because they are a correct account of the conventional start and the conventional start is what most mean-field codes use by default. What changes is which reading to believe.

How the lowest field was found, and what it does not guarantee

The lowest field is the lowest over forty-two starts, not a proof of a global minimum. The thirty-six configuration starts are what make it credible on a ring of four: a weak random density rarely lands in the basin of a strongly ordered solution — two starts in thirty at a repulsion of two and a site energy of 1.2 — while a start that is already one of the orders lands in it. With every such order enumerated, a missed lower solution would have to be one no configuration of two electrons per spin approaches, and on four sites there is not room for many.

The checks are stated where they can fail: that every missed system has a site energy between zero and its repulsion and none has zero; that the collapse moves outward by more than eight per cent at every repulsion from two to eight; that the polarisation’s worst pair on the cross falls from above forty to under three while the control is unchanged; that on the grid the polarisation falls from above forty-five to between two and fifteen while the highest occupied level does not move; that every cheap candidate but the level has a median below its lowest shuffled median, and the level does not; and that the grid’s largest error was a missed system whose corrected error is under a tenth of the miss. The refusal is the reference: at no site energy the two starts must agree exactly, or every correlation energy the square carries would have moved with them.

Still open: a median beside the worst case, and a reference on the band

The general finding is not about this ring. A self-consistent calculation reports the fixed point its iteration reached, and in a problem with competing orders that is a property of the start. Two wrong numbers made a right difference only when both were the same kind of wrong, and a composite method’s error cancellation is exactly that bet; a mean field that lands on a different solution in the target than in the reference breaks it silently.

The obvious open question is whether the median is the score this argument should have used from the start. It is robust to one bad system, it separates a diagnostic that is usually right from one that is never right, and it says the three polarisation-based candidates track the error well in the typical case. A warning, though, is judged by its worst case — a practitioner is hurt by the pair it gets wrong, not reassured by the pairs it gets right — so the honest report is both numbers, and which to act on depends on what failing costs.

The nearer question is the reference itself. Every correction here is carried from one system at a repulsion of eight and no site energy, where the two starts agree. A reference chosen on the band — where the lowest field and the conventional one differ — would carry a correction fitted to a different solution, and whether the square’s scores depend on that choice as strongly as they depended on the start is one run of the same calculation away.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Composite methodCorrelation energyLocal minimumMean-field approximationSelf-consistencySymmetry breakingTransferabilityVariational principle