The arms were the kindest part of the square
Worth reading first: Five failures in five different places · The second number is the error, rearranged.
A composite method computes an expensive correction on a small case and adds it to a cheap calculation on a large one, and whether that transfer holds is the whole of its reliability. Four essays of this argument have looked for a cheap quantity that says in advance when it will not.
Seven candidates have been tried, all of them available from mean fields alone, and the census of their failures established that they fail on different pairs of systems rather than on one pathological corner — so the negative result is general rather than an artefact. Its closing paragraph named the hole in it twice, because the hole had already been deferred twice: every system in it lies on one of two arms. The points were generated by sweeping the repulsion at zero site energy and by sweeping the site energy at one repulsion, so the twenty of them form a cross, and a cross is not a square.
The obvious worry was that a diagonal region — where the repulsion and the site energy move together — might behave like neither arm, and that a diagnostic could be a good function of the error along each arm and a bad one in between.
It is worse than that. The arms are where the diagnostics work.
Filling in the square costs eighty diagonalisations
The calculation is the same one, run over a grid instead of a cross. Nine repulsions from one to twenty, nine staggered site energies from zero to ten, a mean field and an exact ground state at each of the eighty-one combinations, and the reference — repulsion eight, site energy zero — is where the correction is computed and is therefore exact by construction.
The cost is eighty exact diagonalisations of a four-site ring, which is nothing. That is worth stating plainly because it is the whole reason the hole existed: the grid was never expensive, it was simply never run, and each essay that noticed the gap had something else in hand.
The two axes are not arbitrary and the choice of the second one is the reason the square is interesting. A repulsion drives the system towards a state in which the electrons avoid each other by sitting one to a site; a staggered site energy drives it towards one in which they avoid each other by piling onto the low sites. Both of them break the mean field’s spin symmetry and the two symmetry breakings collapse at different places, which is why a cross through them was the natural first scan and why the region where they compete was always going to be the interesting part.
The first thing to check on new ground is the thing that cannot fail. The composite’s error is exactly the change in the correlation energy between the reference and the target, by the algebra of how the composite is assembled, and that holds at every point of the grid to the last bit the arithmetic carries. A grid on which it did not would be computing something other than the cross was.
The interior is worse, and by an order of magnitude
The statistic the census used is a worst case over pairs: find two systems whose diagnostic values agree to within fifteen per cent, and take the ratio of their errors. A quantity that is a function of the error gives a ratio near one. The census took those pairs across the two arms, since that is where its points were.
Taken between two systems neither of which is on an arm, the same statistic gives: mean-field double occupancy 48.03 where the arms gave 4.66; the mean-field energy 43.36 where the arms gave 3.67; the polarisation difference 48.03 where the arms gave 45.74.
Two of the three are worse by a factor of about eleven. A candidate the census reported as wrong by a factor of four is wrong by a factor of forty-eight as soon as the systems it is asked about are not on a sweep axis.
The polarisation difference is the exception and it is instructive. It was already the worst candidate on the arms, at 45.74, and it is 48.03 on the interior — barely moved. A quantity that was already failing badly had nothing left for the interior to expose, and a quantity that looked serviceable had everything. So the effect of the sample’s shape is not a constant multiplier on the score; it is a ceiling that the arms hid and the square removes, and the three candidates converge on it from below.
That reverses the reading of the census’s numbers without touching its conclusion. Its conclusion — no cheap quantity tracks the correlation energy, and the failures are not one corner — stands and is strengthened. Its numbers were charitable, and the charity came from the shape of the sample rather than from anything about the statistic.
It is worth separating those two, because a negative result whose numbers move by a factor of ten is not obviously a negative result that stands. What the census claimed is that each candidate has at least one pair it fails on and that no two candidates fail on the same pair. Both are existence statements, and adding sixty systems can only add failures — so neither can be undone by a larger sample. What it did not claim, and what a reader would reasonably have taken from a table of ratios between three and five, is that the best candidates are wrong by a factor of a few. That is the part the square removes, and it is the part a user of a composite method would have acted on.
Why the arms are kind has a plausible cause and it is worth naming as a guess rather than a result. Along an arm one parameter moves and the other is fixed, so the systems form a one-parameter family and every quantity computed on them is a function of that parameter — including the error and including the diagnostic. Two functions of one variable are functions of each other. The relation breaks as soon as the family has two parameters, and a cross has two one-parameter families rather than one two-parameter one.
The comparison a cross could not make
There is a second thing a grid buys, and it is larger than the first.
The cross’s twenty systems supply at most twelve pairs whose diagnostic values agree to fifteen per cent, and for two of the seven candidates it supplies none at all — which is why the census had to leave two entries blank. The grid supplies 1,232, of which 982 are interior.
Two orders of magnitude in the sample size changes what kind of statement the score is. A worst case over twelve pairs is a worst case over twelve coincidences: the pairs that happen to exist are the ones where two diagnostic values happen to agree, and there is no reason for them to be representative of anything. A worst case over twelve hundred is a worst case.
And it scores the two blanks. The highest occupied level, which the cross could not compare at all, comes out worst of every candidate at a ratio of 627. Polarisation times double occupancy comes out at 48.03, which is the same as polarisation alone — it is a monotone function of it over this range, so it was never a separate candidate.
Neither of those is a rescue and neither was expected to be. What they are is the removal of two blanks from a table, and a blank in a table of numbers reads as the best entry in it rather than as an absence. The candidate with the largest failure in this whole argument spent two essays shown as an empty cell.
Whether it sorts them at all
The grid makes available a question that a dozen pairs cannot support, and it is the more informative one.
A diagnostic that is any increasing function of the error would order the systems the way the error’s magnitude does, whatever the shape of the function. Ordering is a weaker demand than prediction and it is the demand a diagnostic actually has to meet, because a composite method’s user wants to know which of two calculations to distrust rather than by how much.
The polarisation difference reaches 0.783, double occupancy 0.771, polarisation times double occupancy 0.769, the mean-field energy 0.751. The highest occupied level reaches 0.063.
Two readings of that, and the second matters more.
Three quarters is a real relation and it is not a diagnostic. A rank correlation of 0.78 over eighty systems is not noise; there is something in the polarisation that knows about the error. It is also compatible with two adjacently-ranked systems having errors a factor of forty-eight apart, which is what the worst-case ratio says happens. A monotone relation with that much scatter tells a user nothing they could act on.
The something is not mysterious. The mean field breaks its own spin symmetry above a definite repulsion and the break is where the transferred correction starts failing, so anything that measures how broken the mean field is measures a quantity that rises where the error rises. Polarisation, double occupancy and the mean-field energy are three measures of the same break, which is why their three rank correlations agree to a hundredth. They are one candidate wearing three labels, and the census’s seven were closer to three.
And 0.063 is a different kind of failure. The highest occupied level does not order the systems at all. Its blank on the cross was not the band hiding a success; the quantity is unrelated to the error, and it was unscoreable because it separates the two arms rather than because it is subtle.
Where the error is large is not where the diagnostic is large
The rank correlation is one number over the whole square and it hides where the disagreements are, which a grid can show and a cross cannot.
Banding the square into quarters by the error, and separately into quarters by the double occupancy, the two bands agree at forty-six of the eighty systems. If the two quantities were unrelated they would agree at about a quarter of them, so forty-six out of eighty is well above chance and well below a function.
The disagreements are not at the edges. They sit in a band running across the middle of the square, where a large site energy and a modest repulsion produce a large double occupancy difference and a small error, and where a large repulsion at modest site energy produces the reverse. The cross passes through the two ends of that band and samples neither middle.
The mechanism is the competition the two axes were chosen for. A staggered field and a repulsion both move the double occupancy and they move it in opposite directions, so there is a curve through the square along which the double occupancy returns to its reference value while the system has been changed a great deal in both parameters — a curve of zero diagnostic and large error. Nothing on either arm is near it. A diagnostic built from a quantity that two competing parameters both move will always have such a curve, and finding it requires a sample with both parameters moving at once.
That is the most concrete form the finding takes. The arms are two lines through a square, they were chosen because each sweeps one parameter, and the region where the two parameters trade against each other is exactly the region neither line visits and exactly the region where the diagnostics fail.
What the grid is, and what it is not
A four-site ring. Every number here is that system’s, as every number in the essays before it is. The Hilbert space is thirty-six states in the relevant sector and the exact answer is a diagonalisation, which is what makes eighty of them free. Whether the finding survives to a system with more than one kind of correlation in it is the open question the whole argument carries — and the correlation energy of a small ring is already two things rather than one, so a diagnostic that tracked it would have to track a sum whose terms behave differently.
One reference. The correction is computed at repulsion eight and zero site energy and carried everywhere. A different reference moves every error and therefore every score, and that the choice of reference is worth a factor of two hundred and fifty-nine in the correlation energy itself is the finding. What is being tested is the transfer from one stated reference, which is what a composite method does.
The site energy is staggered. The field alternates in sign around the ring, which is the arrangement that competes with the repulsion for the same physics and is why it was chosen as the second axis. A uniform field does nothing at fixed filling; a random one would be a third kind of axis and a different question.
Fifteen per cent is the band and it is inherited. The statistic keeps the census’s own tolerance so that the two numbers are comparable, and the band’s own effect is visible in the control: the quantity that is the error scores 1.18 rather than 1.00, because two systems whose diagnostic values differ by fifteen per cent have errors differing by about that much.
And the errors themselves span three decades across the square, from 4.7 × 10⁻⁴ at systems near the reference to 1.15 at the far corner. A worst-case ratio taken over that range is not dominated by the small errors, because the statistic only compares systems at the same diagnostic value and a diagnostic near zero is excluded by its own floor. What it is sensitive to is the pairing of a large error with a small one at equal diagnostic, which is exactly the failure being looked for — and the same shape of failure the composite shows when it is asked to carry a correction between two rings rather than two parameter values.
A sweep is a sample and its shape is a choice
The habit: a parameter scan is a sample of a parameter space, and sweeping one parameter at a time samples a measure-zero subset of it.
That sounds like a pedantic objection and it is worth twelve-fold here. The reason is specific: along a one-parameter family every quantity is a function of the parameter, so every pair of quantities is functionally related whether or not they have anything to do with each other. A scan built from one-parameter sweeps therefore manufactures relations between the quantities it measures, and a test for a functional relation run on such a scan is asking a question the sampling has already answered.
The corollary is about cost. The grid cost eighty diagonalisations against the cross’s twenty, and in exchange it gave a hundred times as many comparisons, a score for two candidates that had none, and a statistic — the ranking — that a dozen pairs cannot support. Four essays deferred it as a larger calculation. It was four times the calculation and a hundred times the evidence, and the reason it kept being deferred was that the cross kept producing answers.
Who tried what, and when
Composite methods that add a correlation correction computed on a small basis or a small system to a large calculation are the standard practice of quantum chemistry, and the assumption behind them is the additivity this argument has been testing. The four cheap quantities used here as candidates — the mean field’s spin polarisation, its double occupancy, its highest occupied level and its total energy — are the ones a mean-field calculation reports without being asked. The Hubbard ring, the mean field with broken symmetry and the exact diagonalisation are all standard.
What is computed here is the same transferability test over a two-dimensional grid of eighty systems rather than a cross of twenty, the interior’s worst-case ratios against the arms’, the rank correlation a grid makes available, and scores for the two candidates the cross left blank.
The numbers worth carrying are 4.66 and 48.03 — one candidate, one statistic, and the difference between a cross and a square.
Still open: a bigger ring, and a third axis
The obvious open question is the system. A four-site ring has one kind of correlation in it and the composite being tested carries a correction between two of its parameter values, so nothing here distinguishes a failure of transferability from a failure of the ring. A six-site ring costs four hundred diagonalisations rather than eighty on the same grid, which is still minutes, and it has something the four-site ring does not: a filling other than half that is not trivially related to half. Whether the interior stays an order of magnitude worse than the arms, or whether that gap is a property of the smallest ring’s very simple error surface, is the same calculation at a size the census never had a reason to reach.
The nearer question is a third axis. The two parameters here are the repulsion and a staggered site energy, and the grid shows that what matters is the region where they trade. A third — the hopping’s own alternation, which competes with both — would make the sample a cube and would say whether every pair of axes has a bad diagonal or whether this pair is special because both of them drive the same symmetry breaking. Nine values on a third axis is seven hundred diagonalisations, and the reason to do it is that the finding above is currently about one plane and is stated as though it were about sampling.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- A mean field cannot get out of the way — both name correlation energy, model limit, symmetry breaking
- Half of it is given back at one bond — both name correlation energy, model limit, reference state
- The composition that is hard is not the full one — both name model limit, reference state, symmetry breaking
- The distortion that opens the gap — both name model limit, reference state, symmetry breaking
- The give-back that turned into a saving — both name correlation energy, reference state, symmetry breaking
- Two wrong numbers and a right difference — both name correlation energy, model limit, reference state
Named objects
A dashed tag is an object no other essay names yet.
Composite methodCorrelation energyMean-field approximationModel limitReference stateSymmetry breakingTransferability