The second number is the error, rearranged
Worth reading first: The half of the square a ring of four cannot show · The correction that was computed somewhere else.
A test of transferred corrections put a composite method on a square. One axis turns up the repulsion, the other turns up an alternating site energy, and at each point a correction computed on a symmetric reference is transferred to a target that is not symmetric. The diagnostic it tested was the change in the mean field’s spin polarisation — a number that costs nothing, since the mean field has to be run anyway.
It half worked. The diagnostic separates the catastrophic transfers from the good ones on either axis taken alone, and fails to be one rule across both: two points at nearly the same polarisation difference, one from each axis, differ in error by more than tenfold, in the direction that would make a practitioner confident about the worse of the two.
The closing paragraph named the repair and observed that it was free. If the size of a transfer failure needs the change in the correlation energy as well as the change in the polarisation, then the product or ratio of the two ought to collapse the scatter that the polarisation alone leaves — and both numbers are already computed for every point, so the test is arithmetic rather than a calculation.
It is arithmetic. Four lines of it, and they finish the question in a way the paragraph did not anticipate.
Four lines of algebra
The composite is assembled from three pieces. The target’s mean field is cheap and is computed. The reference’s correlation energy — the gap between its exact energy and its mean-field energy — is computed once, on the small symmetric system, and that is the whole point of the method. The composite energy is the first plus the second.
Write the error of that against the target’s exact energy:
The target’s own correlation energy is , and substituting it collapses the expression to
The error of a composite is the change in the correlation energy between reference and target, with the sign reversed. Not correlated with it, not well predicted by it, not approximately equal to it under some condition — equal to it, by the definition of how the pieces were put together.
This is the same identity a method that is not additive ran into from the other side, where an energy was decomposed into pieces that were then recombined and the recombination was found to be the original quantity. A decomposition that is exact is a rearrangement, and a rearrangement can be useful or it can be circular; the difference is whether the pieces are separately obtainable.
That is why the twenty points lie on the diagonal, and why the largest departure across all of them is : the arithmetic’s own last bit, and not a measurement of anything. A diagnostic built from that quantity collapses the scatter perfectly because it is the thing the scatter is scattered in.
What a perfect collapse costs
The consequence is not that the suggestion to use it was a bad one. It is that the suggestion, taken literally, asks for the answer.
cannot be had without , and is the expensive calculation the composite was built to avoid. A practitioner who could compute the change in the correlation energy for their target would not need a transferred correction; they would have the target’s own. So the collapse is real, it is exact, and it predicts nothing that is not already known when it can be evaluated.
This is a specific and recognisable failure mode, and it is not the same as a diagnostic being wrong. A wrong diagnostic gives bad guidance and can be improved. A diagnostic that is the error rearranged gives perfect guidance and cannot be used, and the two are told apart by asking what it costs rather than by asking how well it scores. The original proposal asked how well it would score.
Two wrong numbers and a right difference is the essay this most resembles, and the resemblance is instructive rather than exact. There, two quantities were individually useless and their difference was meaningful, because the errors cancelled. Here the quantity is individually perfect and useless anyway, because it is not obtainable. Those are different diseases with the same symptom — a number that looks excellent on a validation set — and the second is the harder to notice, since nothing about it looks wrong.
The scatter that has to be explained
Removing the exact quantity from consideration leaves the question that actually wanted answering: is there anything a mean field produces on its own that collapses the scatter?
The statistic used throughout is unchanged: across all pairs of points that lie on different axes and whose diagnostic values agree to within fifteen per cent, the largest ratio between their errors. A diagnostic that is one rule scores near one. The polarisation scores , which is worse than the tenfold a sparser grid gives, because the fuller set of points here includes pairs a sparser grid misses.
The choice of a fifteen per cent band is kept unchanged, since altering it would make the numbers incomparable with the polarisation’s original score. It is worth knowing what it does: too tight a band finds no pairs and every candidate returns nothing, too loose a band compares points that are not really at the same diagnostic value and every candidate looks bad. Fifteen per cent is the setting at which the polarisation had twelve pairs to be judged on, which is enough to be a measurement and few enough that one pathological point can dominate.
The worst pair is a site energy of six against a repulsion of four, and where those points sit is the whole of why the diagnostic fails. A site energy of six is the last point at which the broken solution still exists — by eight it has collapsed to the symmetric one — and it is already behaving unlike its neighbours. Its polarisation difference is , roughly double the before it, which reads as an ordinary step along the axis. Its error is , an order of magnitude above the before it and the after; and its mean-field energy has moved back towards the reference while the field grew, which no neighbour does.
So the polarisation sees a system taking an ordinary step and the error sees one about to lose the solution its correction was computed on. Against it stands a repulsion of four, whose polarisation difference is — within fifteen per cent — and whose error is , the smallest on its axis. Two systems, one diagnostic value, and a factor of forty-six between what the method actually does to them.
Everything a mean field knows
A mean-field calculation produces more than a polarisation. It produces the occupations, and from them a double occupancy; it produces the orbital levels, and from them a highest occupied level; and it produces an energy. Each of those is available at no additional cost, and each can be combined with the polarisation.
Seven candidates cost nothing, and the results are flat. The mean-field energy is the best of them at ; the mean-field double occupancy is ; the polarisation divided by the level shift is and multiplied by it is . None of them approaches one, none of them is dramatically better than the polarisation’s neighbours, and the ordering among them has no evident logic.
The floor is worth stating, because it is not one. The statistic compares points whose diagnostics agree to within fifteen per cent, so a diagnostic that is the error can still be caught with two points fifteen per cent apart and score up to . The exact candidate scores , which is the floor and not a coincidence.
The best of the cheap ones
The mean-field energy’s worst pair is a repulsion of one against a site energy of eight. Those are opposite corners of the square in every sense: one is a nearly free system where the mean field is nearly exact, and the other is a strongly polarised one where it is not, and the mean-field energy happens to have moved by a similar amount from the reference in both.
That is the shape of every failure in this table. The cheap quantities all measure how far the target is from the reference in one or another coordinate, and the error measures how far the correlation energy has moved, and the two are not the same journey. A system can be far away in every observable a mean field reports and still have almost the same correlation energy, which is exactly the case that makes a composite method work at all.
That gap between observables and correlation energy is the subject of two kinds of correlation, which separated the part of the correlation energy that grows smoothly with the interaction from the part that appears when a single determinant stops being a good starting point. A mean field can report the first faithfully and be blind to the second, and it is the second that decides whether a transferred correction survives — so a diagnostic assembled from mean-field observables is being asked to see the thing its source cannot see.
Two candidates with nothing to say
Two of the nine return no score, and it matters that the table says untested rather than leaving a blank. The highest occupied level takes values between and along the repulsion axis and between and along the site-energy axis. Those ranges do not overlap, so no pair passes the fifteen per cent test and the statistic has nothing to report.
A blank in that column would read as the best score on the page. It is the absence of evidence, and the two are distinguishable only if the count of comparisons is printed beside the ratio — which is why it is. The same is true of the product of the polarisation and the double occupancy, which also has no comparable pair.
This is a small thing and it is the kind of small thing that decides what a table says. A statistic that can decline to answer must be able to say so.
There is a second reading of the same census that is worth taking. A candidate whose two axes occupy disjoint ranges is not merely untestable by this statistic — it is a candidate that would never be used across the two kinds of change, because a practitioner reading it would see two obviously different regimes and would not attempt to compare them. The census therefore separates the seven cheap candidates into those that make a false comparison possible and those that do not, which is a different and slightly more useful ranking than the one the ratios give. The polarisation, with twelve comparable pairs and the worst ratio on the page, is the most dangerous candidate on both readings.
What was computed, and how
The system is a ring of four at half filling with the reference at an on-site repulsion of eight. Twenty targets: eleven along the repulsion axis and nine along an alternating site energy, each requiring one exact diagonalisation of the full thirty-six-configuration space and one unrestricted mean field allowed to break spin symmetry.
Allowing the mean field to break symmetry is not a detail. A restricted mean field on this system has no polarisation to report at all, so the polarisation diagnostic would not exist; and a mean field cannot get out of the way is the essay establishing that the broken solution is where the interesting failures live. Every candidate here inherits that choice, which is one reason the seven scores are so similar: they are seven views of the same broken solution.
The identity is checked point by point rather than trusted from the algebra, because an identity that holds on paper can fail in an implementation that assembles the pieces differently. At every point the magnitude of the error and the magnitude of the correlation-energy change agree to within , and the check that would catch a broken assembly is the one at the reference, where the composite is exact by construction and both quantities must be exactly zero.
Where the model stops
Four sites is small, and the mean field on four sites has a specific pathology: a ring of four has a degenerate half-filled shell, so it is spin-polarised at every repulsion down to zero. That is why the site-energy axis carries the collapse the worst pair sits on.
More seriously, seven candidates is not all of them. A diagnostic could be built from the mean-field density matrix, from the gap between the two spin channels, from the iteration count the self-consistency took — and a negative result over seven is a statement about seven. The candidates were also chosen by hand, which is the honest description of any such list. They are the quantities a mean-field routine already returns and the two combinations the original proposal explicitly named. A systematic search over combinations would be a different exercise and would need a guard against finding a combination that fits twenty points by accident — the standing hazard a better energy is not a better answer is about, in the setting where the fitted quantity is a variational energy rather than a diagnostic.
What the identity establishes is stronger and does not depend on the count: the quantity that works exactly is the error, so any candidate that reaches the floor without costing an exact calculation would have to be a free quantity that equals a correlation energy, which is a larger claim than a good diagnostic.
And the composite here transfers a correlation energy between two systems of the same size. Real composite methods transfer between basis sets, between levels of theory, and between fragments of different sizes, and each of those has its own error terms that this one does not.
The generalisation
Composite and correction-transfer methods are everywhere in practical electronic structure, and each carries an assumption that its correction transfers. What this says is that the validation of such a method has a trap in it.
The obvious way to build a diagnostic is to take the cases where the method is known to fail, look for a quantity that tracks the failure, and propose it. The trap is that the quantity most strongly tracking the failure will usually be the failure — because a validation set is exactly the set of cases where the expensive answer has been computed, so every quantity derived from that answer is available and none of them will be available in use.
The correction that was computed somewhere else sets out the method tested here, and the trap described here is not a flaw in it — the method is sound and the square shows where it holds. The trap is in the second-order activity of deciding, in advance and cheaply, whether a given case is inside that region.
The test that separates them is not statistical. It is: can this number be computed for a case where the expensive calculation has not been done? A diagnostic that fails that test can score arbitrarily well on any validation set and is worth nothing, and a diagnostic that passes it is worth having at a factor of three.
That last clause is not a concession. The warning a cheap calculation gives argued that a cheap number’s job is to flag the cases worth a second look rather than to predict the answer, and by that standard a factor of three is a usable instrument: the mean-field energy separates the transfers that lose a few per cent from the ones that lose everything, and only fails to rank two cases that are already both bad. What it cannot do is be one rule.
Who found it, and when
Composite methods with transferred correlation corrections go back to the Gaussian-n schemes of Pople and co-workers from 1989 onward, and the observation that their error is the difference in correlation energy between the levels being composed is implicit in every derivation of them; it is not a discovery.
What is new here is the specific consequence for the proposed diagnostic, and the negative result over the seven free candidates on this system. Those are small claims about a small model, and the reason they are worth stating is that the algebra which makes the first of them obvious sits one substitution away from a proposal that did not notice it.
Still open: the corner, and why the candidates fail
The obvious open question is the corner: move along both axes at once and ask whether the region where the composite transfers well is a region rather than the intersection of two intervals. That is a two-dimensional scan with the same method and it is bounded work.
The nearer question is whether the free candidates fail for one reason or several. The seven scores sit between and with no ordering that reads as meaningful, but the pairs that produce them can be looked at directly: if the same two or three points break every candidate, the diagnostics are all failing on one pathological corner of the square and the honest statement is that composites are fine except near a mean-field instability. If each candidate is broken by a different pair, then no cheap quantity tracks the correlation energy and the negative result is general. The pairs are already recorded for every candidate, so this too is arithmetic — the same reason the correlation-energy test was worth running, and a reason to run this one.
What links here
Computed from the collection rather than written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- The give-back that turned into a saving — both name correlation energy, electron correlation, exact diagonalisation, hubbard model, many-electron wavefunctions, reference state, symmetry breaking
- A sign change is not always a zero — both name correlation energy, electron correlation, exact diagonalisation, hubbard model, many-electron wavefunctions, reference state
- Where the electrons are, without subtracting anything — both name correlation energy, electron correlation, exact diagonalisation, hubbard model, many-electron wavefunctions, reference state
- Half of it is given back at one bond — both name correlation energy, electron correlation, exact diagonalisation, hubbard model, reference state
- The reference decides the correlation — both name correlation energy, exact diagonalisation, hubbard model, reference state, symmetry breaking
- A contrast with a closed form — both name electron correlation, exact diagonalisation, hubbard model, many-electron wavefunctions
Named objects
A dashed tag is an object no other essay names yet.
Composite methodCorrelation energyElectron correlationError cancellationExact diagonalisationHubbard modelMany-electron wavefunctionsMean-field approximationReference stateSymmetry breakingTransferability