What is taught wrongly

Five failures in five different places

Seven quantities a mean field produces for nothing have been tried as diagnostics, and none of them is usable. The question left is whether that is one finding or seven — whether the same awkward corner of the square breaks every candidate, or each is broken somewhere else. Each is broken somewhere else. Five candidates, five failing pairs, ten systems, and not one of them appearing twice.

Worth reading first: The second number is the error, rearranged · The half of the square a ring of four cannot show.

The second number is the error rearranged established that the diagnostic a composite method’s validation would want — the change in the correlation energy — is the method’s error rearranged, and so costs the calculation it is meant to replace. Having removed it, it asked what a mean field can supply on its own, tried seven such quantities, and found the best of them leaving a factor of 3.67 between two systems at the same diagnostic value.

Seven numbers between 3.67 and 45.7 with no evident ordering is an unsatisfying way to end, and its closing paragraph said why. The scores are produced by pairs, and the pairs are recorded. If the same two or three systems break every candidate, then the diagnostics all fail on one awkward corner, and the honest statement is a narrower and more useful one: composites are fine except near a mean-field instability. If each candidate is broken somewhere else, then nothing cheap tracks the correlation energy and the negative result is general.

The pairs are all different.

Five candidates, five failures, five different places. Every point of the square, with each cheap diagnostic's failing pair joined by a line. The five tested candidates fail on five different pairs involving 10 different points — no line shares an end with another. Had they all failed on one corner the lines would have converged on it, and the honest conclusion would have been that composites are safe away from that corner.
Fig. 1 Every system on the square, with each cheap diagnostic’s failing pair joined. No line shares an end with another.

Why the disjunction was worth setting up

A negative result can be reported two ways and only one of them is useful.

None of these diagnostics works is a statement a reader can do nothing with. None of these diagnostics works, and here is the region where they all fail is a statement that hands back a usable rule — avoid that region, or check by hand inside it. The difference between the two is entirely a matter of whether the failures coincide, and that is a question about the cases rather than the scores.

The warning a cheap calculation gives is the account of what a cheap number is for: not to predict the answer but to flag the cases worth a second look. A diagnostic that fails in one identifiable region is still doing that job everywhere else. One that fails everywhere is not.

So the disjunction is not a tidying exercise. It decides whether the seven scores leave a usable instrument behind.

Five, and all of them elsewhere

Of the seven candidates, five can be tested — the other two put no system from one axis within fifteen per cent of any system from the other, so the statistic has nothing to compare and returns nothing. Each of the five has exactly one pair exceeding a factor of three, and the five pairs are:

  • the polarisation difference, on site energy 6 against repulsion 4;
  • the mean-field double occupancy, on repulsion 16 against site energy 4;
  • the mean-field energy, on repulsion 1 against site energy 8;
  • the polarisation divided by the level shift, on site energy 10 against repulsion 10;
  • the polarisation multiplied by it, on repulsion 3 against site energy 3.

Ten systems, each appearing once. The two axes contribute five each, so the failures are not even lopsided between them — which they would be if one kind of change were the difficulty.

No point is implicated twice. How many of the cheap diagnostics each point of the square is part of a failure for. If the diagnostics were all breaking on one pathological system, a few points would carry every failure. Instead each of the 10 points appears in exactly 1, so the failures are as spread out over the square as they could be.
Fig. 2 How many failures each system is implicated in. Every one appears in exactly one.

Had the diagnostics been failing on a shared pathology, the lines in the first figure would converge and this second figure would have a few tall bars and many empty ones. It is flat at one. So there is no small set of awkward systems whose removal rescues anything, and the negative result of the seven scores is a statement about cheap diagnostics in general rather than about a corner of one square.

The two that cannot be tested are worth naming rather than passing over: the highest occupied level, and the polarisation multiplied by the double occupancy. Both place the repulsion axis and the site-energy axis in disjoint ranges — the level shift runs 0.019 to 0.151 along one and 0.228 to 2.704 along the other — so no pair of systems from opposite axes is within fifteen per cent, and the statistic has nothing to say. That is not a pass, and the census cannot speak for them.

What that rules out

It is worth being precise about which comfortable conclusion this removes.

The comfortable conclusion would have been: the composite method works, and the cheap diagnostics work, except near the mean-field instability where everything is difficult anyway. That is a common shape for a negative result to take, it is often true, and it would have made the seven scores a measurement of one thing rather than seven.

It is not available. The failures are spread across the square — near the reference and far from it, on both axes, at strong coupling and weak. A practitioner cannot be handed a region to avoid, because the failures are not in a region.

Every failure, and the pair that causes it. Each cheap diagnostic with the pair of systems whose errors disagree most at the same diagnostic value. Reading down the middle two columns is the finding: no pair appears twice and no system appears twice, so there is no small set of awkward cases whose removal would leave the diagnostics working.
Fig. 3 Each failure with the two systems that produce it.

The half of the square a ring of four cannot show built this square precisely to have two axes rather than one, on the argument that a diagnostic tested along a single axis has not been tested. That argument survives here and is strengthened: the failures are distributed over both axes, so a validation run along either alone would have found a diagnostic that works.

How much each score is worth

The census also settles something the scores alone could not, and it makes the table of scores read differently.

How many comparisons each candidate gets, and how many it loses. The number of pairs the statistic can compare for each cheap candidate, with the number that exceed a factor of three drawn inside. The counts differ by sixfold — a diagnostic that spreads the two axes apart offers few comparisons and one that mixes them offers many — so the earlier scores are not all equally well founded, and the two with no comparable pair at all are not scored here either.
Fig. 4 How many pairs the statistic can compare for each candidate, with the failing part drawn inside.

The five tested candidates are judged on 12, 10, 8, 4 and 2 comparable pairs. That is a sixfold spread, and it means a score of 4.90 resting on two comparisons is a different kind of number from a score of 45.7 resting on twelve. Neither is wrong; they are not equally founded, and a chart of seven bars invites a ranking the counts do not support.

The reason the counts differ is structural rather than accidental. A diagnostic that separates the two axes — that gives systematically different values to a repulsion and a site energy — offers few pairs within fifteen per cent of each other and so is barely tested. A diagnostic that mixes them offers many. So the statistic tests hardest the candidates that place the two axes on top of one another, which are exactly the candidates most likely to fail.

That is not a flaw to be corrected but it should be said aloud: the statistic and the candidates are not independent, and a candidate can score well partly by being hard to test.

That interaction between statistic and candidate has a consequence worth drawing out. Two of the seven were untestable because they separate the axes cleanly — which is, on the face of it, a virtue. A quantity that assigns systematically different values to a change in repulsion and a change in site energy is telling the two apart, and telling them apart is what a diagnostic of which kind of change this is would want to do.

But the question here was never which kind of change it is. It was how large the composite’s error will be, and for that a quantity that separates the axes is a quantity that cannot be checked across them. The same property makes a candidate look good on one reading and untestable on another, which is the clearest sign available that the statistic is answering a narrower question than the words “is this a good diagnostic” suggest.

Two systems a model cannot tell apart is the same structure met in a different setting: a diagnostic that assigns two systems the same value has made a prediction about them, and the prediction is checkable. Every failure counted here is one of those predictions coming out wrong, and what is added here is that no two of the wrong predictions are about the same systems.

Whether the line matters

The census depends on where the line is drawn, and says so. The number of failing pairs and the number of distinct points they involve, against the factor at which a disagreement is called a failure. At a threshold of zero every comparable pair counts and the census says nothing; at a factor of three it picks out one pair per candidate. A count that did not move with the threshold would be measuring the threshold rather than the diagnostics.
Fig. 5 The number of failures and the systems involved, against the factor at which a disagreement counts.

A census of failures needs a definition of failure, which is a threshold, and a reader is entitled to ask how much of the answer is the threshold.

At a factor of three each candidate contributes exactly one failing pair — its worst — because the second-worst of every candidate falls below three. At a threshold of zero every comparable pair counts, the census implicates every system on the square, and it says nothing at all. The answer moves, which is what it must do: a count that did not move with the threshold would be reporting the threshold rather than the diagnostics.

What does not move is the finding. At every threshold that separates anything, the failing pairs remain distinct — there is no level at which two candidates start sharing one.

A last check on whether “five different places” is really different. The five pairs involve, between them, systems at repulsions of 1, 3, 4, 10 and 16 and site energies of 3, 4, 6, 8 and 10 — which is close to the full extent of both axes. The repulsion axis runs from 1 to 20 and the site-energy axis from 0.5 to 10, so the failures are drawn from across nearly the whole square rather than from its ends or its middle. That is the sense in which the negative result is general, and it is a stronger statement than the absence of repetition alone.

What was computed, and how

The census, candidate by candidate. For each quantity a mean field produces on its own: how many pairs the statistic could compare, how many exceeded a factor of three, and the worst ratio. The last column is what the original score and the middle two are new — a score of 45.7 and a score of 3.67 are each produced by a single pair, and the pairs are different.
Fig. 6 Every candidate, the comparisons it gets, the failures it has, and which pair produces its score.

Nothing was recomputed. The square is the one the scores were computed on — twenty systems on a ring of four, eleven along the repulsion axis and nine along an alternating site energy, each with one exact diagonalisation and one unrestricted mean field. Every pair and every ratio was already produced in the course of scoring the candidates; here they are read out and counted.

The correction that was computed somewhere else built the composite examined here, and every question since has been about the same gap between what a method needs to know and what it can cheaply find out.

That is worth stating because it is the answer to why was this not done before: it was not expensive and it was not hard, it simply was not asked. The original question was how well each candidate scores, which is a question about the maximum over pairs, and threw away the argument of the maximum.

The refusal is the threshold at zero. A census that reported the same thing however the line was drawn would be measuring its own definition, and the check requires the count to grow by more than fourfold when the line is dropped — which it does.

One more reading the census makes available, and it is the most practically useful thing here. Because each candidate’s score comes from exactly one pair, the score is that pair — the number 45.7 is not a summary of twelve comparisons but a report on one of them, with the other eleven all below three. A reader who wants to know whether the polarisation is usable should look at site energy 6 against repulsion 4 and decide whether those two systems are ones they care about, rather than reading 45.7 as a typical performance.

That is a different relationship to a worst-case statistic than most tables invite, and it is the right one whenever the worst case is a single case. Two wrong numbers and a right difference makes the neighbouring point about summary numbers hiding their constituents.

Where the model stops

Twenty systems on one ring is a small square, and ten distinct systems in five failures is close to the maximum possible spread — with only twenty to choose from and ten slots, some repetition would have been likely by chance if the failures were random. So “no system twice” is a weaker statement than it sounds; what carries the argument is that the five pairs sit in visibly different parts of the square rather than clustering.

The five candidates are also five, and the two that could not be tested might have failed anywhere. A candidate that offers no comparable pair is not a candidate that passed, and the census cannot speak for them.

And a threshold of three is a choice. It was picked because it separates each candidate’s worst pair from its second-worst on this data, which is a property of the data and not a principle.

It is also worth noting what the scoring got right and the census leaves standing. The identity — that the composite’s error is the change in the correlation energy — is exact and is untouched. The floor of 1.176 that the fifteen per cent band imposes is untouched. And the conclusion that no cheap quantity does the job is not weakened by the census; it is strengthened, since the failures turn out not to be attributable to any corner that could be excluded.

A better energy is not a better answer is the standing warning about optimising the wrong quantity, and it applies to the search this census closes: seven candidates were tried and the best scored 3.67, and the temptation at that point is to try an eighth. The census says why that would be misdirected effort — the failures do not share a cause, so there is no structure for a cleverer combination to exploit.

The generalisation

The habit is to keep the argument of the maximum.

A great many summary statistics are maxima or minima over a set of comparisons — the worst case, the largest deviation, the steepest slope, the tightest bound. The number is reported and the case that produced it is discarded, and with it the ability to ask whether several such numbers have a common cause.

Keeping it costs nothing: it is already computed, since the maximum cannot be found without it. And it converts a list of scores into a structure that can be interrogated. Here it turned seven unordered numbers into a single statement — five failures, five places — which is a stronger and more useful conclusion than any of the seven numbers alone.

The same move is available whenever a table of worst cases is produced across several methods, several systems or several settings. If the worst cases coincide, the methods share a weakness and the weakness can be described. If they do not, the failure is general and the reader should stop looking for the special case.

A note on what would have changed the verdict. If the five pairs had shared even one system — if, say, site energy 6 had appeared in three of them — the right conclusion would have been that the mean-field instability near that point poisons every diagnostic, and the seven scores would collapse into one finding with a boundary attached. That is a real possibility and it is what the square was built to be able to show; the reference decides the correlation demonstrates that a single badly-chosen point can dominate a whole comparison.

It did not happen here, and the check took a few hundred subtractions.

Who found it, and when

Nothing here is a discovery about electronic structure. Recording which case produced a worst-case statistic, and asking whether several such statistics share their cases, is ordinary data hygiene and belongs to nobody.

What is new here is the reading on this square: that the five testable cheap diagnostics fail in five different places, that the comparison counts they are judged on differ sixfold, and that both facts were sitting in numbers the scoring had already computed.

Still open: the diagonal of the square, and the two untested candidates

The obvious open question is the corner, still — the two-dimensional scan that has now been deferred twice, in which repulsion and site energy move together. Every failure counted here lies on one axis or the other because every system does, and a diagonal region might behave like neither. That is the same calculation run over a grid rather than a cross, and it is bounded work.

The nearer question is the two that could not be tested. A candidate whose two axes never come within fifteen per cent of each other cannot be scored, and the census has treated that as an absence of evidence — correctly, but not helpfully. There is a way to ask anyway: hold the two axes apart deliberately and compare each system against the nearest system on the other axis rather than only against those inside a band, reporting the diagnostic gap alongside the error ratio. That would give the untestable candidates a score with a caveat attached rather than no score, and it would say whether they are untestable because they are good at separating the axes or merely because they are erratic.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Named objects

A dashed tag is an object no other essay names yet.

Composite methodCorrelation energyError cancellationMean-field approximationModel selectionTransferability