A blank is not a pass
Worth reading first: The arms were the kindest part of the square · Five failures in five different places.
The census of the diagnostics’ failures scored seven cheap quantities and left two of them blank. The rule it scored by is a worst case over pairs: find two systems whose diagnostic values agree to within fifteen per cent, take the ratio of the composite’s errors at them, and report the largest such ratio. A quantity that is a function of the error gives a ratio near one.
Two of the seven never produce such a pair. Their values on the repulsion axis and their values on the site-energy axis do not overlap at all — every value from one family is more than fifteen per cent away from every value in the other — so there is nothing to compare and the rule returns nothing.
The rule is built that way for a reason and the reason is sound. The quantity being predicted is the change in the correlation energy between a reference system and a target, and a diagnostic is a claim that this change is a function of something a mean field already knows. Testing a functional claim means finding two inputs the function would treat alike and checking that the output agrees, and two inputs that differ by forty per cent are not alike. A band is the correct instrument; it simply cannot be applied to every candidate.
The census said so and treated it as an absence of evidence, which was correct. It was also unhelpful in a specific way: a blank cell in a column of numbers between three and forty-six reads as better than any of them. The candidate with the largest failure in this whole argument spent two essays displayed as an empty entry.
Ask anyway, and report what it cost to ask
There is a way to score a candidate that has no pair inside the band, and it is to drop the band.
For each system, take the nearest system on the other axis by diagnostic value, whatever the gap between them, and record two numbers: the error ratio at that pair, and the gap itself. A large ratio at a small gap is a failure of the diagnostic. A large ratio at a large gap is no information about the diagnostic, because the two systems were never claimed to be alike. The pair of numbers says which, and either number alone does not.
The gap is the part that makes this a measurement rather than a looser version of the old rule. A candidate whose nearest cross-family neighbour is always far away is separating the two families — it assigns them disjoint ranges of values. That is a real property of the quantity, and it has two quite different explanations. Either the error separates the families in the same way, in which case the candidate is doing something right and the band was the wrong instrument; or the candidate separates them for reasons of its own and the error does not, in which case the separation is the failure.
Reporting the gap distinguishes those, and the banded rule reported both as the same absence.
Before anything is scored, the rule is checked on the quantity that cannot fail. The composite’s error is exactly the change in the correlation energy, so that quantity is the error wearing a different name, and any scoring rule has to return near one for it. Under the band it scores 1.16. Without the band, over nearest comparisons, it scores 1.84 at a median gap of 13.5 per cent. Both are near one, and the second is larger for the obvious reason: dropping the band admits a few pairs whose values differ by more, and two systems whose errors differ by thirty per cent have errors differing by thirty per cent.
A rule under which the exact quantity scored badly would be measuring the rule. This one does not, so what it says about the others is about them.
Both blanks fail
The highest occupied level’s worst nearest-comparison ratio is 388. Its closest comparison anywhere is 34 per cent apart, and even there the ratio is 167.
Polarisation times double occupancy gives 46. Its closest comparison is 28 per cent apart, and there the ratio is 50 the other way — the system with the smaller diagnostic has the larger error.
Neither is a success the band was hiding, which is the first and least interesting thing the scores say. It is worth having anyway: the alternative reading of two blanks, that the band happened to be badly matched to two otherwise fine quantities, is now excluded rather than merely implausible.
It also settles the arithmetic of the census. That census counted how many candidates fail and on which pairs, and it counted five because two had no score; its finding that no two candidates fail on the same pair was therefore a statement about five of seven. Scored, the two blanks fail on pairs of their own — the highest occupied level’s worst pair involves a system no other candidate’s worst pair touches — so the finding holds across all seven rather than across the five it could reach.
And they are blank for opposite reasons
The interesting half is the gaps, together with a quantity that only a grid supplies.
Both blanks separate the families — median gaps of 74 and 71 per cent, against 14 to 23 per cent for the five that the band could score. So on the gap alone they look alike, and the banded rule’s verdict of unscoreable was the same verdict for the same reason.
The rank correlation with the error separates them completely. Polarisation times double occupancy ranks the systems at 0.769. The highest occupied level ranks them at 0.063.
The two quadrants that matter here are the right-hand ones, and they were invisible under the band because the band collapsed both to a blank.
So one of them is a monotone function of a quantity that had already been scored, and the other is unrelated to anything. Polarisation alone ranks at 0.783 and fails at 45.7; multiplying it by the double occupancy, which ranks at 0.771, produces a quantity that ranks at 0.769 and fails at 46 — the same candidate with a different label, and its separation of the two families is inherited from the product of two quantities that each rise on one axis. The highest occupied level’s separation is its own, and it buys nothing: a quantity with a rank correlation of six hundredths does not order the systems at all.
That is the distinction the closing paragraph of the census asked for and could not make — whether they are untestable because they are good at separating the axes or merely because they are erratic. The answer is one of each, and neither is good.
The band was kind to the five it did score
The same rule, applied to the candidates that did have pairs inside the band, says something about the band itself.
Mean-field double occupancy scores 4.66 inside the band and 45.7 outside it. The mean-field energy scores 3.67 and 36.8. Polarisation divided by the level shift scores 5.56 and 16.6.
The mechanism is not subtle and it is worth naming, because it is the same shape of defect as the one the grid exposed. A band keeps only the pairs where two diagnostic values happen to agree closely, and on a cross of twenty systems there are at most a dozen such coincidences. Twelve pairs is a small sample and it is not a random one: it over-represents the regions where the two families’ values overlap, which are the regions where the diagnostic is behaving most like a single function of one thing. The band selects for the comparisons a diagnostic is best at.
So the two corrections this essay and the one below it make are the same correction arriving by two routes. The grid says the sample’s shape was favourable; the band says the sample’s filter was favourable; and the two together take the best cheap candidate from a factor of four to a factor of forty-six.
They are not independent, and it is worth being clear that they are not being added. Both are consequences of having twenty systems on two lines, and both would be removed by one change — sampling the plane. What the two essays establish separately is that either repair alone is enough to destroy the flattering numbers, so the census’s scores were not marginally optimistic but structurally so. The same lesson arrived one essay earlier from the other direction, when a diagnostic that worked on one family of systems was carried to another and stopped working: a quantity validated on a family is validated on a family.
A ratio alone ranks them wrongly
The last consequence is the one that justifies carrying two numbers rather than one.
Ranked by worst ratio, the best cheap candidate is polarisation times the level shift, at 11.4. Ranked by how close its comparisons are — that is, by how much of a comparison it is making at all — the best is the polarisation difference, at a 14.0 per cent median gap, and its worst ratio is 45.7.
The two orderings disagree at the top, so a table of ratios alone would recommend a candidate whose comparisons are among the weakest available, and a table of gaps alone would recommend one that fails by a factor of forty-six. Neither number is a score. The pair of them is, and what the pair says here is that none of the seven is a diagnostic.
What the rule does and does not establish
It is still a worst case, which is the shape this argument’s diagnostics have used throughout. One bad pair out of seventeen sets the score, which is the right statistic for a warning — a diagnostic that is reliable except in one place is not a warning — and is the wrong one for describing typical behaviour.
The medians are worth quoting beside it, because they are the reading a user who cares about typical behaviour would want and they do not rescue anything. The mean-field energy’s median ratio is 1.91, double occupancy’s 2.20, polarisation’s 2.61, polarisation times the level shift’s 4.90 — against 1.16 for the quantity that is the error. A typical comparison is wrong by a factor of two, which is a different complaint from the worst case’s factor of forty-six and is not a better one for a method whose whole claim is that a correction transfers. The highest occupied level’s median is 37.8, which is not even in the same discussion.
And the pairing is one-sided. Each system is paired with its own nearest neighbour on the other axis, so a system that is nobody else’s nearest neighbour still contributes a pair, and a system that is everybody’s contributes several. That is deliberate — every system gets exactly one comparison, so no region of either axis is weighted more heavily than another — and it means the seventeen pairs are not seventeen distinct pairs of systems. The polarisation’s seventeen pairs reach only ten distinct partners, five of which serve more than one system and one of which serves three — which is what a candidate that compresses one family into a short range does, and is a second reading of the same compression the gap measures.
The nearest comparison is nearest in the diagnostic, not in the system. Two systems that a candidate assigns nearly equal values to may be at a repulsion of twenty and a site energy of half — very different systems that the diagnostic cannot tell apart. That is the point of the test rather than a defect in it: a diagnostic is a claim that systems it cannot tell apart have similar errors.
And the two axes are still two axes. Everything above is computed on the cross of twenty systems rather than on the grid of eighty, deliberately, because the question is what the banded rule was hiding and the banded rule was run on the cross. The grid’s answer to the same question is separate and larger.
The rank correlations are the grid’s, not the cross’s. They are the one quantity here that a cross cannot supply — twenty systems on two lines do not make a ranking worth computing — so the two columns of the quadrant figure come from two different samples. That is stated rather than hidden because it matters for one reading: the claim that polarisation times double occupancy is a monotone function of polarisation rests on their rank correlations agreeing to two thousandths on the grid, and on the cross that comparison would be over twenty points.
The reference is still one reference, at a repulsion of eight and no site energy, and every error and therefore every score is a statement about transferring from it. The quantity being transferred is a correlation energy, which is itself a difference against a mean field, so a diagnostic here is being asked to predict the error in a difference of two things neither of which it sees. Whether a different reference would be flattered differently is untested, and the reason to note it here rather than in the essay before it is that the gap — how far apart a candidate keeps the two families — is a property of the candidate’s values, which do not depend on the reference at all. The gaps in this essay are reference-independent and the ratios are not.
What a working diagnostic would have looked like
It is worth saying what the bottom-left corner of this essay’s first figure would contain, because nothing has ever been in it and the shape of the empty region is informative.
A quantity that worked would have a small median gap — its two families would span the same range of values, so that systems from each can be compared — and a worst ratio near one. The control has both: a median gap of 13.5 per cent and a worst ratio of 1.84. It has them because it is the error, so the two families overlap in exactly the way the errors overlap.
The overlap is the part a cheap quantity would have to reproduce, and it is not obviously the hard part. Three of the candidates do overlap the families: polarisation, double occupancy and the mean-field energy all sit at gaps between 14 and 17 per cent, which is the control’s own region. They fail on the ratio instead. So the failure is not that the cheap quantities cannot be compared across the two axes; it is that where they can be compared, they disagree with the error by factors of tens.
That is a cleaner negative result than the census could state, and it is what the second column buys. Under the banded rule a candidate either had a ratio or did not, and the two ways of failing — no overlap, or overlap and disagreement — were one column apart and looked like the same thing.
An absence is not a value and a table cannot say so
The habit: when a measurement returns nothing, find a weaker measurement that returns something, and report what the weakening cost.
The census was right to refuse a score it could not compute and right to say so. What it could not do is prevent the refusal from being read as a number, because a table of scores has one column and a blank in it is a position in the ordering whether or not anybody means it to be. The repair is not a footnote; it is a second column, and the second column here — the gap — turned out to carry the finding.
The corollary is about what a caveat is worth. Unscoreable is a caveat attached to an absence, and it conveys neither how badly the candidate does nor why it could not be compared. Fails by a factor of 388, at a nearest comparison 34 per cent away, ranking the systems at 0.063 is a caveat attached to a number, and it says both. A weaker measurement with its weakness quantified beats a strong measurement that could not be made.
Who scored what, and when
The worst-case-over-pairs statistic, the fifteen-per-cent band and the seven candidates are this argument’s own, from the essays before it. Spearman’s rank correlation is 1904. The four mean-field quantities, the Hubbard ring, the broken-symmetry mean field and the exact diagonalisation are standard.
What is computed here is a scoring rule that pairs each system with its nearest cross-family neighbour and reports the gap beside the ratio, scores for the two candidates the banded rule could not compare, and the separation of separating from erratic using the ranking the grid supplies.
The numbers worth carrying are 0.769 and 0.063 — two candidates that a banded rule reported identically, and that differ by everything.
Still open: a median rather than a worst case, and a third family
The obvious open question is what the same candidates look like under a statistic that is not a worst case. Every score in this argument is the largest ratio over a set of pairs, which is the right shape for a warning and makes every score depend on one pair. A median ratio, or the fraction of pairs above a stated factor, would say whether the candidates are usually right and occasionally catastrophic or uniformly poor — and those are different things to tell a user of a composite method. Both are available from the pairs already computed, and the reason not to report them alongside was that a median over twelve pairs is not a median; over the grid’s twelve hundred it is.
The nearer question is a third family of systems. The gap measures how far apart a candidate keeps two families, and there are exactly two families here because there are two sweep axes — so separating is currently a statement about a particular pair of one-parameter families rather than about the quantity. A third axis would give three families and three pairwise gaps, and a candidate that separates every pair is doing something quite different from one that separates only the two that compete for the same physics. The highest occupied level is the one to ask it of, since its separation is the least explained thing in this essay.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- A verdict inside its own error bar — both name model limit, reference state, underdetermination
- An end effect with two signs — both name model limit, reference state, underdetermination
- Fifty descriptions of one molecule — both name model limit, reference state, underdetermination
- Half of it is given back at one bond — both name correlation energy, model limit, reference state
- One integer, and everything it changes — both name model limit, reference state, underdetermination
- One number decides which way it breaks — both name model limit, reference state, underdetermination
Named objects
A dashed tag is an object no other essay names yet.
Composite methodCorrelation energyMean-field approximationModel limitReference stateTransferabilityUnderdetermination