What is taught wrongly

A better energy is not a better answer

The variational principle makes the energy a one-way test: lower is closer. It also makes the energy the least sensitive thing a wavefunction gets wrong — second order in the error where every other property is first — so the two diverge without limit as a calculation improves.

Worth reading first: Orbitals are not where the electron is · The smallest many-electron calculation.

The variational principle is the most reliable tool in quantum chemistry and it is reliable in one direction only. For any trial function whatever, the expectation value of the Hamiltonian is at least the true ground-state energy. Lower is better, and no calculation can accidentally come out below the answer.

That one-sidedness is why energies are the currency of the subject. Two methods can be ranked without knowing the answer; a basis set can be enlarged and the energy watched to see whether it has converged; a parameter can be optimised by minimising. None of that works for any other property, and the reason is worth understanding rather than accepting.

The energy is stationary at the exact wavefunction and nothing else is. Make an error ε\varepsilon in the state and the energy is wrong by order ε2\varepsilon^2; every other expectation value is wrong by order ε\varepsilon. The property that makes the energy safe to minimise is the same property that makes it uninformative about everything else.

The energy is the last thing a wrong wavefunction gets wrong. Two errors against the error in the wavefunction, on log axes, for a chain of 2 at U = 4t. The energy's line has slope 2.00 and the double occupancy's has slope 1.01: the first is second order in the error and the second is first order. So the two lines diverge as the wavefunction improves, and the energy stops being evidence about anything else long before it stops improving.
Fig. 1 Two errors against the error in the wavefunction that produced them, on log axes. The energy’s line has slope 2.00 and the double occupancy’s has slope 1.01. The two diverge as the wavefunction improves, without limit.

Making the wavefunction error an input rather than an estimate

The usual difficulty with a claim of this kind is that the error in a wavefunction is not measurable — the exact one is not available, which is why an approximation was being made.

Here it is available. The Hubbard systems used here are diagonalised completely, so the exact ground state is a vector and so is every excited state. A trial function whose error is known can therefore be built by hand:

ψ(ε)0+εm|\psi(\varepsilon)\rangle \propto |0\rangle + \varepsilon|m\rangle

whose overlap with the exact ground state is 1/1+ε21/\sqrt{1+\varepsilon^2}, so the error in the state is ε/1+ε2\varepsilon/\sqrt{1+\varepsilon^2} exactly. Nothing is estimated, nothing is fitted, and there is no variational calculation anywhere in the construction.

The energy of such a state is also exact, because the two components are eigenvectors and the cross term vanishes:

E(ε)=E0+ε2Em1+ε2E(\varepsilon) = \frac{E_0 + \varepsilon^2 E_m}{1 + \varepsilon^2}

E(ε)E0=ε2(EmE0)1+ε2.E(\varepsilon) - E_0 = \frac{\varepsilon^2 (E_m - E_0)}{1 + \varepsilon^2}.

The ε2\varepsilon^2 is visible on the page, which is the whole of the second-order behaviour. An ordinary observable has no such cancellation: its expectation value picks up a cross term 2ε0O^m/(1+ε2)2\varepsilon\langle 0|\hat{O}|m\rangle/(1+\varepsilon^2), which is linear.

Fitting the two errors against ε\varepsilon over a decade returns exponents of 2.0002.000 and 1.0061.006.

The mistake the check made first, which sharpened the claim

The first version of the calculation mixed in state number one — the next state up — and measured the double occupancy’s error at exponent two, apparently refuting the whole argument.

The refutation was real and it was about the dimer’s symmetry rather than about the principle. The first excited state of the two-site system has the opposite parity from the ground state, so 0D^1\langle 0|\hat{D}|1\rangle vanishes identically, the cross term is exactly zero, and the linear contribution is not there to be measured.

The repair is to choose the partner by the overlap that matters rather than by energy: mix in the state the observable actually connects to. For the dimer that is state three, with 0D^3=0.354\langle 0|\hat{D}|3\rangle = 0.354.

That produces a sharper statement than the original. An observable’s error is first order in the part of the wavefunction error that the observable connects to, and second order in the rest. A calculation can therefore be badly wrong in a way that one property cannot see and another can, and which property notices is decided by a matrix element rather than by how big the error is.

Nearly all of the error cancels, and the answer gets worse. For each repulsion: the error a spin-paired mean field makes in the total energy of one four-site system and of two two-site ones with the same number of electrons, and the error left in the difference between them. The cancellation improves from 83 to 97 per cent along the axis. The residue as a share of the quantity being computed goes the other way, from 1 to 423 per cent, because the reaction energy shrinks faster than what survives.
Fig. 2 Nearly all of the error cancels, and the answer gets worse. Two trial functions whose energies differ by very little give observables that differ by a great deal, because the cancellation that protects the energy is exactly the cancellation that leaves the wavefunction free to be wrong. That is the essay’s claim as a measurement rather than as an argument from orders.

What the numbers look like at a realistic error

Exponents are abstract; one worked point is not.

Take ε=0.2\varepsilon = 0.2. The trial state overlaps the exact ground state by 0.98060.9806, so the wavefunction is wrong by about two per cent — the sort of error a decent approximate method makes, and small enough that most people would describe the state as essentially right.

Its double occupancy is out by 53 per cent.

Halve the error to ε=0.1\varepsilon = 0.1. The energy error falls by four and the double occupancy’s by two, so the second is now a quarter of a per cent of what it was relative to the first. Halve again and the ratio doubles again. There is no error small enough for the two to converge together, and the gap between them grows without bound precisely as the calculation gets better.

That is the counterintuitive part and it is worth stating baldly. Improving a calculation makes the energy relatively more trustworthy and every other property relatively less so, in the sense that the ratio of the two errors diverges. A method that gets energies to a millihartree is not a method that gets densities to a comparable accuracy; it is a method whose density error is the square root of its energy error, up to constants.

How often two electrons are in the same place. Double occupancy per site in the exact ground state of a chain of 2, against the repulsion. It starts at the independent-electron quarter and falls to nearly nothing. The second curve is what the state actually pays in repulsion, which peaks and then falls because the electrons have stopped meeting.
Fig. 3 The observable being got wrong. Double occupancy in the exact ground state of the two-site system, against the repulsion — a number the trial functions above miss by up to half while their energies are nearly right.

Where this shows up in real calculations

The behaviour is not a curiosity of a two-site model; it is the reason for several standing pieces of practical advice, and it explains them rather than merely restating them.

Basis sets converged on energies are not converged on properties. A basis is usually enlarged until the total energy stops moving, and the dipole moment computed in it is usually still moving. That is the exponent-one behaviour: the energy has reached its second-order plateau while the wavefunction is still first-order wrong.

Hartree–Fock dipole moments are systematically too large, typically by ten to fifteen per cent for polar molecules, even though Hartree–Fock energies for the same molecules are within a per cent of the exact value. The energy error and the moment error are not comparable numbers and never were.

Recovering a large fraction of the correlation energy is a poor guarantee. A method described as capturing ninety per cent of the correlation energy has captured ninety per cent of a quantity defined as the energy a determinant misses; what fraction of the density error it has captured is a different and generally smaller number.

A method benchmarked against thermochemistry is benchmarked against energies. Atomisation energies, reaction enthalpies and barrier heights are all differences of total energies, so a method tuned to reproduce them has been tuned on the second-order quantity throughout. Its performance on charge distributions, on multipole moments, on spin densities and on anything else is not established by that tuning and has to be tested separately.

The same behaviour in a larger system

One system could be a special case, so the calculation is run on a four-site chain and a six-site chain as well, where the many-body space is twenty-five times and nine hundred times larger.

The exponents come out at 2.0002.000 and 0.9940.994 for the four-chain, and 2.0002.000 and 0.9940.994 for the six-chain. The partner state chosen by the double-occupancy matrix element is number twenty for the first and number ninety-nine for the second — far up the spectrum in both cases, which is itself worth noticing.

The high partner is not an artefact. The double-occupancy operator is diagonal in the configuration basis and connects the ground state to whatever configurations have unusual double occupancy, and those are the high-energy ones. So the part of a wavefunction error that an observable notices most keenly is often the part that costs the most energy — which cuts both ways. A method that controls its energy well is controlling the very part of the error the observable is most sensitive to; and the residual error it leaves is in the part the energy cannot see.

That is the honest reading of the whole exercise. The energy is not blind. It is quadratically insensitive, which is a weaker statement than blind and a much weaker one than proportional.

Why the principle holds at all, and where it stops

The variational bound comes from expanding a trial function in the exact eigenstates. Writing ψ=kckk|\psi\rangle = \sum_k c_k |k\rangle with ck2=1\sum |c_k|^2 = 1,

ψH^ψ=kck2EkE0kck2=E0\langle \psi | \hat{H} | \psi \rangle = \sum_k |c_k|^2 E_k \ge E_0 \sum_k |c_k|^2 = E_0

because every EkE0E_k \ge E_0. That is the entire proof, and reading it shows exactly what it needs.

It needs the exact Hamiltonian. A calculation with an approximate Hamiltonian — a pseudopotential, a truncated interaction, a density functional — has no such bound, and its energy can and does come out below the true value. Density functional energies are not variational in this sense at all, which is why a lower one is not evidence of a better one.

It needs the ground state. The bound applies to the lowest state of a given symmetry, and an excited state calculated variationally can fall below its own answer by mixing in the ground state.

And it needs the energy. There is no analogous bound for any other operator, which is not an oversight in the theory but the direct consequence of the algebra above: the inequality came from EkE0E_k \ge E_0 for every kk, and no other observable has a spectrum ordered the same way as the energy is.

The energy is the last thing a wrong wavefunction gets wrong. Two errors against the error in the wavefunction, on log axes, for a chain of 4 at U = 6t. The energy's line has slope 2.00 and the double occupancy's has slope 0.99: the first is second order in the error and the second is first order. So the two lines diverge as the wavefunction improves, and the energy stops being evidence about anything else long before it stops improving.
Fig. 4 The same measurement on four sites rather than two. The energy is still the last thing a wrong wavefunction gets wrong, and the margin has widened: with more orbitals there are more ways for a trial function to be wrong in directions the energy does not see. Nothing about the argument is a two-site accident.

A calculation that is exactly right about one thing and wrong about another

The clearest chemical case of the pattern is one already computed here, and it is worth revisiting with the exponent argument in hand.

Where molecular orbital theory dissociates shows a restricted single-determinant description of a dissociating bond putting half its weight on ionic configurations no matter how far the atoms are pulled apart. At large separation that wavefunction is qualitatively wrong: it says the molecule dissociates half the time into a positive and a negative ion.

Its energy error, at that point, is a fixed fraction of the interaction — bounded, and by the standards of the subject not enormous. Its double occupancy error is a factor of infinity, since the correct value goes to zero and the determinant’s stays at a quarter.

The two errors are not comparable and they were never going to be. What the second-order argument adds is that this is the generic case rather than the pathological one: any error in a wavefunction shows up more in an observable than in the energy, and dissociation is simply where the error becomes large enough for “more” to be visible without a log plot.

The same reading applies to the correlation question of the hole that is not repulsion. A determinant misses the Coulomb hole entirely and gets the energy roughly right, and the reason those two facts sit together is the exponent.

What follows for the numbers computed here

The discipline this essay is about applies to the calculations here as much as to anybody else’s, and it is worth saying which of them it touches.

The exact diagonalisations are exact, so the argument does not apply to them. There is no trial function in a full configuration calculation on four sites; the vector is the vector.

The one-electron calculations are not variational at all. A Hückel eigenvalue is the eigenvalue of a stated matrix, and the question of whether it is close to something is a question about the matrix rather than about a wavefunction. Hückel theory is explicit that its parameters are absent and its numbers are in units of them.

The one place the argument bites here is the variational exponent of the atom does not bring its own orbital, which optimises a single parameter by minimising an energy. That calculation gets the bond length right and the binding energy wrong by sixteen per cent — and the pattern is exactly the one above, since bond length is a position of a minimum and therefore a first-derivative property while the depth is the second-order quantity. Reading its success on the length as evidence about the wavefunction would be the mistake this essay is about.

The share of the error a transferred correction removes. How much of the mean field's error the composite recipe takes away, against how unlike the two sites are made. Near the reference it removes almost all of it, which is why the practice exists. By ε = 8 it removes a negative share: the correction computed on the symmetric molecule is larger than the one the asymmetric molecule needs, so adding it makes the answer worse than the cheap calculation it was added to. Values below the axis are clipped at −120 per cent.
Fig. 5 What a transferred correction actually removes. If the energy error were a property of the method it would subtract cleanly from one system to the next; the share it removes is plotted here and it is not one, so a correction fitted on one system carries only part of itself to another — which is the practical form of the same defect.

Which properties are safe, and which are not

There is a useful distinction hiding in the algebra, and it separates properties into two classes rather than leaving them all in one.

A property that is a derivative of the energy with respect to something inherits the energy’s good behaviour, at least partly. A dipole moment computed as the derivative of the energy with respect to an applied field is not the same object as a dipole moment computed as an expectation value of the position operator, and for a variationally optimised wavefunction the two are not equal. The first is the one with the better error behaviour, because it is a property of the energy surface rather than of the state.

A property that is a plain expectation value has the first-order error in full. Double occupancy is one. So are spin densities, populations, and every quantity read off a wavefunction directly.

The equality of the two routes — the Hellmann–Feynman theorem — holds only when the wavefunction is fully optimised with respect to every parameter it has, which for a finite basis it is not. The gap between them is a diagnostic in its own right, and a large one is a sign of exactly the error this essay measures.

That distinction cannot be computed in these models, which have no energy derivatives and no applied fields. It is stated because it is the standard practical response to the problem, and a reader meeting the second-order argument for the first time should know that the response exists.

The instrument that is first order on purpose

Everything above is a warning without a remedy: the energy is quadratically insensitive, so a good energy is weak evidence, and the exact answer is not available to check anything against. It is worth asking whether any quantity computable from a trial function alone is sensitive to the wavefunction error at first order. One is, and the construction used here gives it exactly.

Take the variance of the Hamiltonian, σ2=H^2H^2\sigma^2 = \langle \hat{H}^2 \rangle - \langle \hat{H} \rangle^2. It vanishes if and only if the state is an eigenstate, which is a statement about the state rather than about its energy, and nothing outside the trial function is needed to evaluate it. For the mixed states above the algebra closes in one line, because the two components are eigenvectors:

σ(ε)=ε(EmE0)1+ε2.\sigma(\varepsilon) = \frac{\varepsilon\,(E_m - E_0)}{1 + \varepsilon^2}.

The numerator carries a single power of ε\varepsilon. The variance is first order in the wavefunction error, exactly where the energy is second, and it is on the same footing as the observables that were being got wrong.

The two can be put side by side at the worked point. At ε=0.2\varepsilon = 0.2 the energy error is 0.0385(EmE0)0.0385\,(E_m - E_0) and σ\sigma is 0.192(EmE0)0.192\,(E_m - E_0) — five times larger, and within two per cent of the wavefunction error of 0.1960.196 that produced both. The variance is not merely sensitive; scaled by the gap it very nearly is the error in the state.

That is why the quantity is standard equipment in quantum Monte Carlo, where a trial function is optimised on the variance rather than on the energy, and where energies computed at several variances are extrapolated to the zero-variance limit. Both practices are the exponent argument used constructively: minimise the thing that moves at first order, then use the thing that moves at second order to report a number.

It also sharpens what the essay is claiming. The complaint is not that a wavefunction’s quality is unknowable without the answer — it is that the energy does not report it, while a quantity sitting one line further along in the same algebra does.

The instrument that judges a model

The essays on approximation argue that an orbital is a device rather than a thing, and that the ordering of orbital energies is a property of a model rather than of an atom. This one is about the instrument that decides whether a model is any good.

Its conclusion is not that the variational principle is untrustworthy. It is the most trustworthy thing available and it is trustworthy about exactly one number. The mistake is the inference — a good energy, therefore a good wavefunction, therefore good everything — and the inference fails by an amount that has been measured here rather than argued about: exponent two against exponent one, and a ratio that doubles at every halving.

The practical form of the lesson is a habit rather than a formula. A property that was not the one minimised has not been validated by the minimisation, and it needs a test of its own. This site’s habit of giving every claim a check that could refuse it is the same habit applied to its own arithmetic, and the numbers above are the reason it is not optional.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationConvergenceDouble occupancyElectron correlationExact diagonalisationExpectation valueHubbard modelModel limitOne-electron modelsReference state