A better energy is not a better answer
Worth reading first: Orbitals are not where the electron is · The smallest many-electron calculation.
The variational principle is the most reliable tool in quantum chemistry and it is reliable in one direction only. For any trial function whatever, the expectation value of the Hamiltonian is at least the true ground-state energy. Lower is better, and no calculation can accidentally come out below the answer.
That one-sidedness is why energies are the currency of the subject. Two methods can be ranked without knowing the answer; a basis set can be enlarged and the energy watched to see whether it has converged; a parameter can be optimised by minimising. None of that works for any other property, and the reason is worth understanding rather than accepting.
The energy is stationary at the exact wavefunction and nothing else is. Make an error in the state and the energy is wrong by order ; every other expectation value is wrong by order . The property that makes the energy safe to minimise is the same property that makes it uninformative about everything else.
Making the wavefunction error an input rather than an estimate
The usual difficulty with a claim of this kind is that the error in a wavefunction is not measurable — the exact one is not available, which is why an approximation was being made.
Here it is available. The Hubbard systems used here are diagonalised completely, so the exact ground state is a vector and so is every excited state. A trial function whose error is known can therefore be built by hand:
whose overlap with the exact ground state is , so the error in the state is exactly. Nothing is estimated, nothing is fitted, and there is no variational calculation anywhere in the construction.
The energy of such a state is also exact, because the two components are eigenvectors and the cross term vanishes:
The is visible on the page, which is the whole of the second-order behaviour. An ordinary observable has no such cancellation: its expectation value picks up a cross term , which is linear.
Fitting the two errors against over a decade returns exponents of and .
The mistake the check made first, which sharpened the claim
The first version of the calculation mixed in state number one — the next state up — and measured the double occupancy’s error at exponent two, apparently refuting the whole argument.
The refutation was real and it was about the dimer’s symmetry rather than about the principle. The first excited state of the two-site system has the opposite parity from the ground state, so vanishes identically, the cross term is exactly zero, and the linear contribution is not there to be measured.
The repair is to choose the partner by the overlap that matters rather than by energy: mix in the state the observable actually connects to. For the dimer that is state three, with .
That produces a sharper statement than the original. An observable’s error is first order in the part of the wavefunction error that the observable connects to, and second order in the rest. A calculation can therefore be badly wrong in a way that one property cannot see and another can, and which property notices is decided by a matrix element rather than by how big the error is.
What the numbers look like at a realistic error
Exponents are abstract; one worked point is not.
Take . The trial state overlaps the exact ground state by , so the wavefunction is wrong by about two per cent — the sort of error a decent approximate method makes, and small enough that most people would describe the state as essentially right.
Its double occupancy is out by 53 per cent.
Halve the error to . The energy error falls by four and the double occupancy’s by two, so the second is now a quarter of a per cent of what it was relative to the first. Halve again and the ratio doubles again. There is no error small enough for the two to converge together, and the gap between them grows without bound precisely as the calculation gets better.
That is the counterintuitive part and it is worth stating baldly. Improving a calculation makes the energy relatively more trustworthy and every other property relatively less so, in the sense that the ratio of the two errors diverges. A method that gets energies to a millihartree is not a method that gets densities to a comparable accuracy; it is a method whose density error is the square root of its energy error, up to constants.
Where this shows up in real calculations
The behaviour is not a curiosity of a two-site model; it is the reason for several standing pieces of practical advice, and it explains them rather than merely restating them.
Basis sets converged on energies are not converged on properties. A basis is usually enlarged until the total energy stops moving, and the dipole moment computed in it is usually still moving. That is the exponent-one behaviour: the energy has reached its second-order plateau while the wavefunction is still first-order wrong.
Hartree–Fock dipole moments are systematically too large, typically by ten to fifteen per cent for polar molecules, even though Hartree–Fock energies for the same molecules are within a per cent of the exact value. The energy error and the moment error are not comparable numbers and never were.
Recovering a large fraction of the correlation energy is a poor guarantee. A method described as capturing ninety per cent of the correlation energy has captured ninety per cent of a quantity defined as the energy a determinant misses; what fraction of the density error it has captured is a different and generally smaller number.
A method benchmarked against thermochemistry is benchmarked against energies. Atomisation energies, reaction enthalpies and barrier heights are all differences of total energies, so a method tuned to reproduce them has been tuned on the second-order quantity throughout. Its performance on charge distributions, on multipole moments, on spin densities and on anything else is not established by that tuning and has to be tested separately.
The same behaviour in a larger system
One system could be a special case, so the calculation is run on a four-site chain and a six-site chain as well, where the many-body space is twenty-five times and nine hundred times larger.
The exponents come out at and for the four-chain, and and for the six-chain. The partner state chosen by the double-occupancy matrix element is number twenty for the first and number ninety-nine for the second — far up the spectrum in both cases, which is itself worth noticing.
The high partner is not an artefact. The double-occupancy operator is diagonal in the configuration basis and connects the ground state to whatever configurations have unusual double occupancy, and those are the high-energy ones. So the part of a wavefunction error that an observable notices most keenly is often the part that costs the most energy — which cuts both ways. A method that controls its energy well is controlling the very part of the error the observable is most sensitive to; and the residual error it leaves is in the part the energy cannot see.
That is the honest reading of the whole exercise. The energy is not blind. It is quadratically insensitive, which is a weaker statement than blind and a much weaker one than proportional.
Why the principle holds at all, and where it stops
The variational bound comes from expanding a trial function in the exact eigenstates. Writing with ,
because every . That is the entire proof, and reading it shows exactly what it needs.
It needs the exact Hamiltonian. A calculation with an approximate Hamiltonian — a pseudopotential, a truncated interaction, a density functional — has no such bound, and its energy can and does come out below the true value. Density functional energies are not variational in this sense at all, which is why a lower one is not evidence of a better one.
It needs the ground state. The bound applies to the lowest state of a given symmetry, and an excited state calculated variationally can fall below its own answer by mixing in the ground state.
And it needs the energy. There is no analogous bound for any other operator, which is not an oversight in the theory but the direct consequence of the algebra above: the inequality came from for every , and no other observable has a spectrum ordered the same way as the energy is.
A calculation that is exactly right about one thing and wrong about another
The clearest chemical case of the pattern is one already computed here, and it is worth revisiting with the exponent argument in hand.
Where molecular orbital theory dissociates shows a restricted single-determinant description of a dissociating bond putting half its weight on ionic configurations no matter how far the atoms are pulled apart. At large separation that wavefunction is qualitatively wrong: it says the molecule dissociates half the time into a positive and a negative ion.
Its energy error, at that point, is a fixed fraction of the interaction — bounded, and by the standards of the subject not enormous. Its double occupancy error is a factor of infinity, since the correct value goes to zero and the determinant’s stays at a quarter.
The two errors are not comparable and they were never going to be. What the second-order argument adds is that this is the generic case rather than the pathological one: any error in a wavefunction shows up more in an observable than in the energy, and dissociation is simply where the error becomes large enough for “more” to be visible without a log plot.
The same reading applies to the correlation question of the hole that is not repulsion. A determinant misses the Coulomb hole entirely and gets the energy roughly right, and the reason those two facts sit together is the exponent.
What follows for the numbers computed here
The discipline this essay is about applies to the calculations here as much as to anybody else’s, and it is worth saying which of them it touches.
The exact diagonalisations are exact, so the argument does not apply to them. There is no trial function in a full configuration calculation on four sites; the vector is the vector.
The one-electron calculations are not variational at all. A Hückel eigenvalue is the eigenvalue of a stated matrix, and the question of whether it is close to something is a question about the matrix rather than about a wavefunction. Hückel theory is explicit that its parameters are absent and its numbers are in units of them.
The one place the argument bites here is the variational exponent of the atom does not bring its own orbital, which optimises a single parameter by minimising an energy. That calculation gets the bond length right and the binding energy wrong by sixteen per cent — and the pattern is exactly the one above, since bond length is a position of a minimum and therefore a first-derivative property while the depth is the second-order quantity. Reading its success on the length as evidence about the wavefunction would be the mistake this essay is about.
Which properties are safe, and which are not
There is a useful distinction hiding in the algebra, and it separates properties into two classes rather than leaving them all in one.
A property that is a derivative of the energy with respect to something inherits the energy’s good behaviour, at least partly. A dipole moment computed as the derivative of the energy with respect to an applied field is not the same object as a dipole moment computed as an expectation value of the position operator, and for a variationally optimised wavefunction the two are not equal. The first is the one with the better error behaviour, because it is a property of the energy surface rather than of the state.
A property that is a plain expectation value has the first-order error in full. Double occupancy is one. So are spin densities, populations, and every quantity read off a wavefunction directly.
The equality of the two routes — the Hellmann–Feynman theorem — holds only when the wavefunction is fully optimised with respect to every parameter it has, which for a finite basis it is not. The gap between them is a diagnostic in its own right, and a large one is a sign of exactly the error this essay measures.
That distinction cannot be computed in these models, which have no energy derivatives and no applied fields. It is stated because it is the standard practical response to the problem, and a reader meeting the second-order argument for the first time should know that the response exists.
The instrument that is first order on purpose
Everything above is a warning without a remedy: the energy is quadratically insensitive, so a good energy is weak evidence, and the exact answer is not available to check anything against. It is worth asking whether any quantity computable from a trial function alone is sensitive to the wavefunction error at first order. One is, and the construction used here gives it exactly.
Take the variance of the Hamiltonian, . It vanishes if and only if the state is an eigenstate, which is a statement about the state rather than about its energy, and nothing outside the trial function is needed to evaluate it. For the mixed states above the algebra closes in one line, because the two components are eigenvectors:
The numerator carries a single power of . The variance is first order in the wavefunction error, exactly where the energy is second, and it is on the same footing as the observables that were being got wrong.
The two can be put side by side at the worked point. At the energy error is and is — five times larger, and within two per cent of the wavefunction error of that produced both. The variance is not merely sensitive; scaled by the gap it very nearly is the error in the state.
That is why the quantity is standard equipment in quantum Monte Carlo, where a trial function is optimised on the variance rather than on the energy, and where energies computed at several variances are extrapolated to the zero-variance limit. Both practices are the exponent argument used constructively: minimise the thing that moves at first order, then use the thing that moves at second order to report a number.
It also sharpens what the essay is claiming. The complaint is not that a wavefunction’s quality is unknowable without the answer — it is that the energy does not report it, while a quantity sitting one line further along in the same algebra does.
The instrument that judges a model
The essays on approximation argue that an orbital is a device rather than a thing, and that the ordering of orbital energies is a property of a model rather than of an atom. This one is about the instrument that decides whether a model is any good.
Its conclusion is not that the variational principle is untrustworthy. It is the most trustworthy thing available and it is trustworthy about exactly one number. The mistake is the inference — a good energy, therefore a good wavefunction, therefore good everything — and the inference fails by an amount that has been measured here rather than argued about: exponent two against exponent one, and a ratio that doubles at every halving.
The practical form of the lesson is a habit rather than a formula. A property that was not the one minimised has not been validated by the minimisation, and it needs a test of its own. This site’s habit of giving every claim a check that could refuse it is the same habit applied to its own arithmetic, and the numbers above are the reason it is not optional.
What links here
Computed from the collection rather than written here: the essays that point at this one.
- A method that is not additive
- The correction that was computed somewhere else
- Two wrong numbers and a right difference
- A contraction is a decision made once
- A Gaussian is the wrong shape
- The atom does not bring its own orbital
- The half that cannot be computed
- The measurement a basis was not fitted to
- and 10 more
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays naming at least two of the same things, that neither author linked.
- Where the electrons are, without subtracting anything — both name double occupancy, electron correlation, exact diagonalisation, expectation value, hubbard model, model limit, reference state
- Half of it is given back at one bond — both name electron correlation, exact diagonalisation, expectation value, hubbard model, model limit, reference state
- The half of the square a ring of four cannot show — both name approximation, electron correlation, exact diagonalisation, hubbard model, model limit, reference state
- A contrast with a closed form — both name convergence, electron correlation, exact diagonalisation, hubbard model, model limit
- A count rather than an average — both name approximation, convergence, exact diagonalisation, model limit, reference state
- A mean field cannot get out of the way — both name double occupancy, electron correlation, exact diagonalisation, hubbard model, model limit
Named objects
A dashed tag is an object no other essay names yet.
ApproximationConvergenceDouble occupancyElectron correlationExact diagonalisationExpectation valueHubbard modelModel limitOne-electron modelsReference state