Orbitals

The property that gets worse

One Gaussian fitted to hydrogen gets the mean radius exactly right — 1.500000, against an exact 1.5. Two Gaussians get it wrong by 1.4 per cent, which is three hundred and thirty-two thousand times further out, while the energy improves fivefold. The variational principle bounds one number and says nothing whatever about any other, and the sequence of errors in everything else need not even be monotone.

Worth reading first: The measurement a basis was not fitted to · A Gaussian is the wrong shape.

The variational principle is the one guarantee in this subject that comes with a proof: an approximate wavefunction’s energy is above the exact one, so lowering the energy moves towards the truth in that one respect.

Every other respect is unguaranteed, and the sequence of Gaussian bases fitted to hydrogen is small enough to show exactly how unguaranteed.

A basis fitted to an energy does badly on a quantity an experiment measures directly: the Compton profile of a six-Gaussian fit is twenty-five times further out than its energy. This essay asks the question that comes before that one: does a property have to improve at all?

Six bases, and what each one is wrong about

Six sets of exponents, each optimised to give the lowest energy for that number of functions, and five quantities computed from each.

One of these is guaranteed to improve, and it is not the one anybody measures. The relative error in the energy and in four properties of the fitted function, against the number of Gaussians. The energy falls at every step, because that is what the variational principle promises. The mean radius is exact for the single-function basis and 332 thousand times worse for the two-function one, and the density at the nucleus is still 6.2 per cent wrong where the energy is wrong by 0.011 per cent.
Fig. 1 The relative error in the energy and in four properties, against the number of Gaussians. Only the energy is guaranteed to fall, and it does, at every step. The mean radius does not: it is exact for one function and wrong by more than a per cent for two. The density at the nucleus is still 6.2 per cent wrong at six functions, where the energy is wrong by 0.011.

The energy line is the whole content of the variational principle: three orders of magnitude, monotone, by construction. Every other line is what the principle does not cover.

⟨r²⟩ falls monotonically — from 12 per cent wrong to 0.11 — but it is never as accurate as the energy at the same basis size.

⟨1/r⟩ converges fastest of the properties, which makes sense: it is the potential energy divided by the nuclear charge, so it is half of the quantity being minimised.

The density at the nucleus is the slowest. A Gaussian is flat where the exact function has a cusp, so the amplitude at the origin is systematically too small; six functions bring it from 76 per cent wrong to 6.2, and that gap is what an electron spin resonance experiment would be measuring.

And ⟨r⟩ is not monotone at all.

The accident

The single Gaussian is the worst description of hydrogen anybody would write down. Its energy is −0.4244 hartree against an exact −0.5, which is 15 per cent wrong. Its shape is wrong at the nucleus and wrong in the tail. Its mean radius is 1.500000.

That is not a small error. It is the exact answer, to eight decimal places, and the reason is worth following because it is a warning rather than a curiosity.

For a single normalised Gaussian of exponent α\alpha the kinetic energy is 3α/23\alpha/2 and the nuclear attraction is Z8α/π-Z\sqrt{8\alpha/\pi}. Differentiate with respect to α\alpha and set the derivative to zero:

32=Z2πα\frac{3}{2} = Z\sqrt{\frac{2}{\pi\alpha}}

The right-hand side is the mean radius of that Gaussian. So the stationarity condition, written out, is the equation r=3/2Z\langle r \rangle = 3/2Z — and hydrogen’s exact mean radius is 3/2Z3/2Z. The agreement is forced by the optimisation rather than earned by the function.

The wrong shape, fitted as well as it can be. The exact hydrogen 1s orbital and the best sums of one, two, three and six Gaussians, each with its exponents optimised for the energy. Three of them already reproduce the exact function to 99.94 per cent by overlap, which is why the method works at all — and the two places it goes wrong, at the nucleus and far out, are exactly where the other faces of this figure look.
Fig. 2 The functions themselves, against the exact one. The single Gaussian is visibly the wrong shape everywhere: too flat at the nucleus, too fat in the middle, and dead in the tail. It is this function whose mean radius is exact.

The same function’s ⟨1/r⟩ is 8Z/3π8Z/3\pi, which is 0.8488 against an exact 1 — 15.1 per cent low. One function, two properties that are reciprocals of one another, one of them exact to eight figures and the other wrong in the second.

That is the sharpest form of the warning this essay exists for. Agreement on a property is not evidence about a wavefunction, because there is always a property some bad function gets right, and nothing marks which property that is.

What happens next in the sequence

Adding a second Gaussian lowers the energy by a factor of five and makes ⟨r⟩ 1.42 per cent too small. In absolute terms the error goes from 6.4 × 10⁻⁸ to 2.1 × 10⁻², a factor of three hundred and thirty-two thousand.

From there it recovers: 0.57 per cent at three functions, 0.19 at four, 0.064 at five, 0.022 at six. At six functions it is still five thousand times further from the truth than the single Gaussian was.

The recovery is what makes the whole sequence readable. If ⟨r⟩ simply diverged, something would be wrong with the optimiser. It converges, from below, at about the rate the other properties converge at — and the first member of the sequence sits far off that line, on the correct value, for a reason that has nothing to do with being correct.

The energy converges and the shape does not. The variational energy of one electron on one nucleus, in a basis of the stated number of Gaussians with their exponents optimised here rather than quoted. Every point is above the exact −0.5 hartree, as a variational calculation must be, and adding a function always lowers it: -0.42441, -0.48581, -0.49698, -0.49995. What does not improve is the slope at the nucleus, which is exactly zero for every one of them and should be −1.
Fig. 3 The energy alone, which is the quantity every one of these bases was chosen to minimise. This curve is what “converged” usually means, and it is the only curve in this essay with a theorem behind it.

The check that is always satisfied

There is a second quantity in wide use as a convergence check, and the same six bases dispose of it.

The virial theorem says that for an exact eigenstate of a Coulomb Hamiltonian the potential energy is exactly twice the kinetic energy in magnitude, with opposite sign. The ratio −⟨V⟩/⟨T⟩ is therefore 2 for the truth, and is quoted throughout computational chemistry as a check that a calculation is sound.

For every basis in this sequence it is 2.0000000.

The reason is immediate once stated: scaling all the exponents by a common factor is a variation the optimiser is free to make, and the derivative of the energy with respect to that scaling is exactly the combination that the virial theorem sets to zero. So at any stationary point of the energy with respect to the exponents the ratio is two, whether the energy is right to a hundredth of a per cent or wrong by fifteen.

A quantity that is exactly satisfied by a badly wrong answer is not a diagnostic of anything except that the optimiser finished. That is worth stating because the check is cheap, is printed by every program, and is read as reassurance.

What the basis is wrong about, in pictures

The shape errors behind the numbers are the two an exponential cannot be made of Gaussians at either end.

At the nucleus, where the shape is wrong. The first few tenths of a bohr. The exact orbital arrives at the nucleus with a corner — its slope there is exactly −Z, which is Kato's condition and is what cancels the singularity in the potential. Every sum of Gaussians arrives flat, with a slope of exactly zero, because every Gaussian is smooth at the origin and a sum of smooth functions is smooth. More functions raise the peak towards the right height and never produce the corner.
Fig. 4 The nucleus. The exact function has a kink there — its slope is discontinuous, and the size of the discontinuity is fixed by the nuclear charge. Every Gaussian has zero slope at the origin, so every sum of Gaussians does, at every basis size.
Far from the nucleus, on a logarithmic scale. The same functions plotted as the logarithm of their amplitude, scaled to agree at the nucleus. The exact orbital is a straight line here, because an exponential is; every Gaussian sum curves away downward, because e^(−αr²) falls faster than e^(−r) however small α is. At 8 bohr the best six-function fit still has only 58 per cent of the amplitude it should. Adding functions moves the departure further out and does not remove it — which is why properties that depend on the tail converge far more slowly than energies.
Fig. 5 And the tail, where a Gaussian dies as the square of the distance in the exponent and the exact function dies linearly. At five bohr the six-function fit still has amplitude; at eight it does not.

Those two defects are not fixed by more functions — they are properties of the function type — and the properties that weight those regions inherit them. The density at the nucleus weights the cusp entirely; the mean radius weights the middle, where the fit is best; ⟨r²⟩ weights the tail, where it is worst.

That ordering is visible in the first figure, and it is the only systematic thing in it. The non-monotonicity is not systematic: it is an accident of where the optimiser puts one exponent when there is only one to put.

The function all of this is trying to reproduce is hydrogen’s 1s: an exponential with a corner at the nucleus, whose radial density peaks at one bohr. The corner is a consequence of the potential being singular there, so it is a property of the physics rather than of the choice of function — and no sum of smooth functions has one.

Which size, measured five ways

There is a reason the mean radius is the property that behaves oddly, and it is that “the size of an orbital” is not one quantity to begin with.

One of these is guaranteed to improve, and it is not the one anybody measures. The relative error in the energy and in four properties of the fitted function, against the number of Gaussians. The energy falls at every step, because that is what the variational principle promises. The mean radius is exact for the single-function basis and 332 thousand times worse for the two-function one, and the density at the nucleus is still 6.2 per cent wrong where the energy is wrong by 0.011 per cent.
Fig. 6 One of these is guaranteed to improve, and it is not the one anybody measures. The variational principle promises the energy will fall as functions are added and promises nothing whatever about any other quantity — so the two columns are not two views of one convergence, and the second one is free to move the wrong way.

How big is an orbital takes that apart: each measure weights a different part of the same function, so a fit that is good in the middle and bad at the ends will score well on one and badly on another. The mean radius weights the region where a Gaussian fit is at its best, which is exactly why a bad fit can hit it — and why hitting it means so little.

The same point is what makes a stated fraction necessary. An orbital picture here states the fraction of the density it encloses rather than a level or a radius, because the fraction is the only one of those that means the same thing for two different functions. A fitted function’s ninety per cent contour and the exact one’s are comparable; their “sizes” are not.

Where the density actually is settles how much the two failures cost: the cumulative integral reaches half by 1.4 bohr and ninety per cent by 2.7, so the region the fit gets wrong at the nucleus holds almost nothing and the region it gets wrong far out holds a few per cent. That is the whole reason a wrong shape can give a right energy.

What a property is for

The moments computed here are not arbitrary. Each is what some experiment measures, and the list is short enough to give in full.

⟨r⟩ and ⟨r²⟩ appear in diamagnetic susceptibilities and in the form factors that X-ray scattering measures; ⟨r²⟩ is also the size an atom presents to a colliding partner.

⟨1/r⟩ is the electron–nuclear attraction per unit charge and is what a nuclear quadrupole coupling is scaled by. It is also the quantity a screening model adjusts when it replaces a many-electron atom with a hydrogenic one.

The number and position of the radial nodes is not on the list at all, because a fitted 1s has none to get wrong — but for any higher orbital it is the first thing a basis has to reproduce, and a node is a place where the function is exactly zero rather than small, so an approximation either has it or does not.

|ψ(0)|² is the Fermi contact term — the hyperfine coupling an electron spin resonance experiment reads directly, and the quantity that makes a hydrogen atom’s spectrum split at all.

None of them is an energy, and none of them is bounded by any variational statement. The convention that a basis is validated by its energy is a convention about what is easy to compute, not about what is being predicted — the same complaint made about a better energy not being a better answer, here with the properties named, and the a measurement the basis was not fitted to shows the same thing for a momentum-space property with a ratio of errors that grew as the basis improved.

The Compton profile, exact and fitted. The momentum density integrated over the two perpendicular directions, for the exact 1s and for three fitted bases. The exact curve is 8/3π(1 + q²)³ in closed form; the fitted ones are sums of Gaussians and are cusped differently at the origin, which is the position-space cusp showing up as a shape in momentum.
Fig. 7 The momentum-space quantity, for comparison: the Compton profile, which is a measurement of the momentum distribution and is dominated by exactly the regions of space the fit is worst in. Its error at six functions is twenty-five times the energy’s.
Where the electron is, and how fast it is going. The radial distribution in position on the left and in momentum on the right, for the same orbitals. The two run opposite ways: the 1s is the most compact in space and the widest in momentum, and every excited orbital that spreads out in one narrows in the other. Both are normalised, both are the same function, and neither is more fundamental than the other — the transform loses nothing and adds nothing.
Fig. 8 And the exact momentum distributions the profile is built from. A property computed in momentum space weights the cusp — a sharp feature in position is a broad one in momentum — so the errors a Gaussian basis makes are magnified rather than averaged away.

Why the energy may be extrapolated and a property may not

A real calculation has no exact answer to measure against, so what it does instead is run the same quantity in two or three basis sizes and extrapolate the sequence. That practice is universal for energies and it is worth asking what licenses it, because the licence turns out not to extend to anything else in this essay.

The energy’s licence has two parts and both are theorems rather than observations.

Its sequence is monotone by construction. Enlarging a basis can only lower the variational energy, so the points march in one direction and an extrapolation is at worst inaccurate. It cannot point the wrong way, because there is no wrong way to point.

And its rate is derived. The correlation energy of a correlation-consistent sequence falls off as the inverse cube of the cardinal number, and that exponent is not fitted — it comes from analysing how a partial-wave expansion converges on the cusp between two electrons, which is a statement about the shape of the exact wavefunction rather than about any basis. So an extrapolation of the energy is a two-parameter fit to a form whose exponent was supplied from outside, which is a much stronger operation than a three-parameter fit to three points.

It is worth noticing that even the energy carries a warning inside that. The Hartree–Fock part of it converges roughly exponentially with basis size and the correlation part as an inverse cube, so the two halves of one number obey different laws, and extrapolating them together with a single form is already a compromise. Standard practice extrapolates them separately for exactly this reason.

Neither part of the licence is available for a radial moment, a density at a nucleus, or a Compton profile.

There is no monotonicity, and this essay’s own table is the demonstration: the mean radius is exact at one function, wrong by 1.4 per cent at two, and improving thereafter. A sequence that turns round has an extremum in it, and no smooth decaying form has one.

And there is no derived exponent. Nothing analogous to the partial-wave argument exists for a general expectation value, so the functional form would have to be fitted along with the parameters — three unknowns from three points, which is interpolation wearing an extrapolation’s clothes.

The practical consequence is a rule that costs one extra calculation. Compute a property at three basis sizes rather than two, and look at the direction. Two points always define a monotone sequence, so a two-point extrapolation cannot detect the situation this essay is about; three points can, and a non-monotone triple is a refusal rather than a fit. In this essay’s sequence the first three points of the mean-radius line are exact, 1.4 per cent out, and improving — a triple that no decaying form can pass through, and which two points would have shown as a clean, confident, entirely wrong trend.

The rule has a converse that is worth as much. A property whose three points are monotone and whose differences are falling geometrically is behaving as though a convergence law governs it, and extrapolating it is a reasonable risk taken knowingly. What is not reasonable is doing so without looking, which is what two points forces.

None of this makes an extrapolated property useless. It makes it a quantity with an unstated assumption in it — that the sequence is doing what the energy’s sequence is guaranteed to do — and this essay’s whole content is that the guarantee applies to one number and was never about the rest.

Where the model stops

Three things about this exercise are smaller than they look, and saying so is the difference between a warning and an overreach.

This is one electron in a Coulomb field, with an exact answer available for every quantity, which is why every error above is an error rather than a difference between two approximations. In a many-electron calculation the basis error and the correlation error are entangled and neither is separately visible — the smallest system where both can be watched at once has two electrons and no basis error at all, which is the opposite corner of the same difficulty.

The bases here are fully optimised, one to six functions, each at its own variational minimum. Standard basis sets are not: they are contracted, fitted once for a set of atoms, and not re-optimised for the molecule they are used in, which adds a different error on top of these.

Non-monotonicity is not the usual case. Four of the five quantities here converge monotonically, and the one that does not is exceptional because its first member is exceptional. The claim is not that properties routinely get worse; it is that nothing forbids it, and that the diagnostic usually relied on to notice would not.

There is a second cost to adding functions, measured separately: the smallest eigenvalue of the overlap matrix falls as the basis grows, so a basis large enough to fix the tail is a basis approaching linear dependence. Both difficulties are properties of the shape being wrong rather than of the basis being small.

What is quoted, and what is computed

Nothing here is quoted. The exponents are found by optimisation, the energies by solving a generalised eigenvalue problem in the fitted basis, the moments by quadrature on a mapped radial grid, and the exact values from the closed forms for a hydrogenic 1s.

The closed form for the optimal single Gaussian — α=8Z2/9π\alpha = 8Z^2/9\pi, r=3/2Z\langle r \rangle = 3/2Z, 1/r=8Z/3π\langle 1/r \rangle = 8Z/3\pi — is derived rather than quoted, and it is checked against the numerically optimised one-function basis to make sure the two are the same object.

What was checked

The energy falls at every step. Without this the optimiser has not found its minima and no other line can be read.

At least one property is further from the truth in a larger basis than in a smaller one, and the worst such step is worse by orders of magnitude rather than by rounding — a factor of 3.3 × 10⁵.

The virial ratio is two at every basis size, to a part in a million, including the sizes whose energy is badly wrong.

The optimal single Gaussian’s mean radius equals the exact one to machine precision, and its mean inverse radius does not, being out by more than a tenth. Both halves are needed: the first alone would look like evidence that the function is good.

Still open: an error bar, and the contraction’s share

The natural open question is the quantity this essay keeps circling and never computes: an error bar. Every number above is an error against a known exact answer, which is available here and is not available in any calculation anybody would actually run.

What is available in a real calculation is a sequence, and the section above says why extrapolating one is licensed for an energy and not for a property. What it does not supply is the quantity a practitioner actually wants: how far from the truth an extrapolated property is, on a case where the truth is known. Every ingredient for that is in the table here, and what it needs is more basis sizes than six functions of one shape can honestly provide.

There is a second continuation, further off, and it is the one that would make these arguments bite on real calculations rather than on a hydrogen atom. Every claim here is about a fitted basis; the bases in use are contracted once for an atom and used everywhere, and the quantity that would matter is how much of the error in a property is the contraction’s rather than the function type’s. That is a computation within reach, and it needs a second atom to have anything to compare.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBasis setClosed formConvergenceExpectation valueGaussianModel limitMost probable radiusProbability densityRadial momentVariationalVirial theorem