Orbitals

The measurement a basis was not fitted to

Six Gaussians reproduce hydrogen's energy to eleven parts in a hundred thousand and its Compton profile to three parts in a thousand — twenty-five times worse, on a quantity an X-ray scattering experiment measures directly. The gap between the two errors widens as the basis is improved, because the energy is the one property a variational fit is best at.

Worth reading first: A Gaussian is the wrong shape · A contraction is a decision made once.

Every published Gaussian basis set was fitted to an energy. The exponents were chosen by minimising one, the contraction coefficients were frozen at the combination that gave the best one, and the number quoted when a basis is compared with another is almost always an energy.

Three earlier arguments are about what that costs in position space: a Gaussian has the wrong shape at both ends, a contraction freezes a decision made on an isolated atom, and adding a function makes the basis less independent as well as more complete. Each of them compared one calculation with another.

This essay compares a calculation with a measurement, and the measurement is chosen because it is the one that sees exactly what a fitted basis gets wrong.

What a Compton profile is

Scatter an X-ray off an electron and the photon’s energy loss depends on the electron’s momentum along the scattering direction. The measured lineshape is therefore the electron momentum density integrated over the two perpendicular directions,

J(q)=2πqpn(p)dpJ(q) = 2\pi \int_q^\infty p\, n(p)\, \mathrm{d}p

and for a hydrogenic 1s it is 8Z5/3π(Z2+q2)38Z^5/3\pi(Z^2+q^2)^3 exactly. At zero momentum that is 8/3π=0.848838/3\pi = 0.84883.

This is not a derived quantity or a comparison against a better calculation. It is what a spectrometer returns, and it has been measured for hydrogen, helium and the light molecules to a per cent or better since the 1960s.

The Compton profile, exact and fitted. The momentum density integrated over the two perpendicular directions, for the exact 1s and for three fitted bases. The exact curve is 8/3π(1 + q²)³ in closed form; the fitted ones are sums of Gaussians and are cusped differently at the origin, which is the position-space cusp showing up as a shape in momentum.
Fig. 1 The exact profile and three fitted ones. Every curve here integrates to exactly one electron, so the differences between them are shape and nothing else. The one-Gaussian fit is visibly the wrong shape; the three-Gaussian fit looks right at this scale and is not.

Everything drawn is in closed form. The transform of a Gaussian is a Gaussian, the transform of the exact 1s is a Lorentzian raised to a power, and the integral defining the profile is elementary for both — so no quadrature enters anywhere and the comparison is exact arithmetic against an exact expression.

The two errors, side by side

Two errors for one basis, at several sizes. For each of three fitted bases: the error in the energy, the error in the Compton profile at zero momentum, the largest relative error anywhere in the profile, and the ratio of the second column to the first. The energy improves faster than the shape does, so the ratio grows as the basis is made better.
Fig. 2 Three fitted bases, with the error in the energy and the error in the profile in adjacent columns. The energy error falls by a factor of 1,389 from one Gaussian to six; the profile error falls by 43. The last column is the ratio between the two, and it grows.

At six Gaussians the energy is right to 0.011 per cent and the profile at zero momentum is wrong by 0.27 per cent. That is twenty-five times worse, on a quantity that can be measured to about a per cent — which is to say the discrepancy is inside experimental reach while the energy error is not.

The ratio between the two errors is the number that matters, because it is not constant: 0.8 for one Gaussian, 4.8 for three, 24.9 for six. Improving the basis improves both and improves the energy faster, so the two diverge. A basis chosen because its energy has converged is a basis whose other properties have converged less, and by a margin that grows with how carefully the energy was converged.

Why the energy is the flattering one

This is not a fact about Gaussians. It is the variational principle, and a better energy is not a better answer makes the general statement: if a trial wavefunction differs from the exact one by ε, its energy differs by ε², while every other expectation value differs by ε.

So the energy is the property a variational fit is best at, in a precise sense — it is the only one that is second order in the error — and the property everybody uses to judge a basis is the one property guaranteed to understate what is wrong with it.

Two errors for one basis, at several sizes. For each of five fitted bases: the error in the energy, the error in the Compton profile at zero momentum, the largest relative error anywhere in the profile, and the ratio of the second column to the first. The energy improves faster than the shape does, so the ratio grows as the basis is made better.
Fig. 3 Two errors for one basis, at five sizes, in a single table. The energy error falls by a factor of four for every doubling of the basis and the profile error falls by much less — so the two are not two views of one convergence, and a basis chosen by the first is not chosen by the second.

The Compton profile is a first-order property. So the twenty-five-times ratio at six Gaussians is not a defect of the fit — it is the square root of the energy error made visible.

The Compton profile, exact and fitted. The momentum density integrated over the two perpendicular directions, for the exact 1s and for five fitted bases. The exact curve is 8/3π(1 + q²)³ in closed form; the fitted ones are sums of Gaussians and are cusped differently at the origin, which is the position-space cusp showing up as a shape in momentum.
Fig. 4 All five fitted profiles at once over the low-momentum region, where the disagreement lives. The curves approach the exact one from below in an orderly way, and the region they approach it in most slowly is the origin — which is the part of the profile a measurement returns best.

The one thing the energy does fix

The energy is not blind to the profile. It fixes exactly one number in it, exactly, and identifying which one explains everything above.

The second moment of the profile is the mean square momentum along one direction, which is a third of the mean square momentum, which is two thirds of the kinetic energy:

q2J(q)dq=13p2=23T\int q^2 J(q)\, \mathrm{d}q = \tfrac{1}{3}\langle p^2 \rangle = \tfrac{2}{3}\langle T \rangle

and at a variational minimum with fully optimised exponents the virial theorem makes T=E\langle T \rangle = -E. So the second moment of a fitted profile is two thirds of minus its own energy, and that is checked here rather than assumed: for the one-, three- and six-Gaussian fits the moment comes out at 0.282942, 0.331320 and 0.333297 against energies of −0.424413, −0.496979 and −0.499946, agreeing to the last digit computed.

So a basis whose energy has converged has one moment of its profile converged with it, and the exact profile’s own second moment is a third.

Everything else about the shape is free. The profile has an infinite number of features and the energy pins one combination of them; a fit can therefore be right about that combination and wrong about the curve in the way the figures show — high in one place, low in another, and the errors cancelling in exactly the one integral the fitting looked at.

That is a more precise statement than “the energy hides it”, and it is the reason the discrepancy is largest at low momentum: the second moment weights large q, so the region a fit is least constrained in is the region near the origin, which is also the region a measurement reports best.

Where in momentum the fit is wrong. The Compton profile of each fitted basis minus the exact one, as a fraction of the exact one, across momentum. The curves cross zero several times rather than approaching it from one side: a fit is a compromise, and a compromise is right in some places by being wrong in others.
Fig. 5 Where in momentum the fit is wrong, rather than by how much. The error is not spread evenly: it is concentrated at high momentum, which is the region a position-space energy weights least and a scattering experiment weights most — and that mismatch is the whole reason the two errors converge at different rates.

A tight Gaussian is a broad one

The other half of the account is what the transform does to the functions themselves, and it is simple enough to be worth stating.

A Gaussian of exponent aa transforms into a Gaussian of exponent 1/4a1/4a. So the tightest function in a basis — the one put there to describe the region near the nucleus — becomes the broadest function in momentum, and the most diffuse one becomes the narrowest. A basis spanning four orders of magnitude in position spans four in momentum, reversed.

Where the electron is, and how fast it is going. The radial distribution in position on the left and in momentum on the right, for the same orbitals. The two run opposite ways: the 1s is the most compact in space and the widest in momentum, and every excited orbital that spreads out in one narrows in the other. Both are normalised, both are the same function, and neither is more fundamental than the other — the transform loses nothing and adds nothing.
Fig. 6 The exact momentum functions of a 1s and a 2s. Neither has a cusp — the cusp is in position — and both die as a power rather than exponentially, which is the feature no finite sum of Gaussians has, because every Gaussian dies faster than any power.

The consequence for the fit is that the two spaces make opposite demands on the same set of exponents. The exact function’s cusp needs arbitrarily tight Gaussians and its tail needs arbitrarily diffuse ones; in momentum those two requirements swap places, and the same finite set has to serve both. A fit optimised for the energy resolves that competition in the way the energy prefers, which is by describing the region that dominates T\langle T \rangle and V\langle V \rangle — and that is not the region the profile’s own shape is decided in.

The case in the textbooks

Three Gaussians is not an arbitrary choice. STO-3G — three Gaussians fitted to one Slater function — is the basis a generation of chemists met first, is still what a quick calculation defaults to in several programs, and is the entry in every basis-set table that everybody has seen.

Its energy for hydrogen is 0.6 per cent from exact, which is respectable for three functions and is how it is usually reported. Its profile at zero momentum is 2.9 per cent low, its worst relative error inside q<4q < 4 is 16.7 per cent, and the error changes sign twice on the way.

Neither number is a criticism of the basis: three Gaussians is three Gaussians. What they show together is that the single number by which the basis is universally described understates its error in the thing it is being used to compute, by about a factor of five — and that the factor is not a property of STO-3G but grows for every better basis in the same family.

Where in momentum the fit goes wrong

An average error hides the shape of the disagreement, and the shape is the interesting part.

Where in momentum the fit is wrong. The Compton profile of each fitted basis minus the exact one, as a fraction of the exact one, across momentum. The curves cross zero several times rather than approaching it from one side: a fit is a compromise, and a compromise is right in some places by being wrong in others.
Fig. 7 The relative difference between each fitted profile and the exact one, across momentum. The curves cross zero several times rather than approaching it from one side, because a variational fit is a compromise: it buys agreement in the region that dominates the energy by paying for it elsewhere.

The three-Gaussian curve is 3 per cent low at the origin, crosses zero near q=1.6q = 1.6, is 5 per cent high around q=2q = 2, and reaches 17 per cent at its worst inside this range. The energy sees a single number and each of these oscillations very nearly cancels in it, which is what “the energy hides it” means arithmetically rather than as a slogan.

The origin is the region worth watching. J(0)J(0) is the total density of electrons with zero momentum along the scattering direction, and it is dominated by the outer part of the wavefunction — the region a Gaussian gets wrong by dying too fast. High momentum, on the other side, is the region near the nucleus where the exact function has its cusp.

The wrong shape, fitted as well as it can be. The exact hydrogen 1s orbital and the best sums of one, two, three and six Gaussians, each with its exponents optimised for the energy. Three of them already reproduce the exact function to 100.00 per cent by overlap, which is why the method works at all — and the two places it goes wrong, at the nucleus and far out, are exactly where the other faces of this figure look.
Fig. 8 The same fit in position space, where the two failures are the two ends of the curve. The cusp at the origin cannot be reproduced by any sum of Gaussians, because every Gaussian has zero slope there; the tail cannot either, because a Gaussian dies as exp(−ar²) and the exact function dies as exp(−Zr).

The momentum picture and the position picture are the same information transformed, so it is worth saying which way round it goes: the near-nucleus region in position is the far tail in momentum, and the far tail in position is the near-origin region in momentum. So the two ends a Gaussian gets wrong appear at both ends of the profile — and the origin of the profile, which is the easiest part to measure, is set by the part of the wavefunction furthest from the nucleus.

The measurement this makes possible

A basis set’s error is usually estimated by comparing it against a bigger calculation, which is a comparison of two models. The profile makes a different check available: compare against an experiment.

The three quantities to compare are of quite different kinds and it is worth separating them.

The energy is not measurable for a basis set at all. What is measurable is a molecule’s total energy, and the basis-set error in it is entangled with correlation, relativity and the Born–Oppenheimer separation, each of which is larger than the basis error for a good basis.

The profile is measurable directly and at the per cent level. For hydrogen, and for helium and the light molecules, that is a check that a basis passes or fails on its own.

The tail is the quantity most sensitive to what a contraction freezes, and it is the region a profile measures best. A contraction is a decision made once showed that freeing the outermost coefficient recovers most of what contraction costs in energy; the profile at low momentum is the observable that decision moves most.

What a contraction costs in energy is worth putting beside these profiles for one reason: the same freedom that buys three quarters of the energy back is the freedom that sets the shape of the tail. Freezing it is cheap in the currency the basis was fitted in and expensive in the one it was not.

What was computed, and what was checked

Nothing here is quadrature. The transform of a sum of Gaussians is a sum of Gaussians with inverted exponents, the radial integral defining the profile is elementary for a product of two of them, and the exact profile has a closed form — so every curve above is an expression evaluated rather than an integral estimated, and the disagreements are not convergence.

Four checks run while the figures are drawn.

The exact profile integrates to one electron, to a part in a million. That is what makes every error above a fraction of something real rather than of an arbitrary normalisation.

Every fitted profile integrates to one electron too, by construction: the profile is divided by the norm of its own function. So two curves that differ, differ in shape, and cannot differ by an overall factor that would flatter or flatten the comparison.

Both errors fall as the basis grows. A comparison in which one of them did not would be measuring the exponent search rather than the shape, and the search is a descent that could in principle stall.

The second moment of each profile is two thirds of minus its own energy, which is the sum rule above and the sharpest of the four: it connects two calculations done by completely different routes — an eigenvalue from a three-by-three matrix, and an integral over a curve — and requires them to agree to five decimal places.

The last one is also the check that would catch a sign or a factor anywhere in the transform, because a wrong constant would break it while leaving every curve looking entirely plausible.

There is a second cost to adding functions, and it is measured separately: the smallest eigenvalue of the overlap matrix falls as a basis grows. A basis good enough to reproduce a profile has functions close enough together to be nearly linearly dependent, so the two difficulties arrive together.

What this comparison cannot support

One electron, one atom. Everything above is hydrogen’s 1s and a fit to it. A real Compton experiment measures a many-electron system, and the measured profile is a sum over occupied orbitals with correlation corrections that are not small at the level of accuracy being discussed. The isolated statement — a Gaussian fit is worse in momentum than in energy, by a factor that grows — survives that, and the specific numbers do not.

No basis set here is a published one. The exponents are found by descent rather than quoted, which is the choice a Gaussian is the wrong shape makes so that the fit is a computation rather than a table. Published sets are fitted with additional constraints — to several atoms at once, to molecular rather than atomic properties — and would give different numbers with the same shape.

The profile is not the whole momentum density. It is one projection of it, and directional profiles in an oriented crystal measure more. Nothing here needs the difference, because a 1s has nothing to be directional about.

What the field did about it, which is to give up on one basis

If a basis fitted to an energy is worst at everything else, the obvious repair is to fit it to the thing being computed instead. That is what happened, and the shape of the result is the strongest evidence for this essay’s claim — because if the energy were a decent proxy, one family of basis sets would have sufficed, and there are several.

The families are built to converge different quantities and they differ in exactly the places the quantities are sensitive to. One is designed so that the correlation energy converges smoothly with size, which requires functions of high angular momentum in a particular ratio, because that is what the correlation cusp between two electrons needs. Another is designed for density-functional energies and converges differently, because the quantity it is chasing has no such cusp in it.

Then come the property families, and they are the interesting ones. A basis for magnetic shielding needs its inner functions loosened. A basis for nuclear spin–spin coupling needs tight s functions added, which is the repair the cusp condition alone identifies — the coupling is read off the density at the nucleus, and no energy-driven optimisation will ever put functions there because they buy no energy. A basis for polarisabilities and for anything involving a weak long-range interaction needs diffuse functions, which are the opposite end and are also energetically nearly free.

So the practice has three shapes of correction — tighter, looser, and more angular — and which one a calculation needs is decided by which region of space the property lives in rather than by how accurate the calculation is supposed to be. A very expensive basis of the wrong family is worse than a cheap one of the right family, and no amount of watching the energy converge will indicate which is which.

That is the honest summary of the position this measurement puts a reader in. The energy is not a bad number; it is the wrong number to converge on unless the energy is what is wanted. The field discovered that by measuring something a basis had not been fitted to, and its response was not to fit better but to stop believing that one basis could serve.

Who measured this, and when

Compton profiles became a practical probe of wavefunctions in the 1960s and 1970s, when synchrotron sources made the measurement fast enough to compare against calculations rather than the other way round. The literature of that period contains exactly the argument above, made as a complaint: basis sets tuned on energies gave profiles that were visibly wrong at low momentum, and the fix was to include the momentum data in the fitting.

That is where the modern practice of validating a basis on properties rather than on energy comes from, and it is worth noticing that it took a measurement to produce it. No amount of comparing calculations with better calculations reveals a systematic error that every calculation in the family shares.

Still open: the Compton profile of a bond

The obvious open question is a molecule. A Compton profile of a bonded pair is not the sum of two atomic profiles — the difference is the profile of the bond, and it is the quantity in this whole subject that comes closest to a direct measurement of what bonding does to the electrons.

Computing it needs the two-centre momentum density, which is available here in principle: the transform of a function on a displaced centre is the transform of the function times a phase factor, so the cross term is an oscillating integral rather than a new kind of object. What it would need is a two-centre wavefunction good enough to be worth transforming, and the ones used here are hydrogenic and one-electron — which is the standing limit and the reason the calculation is described rather than done.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBasisClosed formConvergenceExpectation valueGaussianModel limitMomentumNormalisationObservableVariationalWavefunction