Orbitals

A contraction is a decision made once

Every published basis set freezes its primitive functions into fixed combinations, on an isolated atom, before any molecule is in sight. Freeing one coefficient recovers three quarters of what that costs — and freeing it at the other end of the basis recovers one per cent.

Worth reading first: A Gaussian is the wrong shape · Hybrids are a basis.

The word contraction means two different things in the description of an orbital, and both of them appear in this collection. One is what a molecule does: an orbital shrinks when a bond forms, because the electron is now attracted by two nuclei rather than one, and the exponent the hydrogen molecule ion chooses comes out at 1.238 rather than the free atom’s 1. The other is a piece of bookkeeping in every electronic-structure program written since 1969: a contracted basis set is a collection of primitive Gaussian functions which have been frozen into fixed linear combinations, once, and will only ever appear in those combinations again.

The two meet, and the meeting is the whole of what follows. A contraction fixed on the free atom cannot represent the contraction the molecule wants, because the ratio it froze was the free atom’s answer.

What is actually frozen

A primitive Gaussian is gi(r)=eαir2g_i(r) = e^{-\alpha_i r^2}, and a sum of them is the wrong shape at both ends — no cusp at the nucleus, too fast a decay far out — for reasons that no number of them repairs. What they buy is that every integral has a closed form, and that is enough to have carried the whole subject for sixty years.

Optimising six of them for a hydrogen atom gives exponents spanning three orders of magnitude, from 0.0900 to 82.9, and a set of coefficients that ties them together.

Six primitives and the one function they are frozen into. The six optimised Gaussians, each scaled by its own contraction coefficient, and their sum. The exponents run from 0.090 to 82.9 — a range of 922 — so the tightest describes the cusp region and the most diffuse the tail. A contraction says that these six will only ever appear in this ratio, which is exactly right for the atom they were fitted to and is a decision about every other atom as well.
Fig. 1 The six optimised primitives, each drawn at its own contraction coefficient, and the function they add up to. The tightest describes the region near the nucleus where the exact orbital has its corner; the most diffuse carries the tail. A contraction is the statement that these six will only ever appear in this ratio.

A calculation using them uncontracted has six linear coefficients to choose. A calculation using them contracted into a single function has one, and the other five have already been decided. The primitives are the same primitives; the integrals are the same integrals — a contracted matrix element is a double sum over the primitives of the two shells, so contraction saves nothing at all at the integral stage. What it saves is everything after it: the matrix to be diagonalised, the number of two-electron integrals to be transformed, the size of the correlation problem. That is why it exists.

The wrong shape, fitted as well as it can be. The exact hydrogen 1s orbital and the best sums of one, two, three and six Gaussians, each with its exponents optimised for the energy. Three of them already reproduce the exact function to 100.00 per cent by overlap, which is why the method works at all — and the two places it goes wrong, at the nucleus and far out, are exactly where the other faces of this figure look.
Fig. 2 The exact hydrogen 1s and the best sums of one, three and six Gaussians. Six of them reproduce the function to better than 99.99 per cent by overlap over the range where the density is, which is why a minimal contracted set is a usable description of an atom — and it is an atom that is being described.

What a contraction actually buys

It is worth being exact about the saving, because the usual summary — a contracted basis is smaller — gets it wrong in a way that matters for reading the cost.

The integrals are unchanged. A matrix element between two contracted shells is

AO^B=iAjBcicjgiO^gj,\langle A | \hat{O} | B \rangle = \sum_{i \in A} \sum_{j \in B} c_i c_j \langle g_i | \hat{O} | g_j \rangle,

so every primitive integral still has to be computed, and there are exactly as many of them as there would be with no contraction at all. What falls is the number of basis functions the rest of the calculation sees. The matrix to diagonalise is smaller; the transformation of two-electron integrals from the primitive index to the orbital index — which scales as the fifth power of the number of functions — is smaller by that power; and any correlated method built on top of the orbitals is smaller again.

For a molecule of any size that is the entire cost of the calculation, and the contraction is buying it with a variational freedom whose value is what this essay measures. The trade is real. The question is only whether anybody knows what they paid.

The atom it was fitted to is the one place it costs nothing

At the atom the contraction was built for, freezing the ratio costs nothing: the excess energy is 6×1015-6 \times 10^{-15} hartree, which is the arithmetic and not a quantity. That has to be so, and it is worth saying why rather than merely observing it. The contraction coefficients are the answer to the atom’s own variational problem. Handing a calculation an answer it was going to arrive at anyway takes nothing away from it.

This is the reason the defect is invisible to anybody testing a basis set on atoms, and it is not a small class of people. A basis set is published with atomic total energies, and those energies are exactly the quantity a contraction is free in.

The other end of the knob

A one-electron problem has one knob for a different environment, and it is the nuclear charge. That is a cruder change than forming a bond, and it is also the same change measured in the other sense of the word: an exponent carries a length, so an orbital that has contracted by a factor is an atomic orbital at an effective charge larger by that factor. The hydrogen molecule ion settles at ζ=1.238\zeta = 1.238 at its own bond length, so 1.238 is a computed estimate of what bonding does to a basis, computed rather than assumed.

The exponent the molecule chooses, and what it buys. The 1s exponent that minimises the energy of a one-electron diatomic, against the separation of the nuclei, with the binding curves at that exponent and at the free atom's. Held at ζ = 1 the bond comes out at 2.49 bohr and binds 0.0648 hartree; with the exponent free it comes out at 2.00 bohr at ζ = 1.238 and binds 0.0865. The exact answer for this molecule is 2.00 bohr and 0.1026.
Fig. 3 The exponent the hydrogen molecule ion chooses at each separation, from a one-electron calculation with every integral in closed form. It rises to 1.39 at short range, passes through one near five bohr and returns to the free atom’s value at long range. The value at the equilibrium separation, 1.238, is the charge the basis test below is run at.

Every basis in the comparison holds exactly the same six primitives. The only difference between them is how many of the linear coefficients the calculation may choose for itself, and the whole of the cost reported is the cost of having chosen them somewhere else.

The cost of a ratio decided somewhere else. Four bases containing exactly the same six primitives, differing only in how many of the linear coefficients the calculation may choose. The horizontal axis is the effective nuclear charge, which is this one-electron problem's only knob for a different environment; the contraction was fitted at one. At 1.238 — the exponent H₂⁺ chooses when a bond forms — the fully contracted basis is 0.0283 hartree above what the same six functions could give, and one freed coefficient removes most of it.
Fig. 4 The energy penalty against the effective nuclear charge, on a logarithmic scale. At the charge the contraction was fitted to it is zero. At 1.238 the fully contracted basis is 0.0283 hartree above what the same six functions could give — seventy-four kilojoules a mole, which is a bond energy — and it climbs from there. One freed coefficient removes most of it; two remove nearly all.

Seventy-four kilojoules a mole is the number worth carrying away. It is not an error in a small correction; it is the size of the quantity a chemist wants the calculation for. And it arises on an atom that has not changed, in a basis containing exactly the same functions, purely because a linear coefficient was fixed on somebody else’s problem.

Which end to unfreeze, and a reason that was wrong twice

A split-valence basis set leaves the outermost primitive — or two — free, and contracts the rest. The convention is universal and the justification usually given is chemical: the valence region is where bonding happens, so the valence functions need freedom.

Two natural expectations both fail. One is that freeing a single primitive recovers more than half of what the contraction lost; it recovers 29.5 per cent. The other is that the diffuse end must be the wrong one to free — a larger nuclear charge pulls the orbital inward, so surely the tight coefficients are the ones needing to move — and freeing the tight end recovers 1.1 per cent against the diffuse end’s 29.5.

The convention is right and both readings of it were wrong.

Which end of the basis to unfreeze. At an effective nuclear charge of 1.238 — the contraction still frozen at the value fitted for a charge of one — the cost of the freezing and what each way of relaxing it recovers. Freeing one coefficient at the DIFFUSE end recovers 78 per cent of the loss; freeing one at the tight end recovers 1 per cent. A split-valence basis frees the valence end, and this is a one-electron atom's own reason for the convention.
Fig. 5 At the effective charge a bond gives an orbital: what the frozen ratio costs, and what each way of relaxing it recovers. One freed coefficient at the diffuse end recovers 78 per cent; one at the tight end recovers 0.9 per cent. Two at the diffuse end recover 98.8 per cent, which is why a basis set has more than one split.

What a frozen ratio gets wrong is how much tail it carries. A contracted function fitted at one charge and used at a larger one is too diffuse, and one free coefficient on the most diffuse primitive can subtract that excess directly — the five primitives held together are then already close to the right shape. Freeing the tight end instead leaves five diffuse primitives locked in one badly wrong combination, and no amount of the sixth repairs it.

So the convention frees the end that can be corrected, not the end that has changed. Those are different criteria, and only one of them is the reason. The chemical story happens to name the same functions.

Which end of the basis to unfreeze. At an effective nuclear charge of 3 — the contraction still frozen at the value fitted for a charge of one — the cost of the freezing and what each way of relaxing it recovers. Freeing one coefficient at the DIFFUSE end recovers 29 per cent of the loss; freeing one at the tight end recovers 1 per cent. A split-valence basis frees the valence end, and this is a one-electron atom's own reason for the convention.
Fig. 6 The same comparison at an effective charge of three, far outside anything a bond produces. The frozen ratio now costs 1.99 hartree, one split recovers 29.5 per cent of it and two recover 71 per cent — so the fraction one split recovers falls the further the environment is from the atom the ratio was frozen at. A basis set has a range of validity, and this is what it looks like.

The comparison a chemist actually makes

An energy on its own is not a chemical statement. What matters is whether the error cancels between two situations being compared, and the shape of the curve above says that it does not.

A frozen ratio costs nothing at charge one, five thousandths of a hartree at 1.1, and twenty-eight thousandths at 1.238. So a calculation comparing a free atom with the same atom in a bond is comparing one number that is exact for its own problem with another that is seventy-four kilojoules a mole too high. The error does not cancel; it is entirely on one side. That is the worst arrangement a systematic error can have, and it is the arrangement a contracted basis produces for exactly the quantity — a bond energy — that the calculation was run for.

The repair the field settled on is not a better contraction but more freedom, and the numbers above say how much freedom buys what. It is the same shape of answer as the convergence of the energy with basis size: each added function lowers the energy from above, none overshoots, and the returns fall off. Freedom in the linear coefficients behaves the same way, because it is the same variational principle acting on the same space.

The cost of a ratio decided somewhere else. Four bases containing exactly the same six primitives, differing only in how many of the linear coefficients the calculation may choose. The horizontal axis is the effective nuclear charge, which is this one-electron problem's only knob for a different environment; the contraction was fitted at one. At 1.238 — the exponent H₂⁺ chooses when a bond forms — the fully contracted basis is 0.0283 hartree above what the same six functions could give, and one freed coefficient removes most of it.
Fig. 7 The same cost measured at three bohr instead of at the bond length, which is the check that it is a cost of the contraction rather than of the geometry. It is larger there, not smaller: the further apart the atoms are, the more the frozen ratio is the wrong ratio, because it was fixed at a separation the atoms are no longer at. A curve of atomic energies against the number of primitives — the curve a basis set is published against — cannot see any of this, because every point on it is an atom on its own.

The ratio is the same for every atom

A result worth stating because it says what the loss is not. The one-electron problem scales exactly: solving it at charge ZZ gives the charge-one answer with every exponent multiplied by Z2Z^2 and every energy by Z2Z^2. So the optimised contraction coefficients at charge three are the optimised coefficients at charge one, to two parts in a hundred. Only the exponents move.

The loss measured above is therefore not the loss of a ratio that should have been different. It is the loss of exponents that should have been different, with a ratio holding them together. A contraction cannot scale its own exponents, and that is the whole of its rigidity.

The check that the fitted ratio is nevertheless doing work is a refusal: two other ratios over the same six primitives — a flat one, and the fitted one reversed — are both worse at the atom the fit was made for, by amounts far outside any numerical tolerance. So the coefficients carry the fit, and the fit is worth something; it is just worth it in one place.

Where the numbers came from

Nothing here is quoted from a table of basis sets, and that is deliberate: quoting the exponents would put the one thing this essay is about — that they were fitted, once, by somebody, to one problem — behind a reference. So the six exponents are found by a deterministic descent on their logarithms, the coefficients fall out of the generalised eigenvalue problem HC=SCEHC = SCE, and the non-orthogonality of Gaussians on one centre is removed by symmetric orthogonalisation — the same S1/2S^{-1/2} that the localisation transformation uses on four bond orbitals, where what is being removed is the overlap between hybrids rather than between primitives.

The eigenvalues then come from a Jacobi eigensolver checked against closed forms. Two internal checks make the comparison meaningful rather than merely consistent: an uncontracted basis solved as one-primitive shells reproduces the primitive solver’s own answer to twelve decimals, so the two routes are solving the same problem; and no contracted basis anywhere in the scan beats the uncontracted one it was built from, which is a theorem rather than a measurement — its functions span a subspace of theirs — and is therefore the check that would catch an error in the arithmetic rather than in the physics.

An error that varies along a bond is a force

Seventy-four kilojoules a mole is quoted above as a cost, and a cost is a number a chemist immediately wants to see cancel. Most errors of this kind do: a defect present in the same amount on both sides of a comparison subtracts out, which is the whole of why absolute energies are never quoted and differences always are.

This one does not cancel, and the reason is contained in how it was measured. The contraction was frozen at the free atom’s nuclear charge and evaluated at the charge a bond gives an orbital. At the first it costs nothing whatever — it is exact there, by construction. At the second it costs seventy-four. So the error is zero at large separation and of chemical size at bonding separation, which is to say it is a function of the internuclear distance rather than a constant.

An energy error that varies along a coordinate is not an offset. It is a contribution to the shape of the curve, and its derivative is a contribution to the force. This one rises as the atoms approach, so it opposes their approach: a contracted basis under-binds at the bond length and does not under-bind at infinity, which pushes the computed minimum outward and flattens the well around it. A calculation in such a basis therefore returns a bond that is slightly too long and a force constant slightly too small, and neither error is visible in any single number the calculation reports.

The direction is worth keeping because it is the direction the errors actually run. Bond lengths from small contracted basis sets are systematically long and their harmonic frequencies systematically low, and the usual explanation offered is basis incompleteness in general. Part of it is specifically this: a ratio decided on an isolated atom, applied to an orbital that is no longer on an isolated atom, and applied more wrongly the closer the atoms come.

It also identifies which comparisons are protected and which are not. Two molecules of similar bonding, computed in the same basis, have similar contraction errors and the difference between them is safe. A single molecule at two geometries does not, because the contraction error is one of the things that changes between them — and a potential energy surface is nothing but one molecule at many geometries.

What this cannot say

Three limits, and they matter for reading the number.

The environment is a nuclear charge, not a molecule. A real basis set is used in a field with more than one centre in it, and the redistribution a bond causes is not a uniform contraction. Modelling it as one is a caricature that gets the sign and the rough size right and nothing else.

There is one electron. No repulsion, no self-consistency, no correlation. A self-consistent mean field exists and is not used here, because bringing it in would replace a number that can be checked against an exact answer with one that cannot.

Only s functions. A real molecular basis carries p and d primitives whose contraction raises the same question with more indices, and the polarisation functions that make a basis usable for anything anisotropic are usually left uncontracted precisely because there is no atomic answer to freeze them at.

None of the three changes the shape of the result, which is that a fixed linear coefficient is a decision, that the decision is invisible on the problem it was taken for, and that its cost elsewhere is of chemical size.

Where the argument goes next

The next thing to ask of a basis is what it does to a property rather than to an energy, and the answer is uncomfortable: an energy can be excellent while the function is the wrong shape, and it is properties that depend on the tail — polarisabilities, long-range interactions, anything that has to reach — which converge slowest. A contracted tail is a fixed tail. The measurement above says how much energy that costs; it says nothing about how much shape it costs, and shape is the thing the cusp and the tail were already wrong about before anything was frozen.

The other direction is still open: two electrons, self-consistently, in a Gaussian basis. Every other ingredient is in hand — the primitive integrals in closed form, the symmetric orthogonalisation, the eigensolver, the mean field with a checked sign. What is missing is a two-electron integral over Gaussians on different centres, which is the one ingredient that turns all of this into quantum chemistry rather than a very careful account of one atom.

There is also a question this raises about pictures rather than about energies. A contracted orbital has a fixed shape by construction, so a contour drawn at a stated enclosed fraction is a contour of a function whose outer form was decided on an atom. The picture will be perfectly self-consistent and the fraction it encloses will be exactly what is claimed; what it will not be is a picture of the orbital in the molecule. That is a different kind of error from the ones the isovalue essay is about, and it is invisible in the same way — every number is right, and it is a number about the wrong object.

The claim to hold on to meanwhile is small and firm. A published basis set is a table of numbers, half of which are answers somebody else computed to a question about an isolated atom. Using it is agreeing to those answers. On the atom they are free; at a bond length they are worth seventy-four kilojoules a mole; and the one coefficient it is worth unfreezing is at the end of the basis that can absorb the error, rather than the end where the error came from.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBasisBond lengthConvergenceEffective nuclear chargeEigenvalueModel limitOne-electron modelsOrthogonalityVariational