Bonding models

A method that is not additive

Put a molecule next to a copy of itself, far enough away that they do not interact, and the exact energy doubles exactly. A truncated calculation does not — because the truncation forbids both halves being excited at once, which is something the pair can do and neither half can.

Worth reading first: The smallest many-electron calculation · A better energy is not a better answer.

A better energy is not a better answer establishes that the variational principle makes the energy a one-way test — lower is closer — and that this makes the energy the least sensitive thing about a wavefunction. This essay is about a defect in a class of methods that the one-way test cannot see at all, because it shows up only when the same method is applied twice.

The defect has a name and it is a strange one: size consistency. A method is size consistent when the energy it gives for two non-interacting systems equals the sum of the energies it gives for each alone. Written out, that sounds like a property no reasonable method could lack.

Truncated configuration interaction lacks it, and this essay is the smallest demonstration of why. The model is the one the smallest many-electron calculation built and two pictures, one plane used to measure the ionic weight that this truncation is about to forbid.

The setup: two molecules that ignore each other

The system is Hubbard dimers with no bonds between them at all. Two sites per dimer, one electron of each spin per dimer, a cost UU for double occupancy, and nothing whatever connecting one dimer to the next.

Nothing physical connects them, so the exact answer must be additive, and the first check is that it is:

units exact energy one dimer × n
1 −0.828427 −0.828427
2 −1.656854 −1.656854
3 −2.485281 −2.485281

To nine decimal places. That is not a surprise and it is not a formality either: it is the check that makes everything below a statement about the truncation rather than about the model.

The truncation, and what it is the analogue of

The exact calculation keeps every configuration. A truncated one keeps some of them, and the usual rule is to keep configurations that differ from a reference by at most a stated number of excitations — singles and doubles, which is the workhorse of the subject.

The Hubbard model’s own form of that rule is to bound the number of doubly occupied sites. A site holding two electrons is a charge fluctuation away from the neutral arrangement, so allowing at most kk of them is allowing kk-fold excitations of that kind. The analogy is not exact and it is close enough to make the point, because the structure of the failure is identical.

Truncate at one doubly occupied site.

For a single dimer this throws nothing away. Two electrons cannot doubly occupy two sites, so the truncated space is the full space and the truncated energy is exact: 0.828427-0.828427, error zero.

That is the situation in which methods are usually tested — one molecule, and a truncation whose error is small or, as here, absent.

The same dimer, one, twice and three times over. Independent Hubbard dimers with nothing between them, solved exactly and solved in a space with the configurations that make more than one of them ionic thrown away. The exact energy is exactly additive; the truncated one is exact for a single dimer, because there is nothing there to throw away, and falls behind by 0.193 for two and 0.485 for three. The error per dimer grows, which is what makes a method size-inconsistent rather than merely approximate.
Fig. 1 The same dimer, once, twice and three times over. The exact energies are exactly additive. The truncated ones are not: exact for one unit, 0.19 short for two, 0.49 short for three — and the error per unit grows rather than staying fixed.

What the truncation forbids

For two dimers, the truncated space keeps 30 of the 36 configurations. The six it removes are exactly the ones in which both dimers are ionic at once.

That is the whole mechanism. Each dimer, left alone, spends some of its time with both electrons on one site — at U=4tU = 4t, about fifteen per cent of it, which is the number two pictures, one plane computes as the ionic weight. Two independent dimers therefore spend about two per cent of their time with both of them in that state, and a truncation at one doubly occupied site says that configuration does not exist.

The consequence is an energy 0.19280.1928 too high for two dimers and 0.48530.4853 too high for three. Per dimer: 00, 0.09640.0964, 0.16180.1618. Growing.

The direction is fixed by the variational principle — throwing configurations away can only raise the energy — and the growth is the thing that matters. An error that stayed at a fixed value per unit would be a systematic offset, and offsets cancel in the differences chemists actually use. An error that grows per unit does not cancel in anything.

Why this makes a method useless for large molecules

Consider computing the energy of a reaction in which a large inert fragment sits on both sides of the equation. With a size-consistent method the fragment’s contribution cancels exactly. With a size-inconsistent one it does not, because the fragment’s energy computed as part of the reactant and as part of the product both suffer errors that depend on the whole system’s size — and the two are not the same.

Worse, the size-inconsistency error grows without limit while the chemistry does not. Add enough non-interacting units and the truncated method’s total error exceeds any bond energy the calculation was for, which is the sense in which the method fails at scale rather than being merely approximate. That is a different failure from the one the lattice sum a finite calculation cannot do records — there a sum genuinely does not converge, here a perfectly convergent method converges to the wrong thing at a rate that depends on how much of the system is being ignored.

This is the reason coupled-cluster methods displaced truncated configuration interaction for anything larger than a few atoms. A coupled-cluster wavefunction is written as an exponential of excitation operators — a change of parameterisation rather than of basis, which is the same kind of move as hybrids are a basis and with a consequence that is not cosmetic — and the exponential’s expansion generates the “both units excited at once” configurations automatically from the “one unit excited” ones. It is size consistent by construction, and it is size consistent for a reason that this model makes visible: the missing configurations are products of the ones that are kept, and an exponential is what turns a sum into products.

The refusal that makes the claim a claim

There is an obvious way to be fooled here: a truncation whose error happened to vanish for small systems and grow for larger ones might be describing the model rather than the method.

So the same computation is run with the repulsion switched off, where the exact ground state is most ionic — a single determinant with the ionic configurations at their maximum weight. The truncation should be at its worst there, and it is: the error is 1.17161.1716 for two dimers and 2.53592.5359 for three, against 0.19280.1928 and 0.48530.4853 at U=4tU = 4t.

That is the right direction and it is worth saying why it is the interesting one. Setting U=0U = 0 makes the exact problem trivial — the ground state is a single determinant, which is the easiest case there is — and the truncation is at its worst there anyway, because the truncation is defined in a basis of site configurations rather than of molecular orbitals. A statement about a method has to survive the case where the physics is easy.

How often two electrons are in the same place. Double occupancy per site in the exact ground state of a chain of 2, against the repulsion. It starts at the independent-electron quarter and falls to nearly nothing. The second curve is what the state actually pays in repulsion, which peaks and then falls because the electrons have stopped meeting.
Fig. 2 The quantity being thrown away, for one dimer: how much of the time both electrons sit on the same site, against the repulsion. At U = 0 it is a half and at U = 8t it is a twentieth. Two dimers’ worth of that, squared, is what the truncation forbids.

The fraction recovered is not a constant either

A percentage is the way approximate methods are usually reported — “recovers 95 per cent of the correlation energy” — and the numbers above say that the percentage is not a property of the method.

At U=4tU = 4t the truncation misses 0.190.19 for two dimers, out of a correlation energy that is itself a function of UU. At U=1U = 1 it misses 0.750.75; at U=2U = 2, 0.470.47; at U=8U = 8, 0.0450.045. The error falls with UU because there is less ionic weight to forbid, and it falls at a different rate from the correlation energy itself, so the ratio moves.

So there is no single figure of merit. What the method’s error depends on is how much charge fluctuation the exact state has and how many independent places it can have it in — two quantities that vary from molecule to molecule and, worse, from geometry to geometry within one molecule.

How often two electrons are in the same place. Double occupancy per site in the exact ground state of a chain of 4, against the repulsion. It starts at the independent-electron quarter and falls to nearly nothing. The second curve is what the state actually pays in repulsion, which peaks and then falls because the electrons have stopped meeting.
Fig. 3 The same quantity on four sites rather than two, which is where the replicated units start interacting with the truncation rather than only with each other. The curve has the same shape and a different scale; what the truncation removes are the configurations in which two units are excited at once, and there are none of those until there are two units.

The variational principle certifies nothing about this

It is worth being explicit about why a method with this defect can pass every check a variational calculation offers.

The variational principle says the computed energy is an upper bound. Every number in the table above satisfies it: 1.4641-1.4641 is above 1.6569-1.6569, and 2.0000-2.0000 is above 2.4853-2.4853. Nothing is violated, no warning is produced, and the calculation converges beautifully.

The principle also says lower is better, and by that test the truncated two-dimer calculation is worse than the truncated one-dimer calculation scaled up — but nobody does that comparison, because there is no reason to. The comparison that would expose the defect is between one calculation and two, and a single run cannot make it.

This is a specific instance of the general shape a better energy is not a better answer draws: the energy is the quantity a variational method optimises and is therefore the quantity least sensitive to what the wavefunction is doing wrong. Here it is worse than insensitive — it is systematically reassuring, since the errors are all in the direction the principle guarantees.

The check that does expose it costs one extra calculation: run the fragment alone, multiply, compare. That is a two-route check of exactly the kind this site runs on, and it is not part of any standard output.

Where the electrons are, without subtracting anything. The opposite-spin pair distribution of a half-filled ring of 6 at six repulsions, by separation, each divided by what uncorrelated electrons of the same density would give. At no repulsion it is one everywhere; at a repulsion of 16 the chance of finding two electrons on one site is 0.0430 of that, and what is missing has turned up next door. Nothing here is a difference between two calculations.
Fig. 4 Where the electrons actually are, with no subtraction anywhere in it. This is the quantity a truncation gets wrong first and an energy reports last: an energy is second order in a wavefunction’s error where a property is first order, so a description that is visibly wrong here can still produce an energy that looks nearly right — which is the mechanism by which a size-inconsistent method passes for healthy.

The same failure with more units per system

Two dimers is the smallest case that shows anything, and it is worth seeing that the mechanism is about independent units rather than about the number two.

A chain of four sites is not two independent dimers — its sites are all connected — and its double occupancy behaves quite differently: the electrons can spread, the ionic weight per site is lower, and there is no sense in which the system decomposes. That is the case where a truncation by charge fluctuations is a reasonable approximation rather than a structural error, because there is only one system for the fluctuations to be in.

How often two electrons are in the same place. Double occupancy per site in the exact ground state of a chain of 4, against the repulsion. It starts at the independent-electron quarter and falls to nearly nothing. The second curve is what the state actually pays in repulsion, which peaks and then falls because the electrons have stopped meeting.
Fig. 5 A connected chain of four, for contrast with the disconnected pair. Its double occupancy per site falls smoothly with the repulsion and there is nothing here that decomposes into independent pieces — which is why the size-consistency question does not arise for a single molecule, and why testing on one is not a test.

The failure needs separability: a system that genuinely splits into parts, so that the exact wavefunction is a product and the truncated one cannot be. Real chemistry supplies separability constantly — a solvent molecule far from a solute, a substituent far from a reaction centre, two ends of a long chain — which is why the defect matters in practice rather than only in a constructed example.

Where the model stops

The truncation is not literally CISD. Configuration interaction with singles and doubles is defined by excitation level from a reference determinant in a molecular-orbital basis; the truncation here is by the number of doubly occupied sites, which is the natural analogue in this model and is not identical to it. What transfers is the structure of the failure — a space that contains one unit’s excitations and not two units’ at once — and that is the structure that produces size inconsistency in the real method.

The dimers are identical. Non-interacting fragments of different kinds would show the same failure and the arithmetic would be less clean; nothing about the argument needs them to be alike, and the check that the exact energies add is what rules out the alternative explanation.

On-site repulsion only, and no basis. Everything the smallest many-electron calculation says about the limits of this model applies here: one orbital per site, no long-range repulsion, and no relation to any real distance.

Three units is not the thermodynamic limit. The trend in the error per unit is established over three points, which is enough to show that it grows and not enough to establish a law. The exact solver here is limited to six sites deliberately, which is three dimers.

Size consistency is not size extensivity. They are usually run together and they are different: extensivity is about the energy scaling linearly with the number of interacting particles, consistency is about non-interacting fragments adding. A method can have one and not the other, and the model here tests only the second.

Two properties, and this essay measured both of them

The demonstration above ran two different tests and gave them one name. They are usually distinguished in the literature, the distinction is worth having, and the essay’s own numbers are the cleanest illustration of it.

Size consistency is the property tested by putting two units infinitely far apart and asking whether the energy of the pair equals the sum of the energies of the parts. It is a statement about a limit — a dissociation limit — and it is what decides whether a method can describe a molecule coming apart, or a reaction between two fragments, without the answer depending on whether the fragments were computed together or separately.

Size extensivity is the property tested by making the system bigger and asking whether the energy grows in proportion. It is a statement about scaling, it does not involve separating anything, and it is what decides whether a method’s error stays a fixed amount per atom as a molecule grows.

The two are not the same requirement. Extensivity is about the leading behaviour in the number of units; consistency is about an exact equality at one geometry. A method can satisfy one and fail the other, and the standard perturbation treatments are the case in point: they are extensive by a theorem about which diagrams survive, and they still misbehave badly at a stretched bond for reasons that have nothing to do with size.

Truncated configuration interaction fails both, and this essay measures both failures separately without saying so. That the exact energy of two dimers doubles to nine decimal places while the truncation is 0.19 out is the consistency test. That the error per unit grows as units are added, rather than staying put, is the extensivity test — and it is the more damaging of the two, because it says the method’s accuracy per atom degrades with molecular size, which is exactly backwards for a tool intended to be applied to larger systems than the ones it was benchmarked on.

Naming them separately also says which repair is needed for which purpose. A method for dissociation curves must be consistent; a method for large molecules must be extensive; and the exponential parameterisation supplies both at once, which is why it displaced the truncated sum rather than competing with it.

What the exponential does, in this model’s terms

The repair deserves a sentence of arithmetic rather than a claim, because the model makes it visible.

Write the truncated wavefunction as the reference plus one unit’s excitations: (1+T)Φ(1 + T)|\Phi\rangle, where TT generates the configurations with one ionic site. For two independent dimers the exact wavefunction is a product of each dimer’s, and a product contains a term with both excited — which (1+T)(1 + T) does not, since TT acting once cannot produce two excitations.

Now write it as eTΦ=(1+T+12T2+)Φe^{T}|\Phi\rangle = (1 + T + \tfrac12 T^2 + \dots)|\Phi\rangle. The T2T^2 term is exactly the missing configuration, with exactly the coefficient the product requires — because the exponential of a sum is the product of the exponentials, and two non-interacting units contribute independent terms to TT.

So size consistency is not an extra property bolted onto coupled-cluster theory; it is what an exponential parameterisation is. The same statement in the model above: the truncated space keeps 30 of 36 configurations for two dimers, the exponential generates the other 6 from the ones it kept, and it generates them with coefficients that are products rather than free parameters.

That last clause is what makes it cheap. The missing configurations are recovered without being variationally optimised, so the calculation stays the size of the smaller space and gives an answer that behaves like the larger one.

Where size-consistency sits among the correlation arguments

The first essays on correlation introduce a model with repulsion in it at all, and separate the correlation that comes from exchange — which a single determinant already has — from the correlation that comes from repulsion, which it does not.

This one is about a defect that appears only when a method is used more than once. The exact solution of two non-interacting units is additive to nine decimal places; a truncation that is exact for one unit is 0.190.19 wrong for two and 0.490.49 wrong for three; and the error per unit grows rather than staying put. None of that is visible in a calculation on a single molecule, which is where a method’s error is normally quoted, and none of it is caught by the variational principle, which certifies only that the answer is too high.

Still open is the repair — what an exponential ansatz does that a truncated sum cannot, which is the one structural idea separating the methods that scale from the methods that do not.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationConfiguration interactionConvergenceDouble occupancyElectron correlationExact diagonalisationHubbard modelMany-electron wavefunctionsModel limitOn-site repulsion