Orbitals

The correction that gets harder to assemble

Summed across a trimer, pairwise counterpoise corrections come to fifteen per cent more than the trimer's own. Three Gaussians a centre is a small basis, so the natural question was whether the fraction shrinks with a better one or stays put. It does neither. The correction falls by three orders of magnitude and the fraction grows fivefold.

Worth reading first: The assembly that counts one share twice · The basis the other atom lent.

The assembly that counts one share twice put three atoms on a line and asked whether the counterpoise correction for the trimer is the sum of the corrections for its pairs. It is not: the assembly overshoots by 15.7 per cent at a separation of two bohr, and it is wrong in two directions at once — the heavy central fragment over-corrected, the light outer ones under-corrected, with the total taking the centre’s sign because the centre’s correction is thirty times the others’.

Every number there used three Gaussians a centre. That is a small basis, and a small basis is one with a large superposition error to begin with, so the closing question was the obvious one: does the fraction shrink as the basis improves, which would make this an artefact, or stay put, which would make it a feature?

Neither. It grows — and it grows at every separation tested, by factors between four and thirteen, while the quantity it is an error in falls by three orders of magnitude.

The correction collapses and the error made assembling it grows. At a separation of 2 bohr, two quantities against the number of Gaussians a centre. The trimer's own counterpoise correction falls by a factor of 2717 from one function to six — a bigger basis has less to borrow. The fraction by which summing the pairwise corrections overshoots it rises from 5.3 per cent to 38.0. Improving the calculation makes the assembly proportionally worse.
Fig. 1 At a separation of two bohr: the trimer’s own correction against the basis size, and the fraction by which assembling it from pairs overshoots. The two go in opposite directions.

Why a bigger basis has less to borrow

The mechanism behind the first half of the expectation is worth stating, because the second half fails despite it rather than against it.

The basis the other atom lent established what the correction is correcting: a fragment computed inside the whole system’s basis set can use functions centred on its neighbours, and so comes out lower than the same fragment computed alone. The difference is not physics — the neighbour is not there in the isolated calculation — so it is subtracted off.

How much a fragment gains that way depends on how badly its own basis describes it. With one Gaussian a centre the description is poor and a neighbour’s function is a large improvement; with six it is good and the neighbour has little to add. So the correction must fall as the basis grows, and it falls steeply.

Everything about that is right. What does not follow — and what the closing question assumed would follow — is that the error made in estimating the correction falls with it.

Both halves move, and they move apart

The first half of the expectation is right, and it is right emphatically. The counterpoise correction is what one fragment gains by borrowing another’s basis functions, and a fragment with a good basis of its own has little to gain. Across the sweep from one Gaussian a centre to six, the trimer’s correction at a separation of two falls from 3.83×1023.83 \times 10^{-2} hartree to 1.41×1051.41 \times 10^{-5} — a factor of 2,717.

The second half is wrong. Over the same sweep the overshoot goes 5.25, 7.10, 15.69, 23.16, 32.27, 37.98 per cent. The quantity being corrected for is disappearing, and the error made by assembling it pairwise is taking a larger and larger share of what remains.

It grows at every separation, and fastest where it starts smallest. The assembly's overshoot against the basis size at four separations. Every curve rises. The near separations start high and roughly quintuple; the far one starts under one per cent and grows by a factor of 13, so a geometry where the correction was negligible with a small basis is one where the assembly is badly wrong with a large one.
Fig. 2 The overshoot against the basis size at four separations. Every curve rises, and the one that starts smallest rises fastest.

A useful way to hold the two numbers together: at one Gaussian a centre the assembly is wrong by about 2×1032 \times 10^{-3} hartree, and at six by about 5×1065 \times 10^{-6}. In absolute terms the approximation is improving by a factor of four hundred. It is only as a share of what it is approximating that it is getting worse — which is the honest statement, and which of the two matters depends on what the number is being compared against.

It is not a property of one geometry. At 1.6 bohr the overshoot goes from 11.6 to 52.9 per cent, at 2.0 from 5.3 to 38.0, at 2.5 from 2.3 to 20.8, and at 3.0 from 0.7 to 9.5. The growth factor is largest where the starting value is smallest — thirteenfold at the widest separation against fourfold at the narrowest — so the geometries where a small-basis calculation would have said the assembly is fine are exactly the ones where a larger basis says it is not.

What makes the fraction climb

The two movements have a common cause and it is not mysterious. The correction a fragment gains is roughly what its neighbours’ functions add to its own description. The error the pairwise assembly makes is the part of that gain which two neighbours supply together — the share the central fragment would have got from either one alone, counted once for each.

As the basis improves, the part a single distant function can add falls quickly, because the fragment’s own basis already covers what that function was providing. The part that only the two together supply falls more slowly, because it is the residue neither alone accounts for. A ratio whose denominator falls faster than its numerator rises, and that is the whole of the effect.

That reading predicts what the fragment figure shows — the growth belongs to the centre, the only atom with two neighbours — and it predicts the separation trend too: the wider the fragments are apart, the more of the correction is the jointly-supplied residue rather than the singly-supplied bulk, so the fraction starts smaller and climbs faster. Both hold.

Which fragment is doing it

The heavy fragment carries the error, at every basis size. Each of the three fragments' own overshoot against the basis size, at R = 2. The heavy central fragment is over-corrected throughout and its error grows steadily; the two light outer ones stay near zero and sometimes go the other way. The total is the centre's, because the centre's correction is much the largest — so the growth in the previous figure is one fragment's growth, not three.
Fig. 3 Each fragment’s own overshoot across the sweep. The heavy centre carries all of it.

Splitting the total by fragment shows the growth is one fragment’s. The heavy central atom — the one with two neighbours to borrow from — is over-corrected at every basis size and increasingly so. The two light outer atoms sit near zero throughout and cross it in both directions.

That is the same structure the three-atom assembly found, and it survives the sweep unchanged, which is worth stating because it is the part that does not move: the mechanism is not basis-dependent, only its size is. A fragment with two neighbours has its pairwise corrections counted against a basis it has already partly been given, and counting the same share twice is a statement about geometry rather than about Gaussians.

There is a small irregularity worth not smoothing over. At the two wider separations the sequence is not monotone: at 2.5 bohr the overshoot dips from 2.28 per cent at one function to 1.54 at two before climbing, and at 3.0 it falls from 0.74 through 0.67 to 0.58 before rising to 9.54. So the first one or two members of the series do not follow the trend at long range. That is the region where the correction is tiny and the fragments barely overlap, and a one-Gaussian description is so poor that its errors are not yet dominated by the sharing this measures. The claim checked is the one the data supports — larger at six than at one, at every separation — rather than a monotone rise.

The direction that matters for practice

The smaller the correction, the worse the assembly of it. Every basis size at every separation, with the overshoot against the size of the correction being assembled. The relation runs the wrong way for a practitioner: the points on the left are the ones where the correction is small enough to be tempting to approximate, and they are the ones where approximating it pairwise is worst. Marks are coloured by separation.
Fig. 4 Every basis size at every separation, plotted against the size of the correction being assembled. The relation runs the wrong way.

Put the two quantities against each other and the practical shape appears. The points at the left of that figure are the ones where the correction is small — a few times 10510^{-5} hartree, a hundredth of a kilocalorie, the regime where a practitioner would reasonably decide the correction is a detail and any sensible estimate of it will do. Those are the points where the pairwise estimate is worst, by a factor of five over the points on the right.

A Gaussian is the wrong shape is where this collection starts, and it is worth recalling here: the whole reason a correction like this exists is that the functions are the wrong shape and more of them are needed to compensate. Adding functions fixes the description and does not fix the fact that neighbouring fragments’ functions overlap — so the thing the correction exists for goes away while the thing that makes it hard to assemble does not.

The reasoning that fails here is not careless. It is: this correction is small, so an approximate treatment of it is safe. That is sound whenever the approximation’s error is a fixed fraction of the thing approximated, and it is exactly what does not hold. The error’s fraction and the correction’s size move in opposite directions.

What is not in doubt

Three things survive the sweep unchanged, and listing them is the honest way to bound what has been shown.

The sign is one: the assembly overshoots at every basis size and every separation tested, so the pairwise sum is always an over-correction and never an under-correction. The mechanism is another: the heavy centre is over-corrected and the light ends are not, at every basis size, so the structure identified on the trimer is not a small-basis feature either. And the geometry dependence is the third — the overshoot falls with separation at fixed basis size, as it must, since fragments far enough apart share nothing.

What changed is only the answer to the question that was asked, which was whether the size of the effect goes away with a better description. It does not, and the direction of the failure is the one that costs a practitioner rather than the one that reassures them.

What a converged basis would say

Where the sweep is heading, and why it cannot be read off. The same overshoots against the reciprocal of the basis size, which is where a limit would show. None of the four curves is straightening towards the axis: each is still rising at the largest basis and rising faster in these coordinates than in the last figure. Six functions is not near a converged basis for this quantity, and nothing here says what the converged answer is.
Fig. 5 The same overshoots against the reciprocal basis size, where a limit would show as a flattening towards the left edge. None of them flattens.

The honest limit here is that nothing can yet say where the sweep ends. Plotted against 1/n1/n — the coordinate in which a converged quantity approaches its limit linearly, and where a flattening towards the axis would show a limit being reached — all four curves are still climbing at six functions, and climbing more steeply than they were.

So the sweep establishes a direction and not a destination. Whether the overshoot approaches a finite fraction, or continues to a complete-basis limit where the correction is zero and the fraction is undefined, is not decidable on six functions. The second possibility is not absurd: at a complete basis there is no superposition error, the correction is exactly zero, and a fraction whose numerator and denominator both vanish has no value to approach.

The property that gets worse is the essay this most resembles in shape: there, improving a basis improved the energy and degraded a property computed from it, because the two respond to different parts of the description. Here it is not a property that degrades but an approximation made on top of the calculation, and the reason is the same in outline — the quantity being approximated shrinks faster than the approximation’s error does.

What was computed, and how

Three centres on a line, one electron each, nuclear charges 1, 2 and 1 with the heavy one in the middle. The basis at each centre is nn Gaussians fitted for that centre’s charge, and every energy comes from the same variational solver on the same line — so the comparison across basis sizes is a comparison of one method at six resolutions rather than of six methods.

The whole sweep, both quantities, four separations. For each basis size: the trimer's own counterpoise correction at a separation of two, and the percentage by which the sum of pairwise corrections exceeds it, at each of four separations. Reading down the second column and then across the rest is the finding — three orders of magnitude of improvement in the quantity, and a steadily worse job of assembling it.
Fig. 6 The whole sweep: the correction at a separation of two, and the overshoot at four separations, for each basis size.

The correction for a fragment is its energy alone minus its energy computed in the full set of basis functions with the other nuclei’s charges switched off, which is the standard construction; the pairwise assembly repeats that with one neighbour at a time and adds the results. Nothing is fitted and nothing is cached — the whole sweep is under a second, which is why four separations were run rather than one.

The refusal is a separation where nothing is shared. At twelve bohr the fragments’ basis functions do not reach one another, the correction is numerically zero, and the calculation has to report that rather than a ratio of two vanishing quantities.

Where the model stops

One electron a centre is the model’s largest simplification, and it is the one that matters most here. Real superposition error is a many-electron effect, and a real fragment’s correction depends on how much of the borrowed basis its occupied orbitals can use. Whether the fraction grows with basis size in a many-electron system is not answered by this and could go either way.

The basis series is also one family — even-tempered Gaussians fitted per centre — swept in size alone. A real basis series varies in more than count: polarisation functions, diffuse functions, contraction schemes. A sweep along a real series such as the correlation-consistent family would be the honest test of the same claim, and it is not available here.

And three centres is the smallest system in which the question exists. With four the pairwise assembly has six pairs and four triples to leave out, and a correction that is two functions is about what the correction is made of rather than how many pieces it is cut into.

Two further limits belong to the model rather than to the method. A correction computed at one length established that the correction is a function of geometry and not a constant to be carried between structures, and that stands here — every number above is at a stated separation. And the half that cannot be computed is about what the counterpoise construction leaves undefined, which no amount of basis improvement settles.

The generalisation

Fragment and many-body methods are built on a premise this measures: that a correction computed on pieces adds up to the correction for the whole, well enough. The premise is usually defended by size — the neglected terms are small — and the defence is usually checked in one basis.

What this shows is that the two variables are not independent in the convenient direction. Improving the basis improves the calculation and degrades the approximation made on top of it, so a fragment method validated in a modest basis has been validated in the regime where it works best. That is the opposite of the usual assumption, which is that a modest basis is a hard test.

The check is one extra sweep and it is the cheapest kind: run the fragment approximation at two basis sizes rather than one, and look at the ratio rather than the difference. A ratio that is stable licenses the extrapolation everyone is implicitly making; a ratio that grows says the validation does not transfer to the basis anybody will actually use.

Who found it, and when

The counterpoise correction is Boys and Bernardi’s, from 1970, and the observation that many-body counterpoise schemes for clusters are not simply additive has been in the literature since Valiron and Mayer’s work in the 1990s — the distinction between site-site functional counterpoise and the full scheme is exactly this problem, and their conclusion, that the pairwise version is not a sound approximation to the whole, is not overturned here.

What is added here is the basis dependence of the discrepancy, on a model small enough to sweep, and the practical corollary: that the discrepancy’s fraction grows as the basis improves, so the error and the correction cannot be assumed to shrink together.

Still open: a fourth centre, and the absolute error

The obvious open question is the fourth centre. With four fragments the pairwise assembly leaves out four three-body terms as well as the four-body one, and the question is whether the overshoot grows with the number of fragments as fast as it grows with the basis — which decides whether a pairwise-corrected calculation on a cluster of ten is usable at all. The calculation generalises directly and the cost is a larger linear solve.

The nearer question is the numerator. Everything above reports the fraction, because that is what is usually reported and what a practitioner would use. But the absolute difference between the assembled correction and the true one is what actually enters an energy, and it is falling across the sweep even as the fraction rises — from about 2×1032 \times 10^{-3} hartree at one function to 5×1065 \times 10^{-6} at six. Which of the two a reader should care about depends on whether the correction is being compared against a chemical accuracy target or against itself, and the two answers point in opposite directions. Plotting both against the basis size, with a chemical accuracy line drawn across it, would say at what basis size the assembly’s error stops mattering in absolute terms — and that is a number a practitioner could use, where a growing percentage is only a warning.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

Basis setBasis set superposition errorCounterpoiseFragment methodGaussian basisMany-body expansionVariational