Orbitals

A repair that costs more than the whole

Four fragments are the first system in which a counterpoise correction can be assembled from something between pairs and the whole. Adding the three-body increments leaves what is still missing below a kilocalorie a mole on every cell, widens the usable range for every arrangement and turns a single usable basis size into two — and under a cubic model of cost it is dearer than the full calculation it stands in for until the cluster has eleven fragments. Keeping only the consecutive triples, which pays from five, works for two arrangements and does worse than pairs for the third.

Worth reading first: The overshoot was one arrangement · The line was holding the answer up.

The overshoot was one arrangement found that adding a fourth fragment to a chain changes the window in which a pairwise counterpoise assembly can be trusted, in a direction that depends on where the heavy centres sit — wider for two arrangements, narrower for one — and that the assembly’s error changes sign between arrangements and even between fragments of one chain. It ended by naming the thing four fragments are the first to make possible.

With three fragments there is nothing between pairs and the whole. Each fragment’s correction can be assembled from the two pairs it belongs to, or computed in the full basis of all three, and a “three-body” assembly of three fragments is just the full calculation under another name. With four fragments that stops being true. Each fragment belongs to three pairs and to three triples, and the triples are real calculations in smaller bases than the whole. Restoring what the triples add to the pairs is a genuine intermediate: the next order of the expansion.

This essay measures what that order buys, and what it costs. It buys almost everything the pairwise assembly lacked. It costs more than the calculation it replaces.

Between pairs and the whole

For one fragment the construction runs as follows. Its correction from a set of ghost functions is how much its energy falls when those functions are added with their charges removed. The two-body terms are its corrections from each other fragment’s functions alone. A three-body increment is its correction from two others together, less the two corrections those two give separately — which is the part that exists only because both are present. The three-body assembly is the sum of the pairs plus the sum of those increments, and what it still misses of the full correction is the four-body remainder.

For three fragments that remainder is zero by construction, since the one triple is the whole, and the calculation confirms it to a part in 10¹² on every cell. That identity is the check that the increments are being built as defined. For four fragments the remainder is a number, and the question is how big.

Everything else is held where the four-centre comparison put it: one electron and one set of even-tempered s Gaussians a fragment, one to six functions a centre, four separations from 1.6 to 3 bohr, and the three four-centre arrangements — every centre alike, the heavy centres inside, the heavy centres outside.

What the triples leave out

What pairs miss, and what triples still miss. For every cell of the three four-centre arrangements, the size of the pairwise assembly's error and of the four-body remainder the three-body order leaves, in kilocalories a mole on a logarithmic scale. The pairwise error is above the line in 1, 2, 3 cells; the remainder is below it in every one of the seventy-two, and is at worst 0.612 kcal/mol.
Fig. 1 For every cell, the size of the pairwise error and of the remainder after triples, on a logarithmic scale against a kilocalorie a mole.

The four-body remainder is below a kilocalorie a mole on every one of the seventy-two cells. Its worst value is 0.153 kilocalories a mole with every centre alike, 0.325 with the heavy centres inside, and 0.612 with them outside — all at one Gaussian a centre and the closest separation, where every error in this problem is largest.

The pairwise error, by comparison, is above the line on one, two and three cells respectively. That sounds like a modest number to have fixed, and it is exactly the number that mattered: a window is set by its worst cell, so a single cell above the line removes a basis size from the usable set at every line below it. The triples do not have to improve the assembly everywhere to change the answer. They have to bring the worst cells down, and they do.

Most of the remainder figure is a gap between two clouds of points — the pairwise errors above, the remainders below, often by a factor of ten or more. It is not a uniform gap, and that turns out to matter.

A window’s top cannot move

The window for a basis size runs from its floor — the worst assembly error at any separation — to its top — the smallest correction at any separation. The top is the correction itself, and no way of assembling a correction changes what the correction is. So whatever the triples do, they cannot raise a single window’s top, and every change they make to the usable range is made by moving floors.

That is a strong constraint and it has an immediate consequence. A basis size whose correction has fallen below the line cannot be rescued by any assembly at all, however good. The largest basis sizes, whose corrections are hundredths or thousandths of a kilocalorie a mole, will stay unusable at the standard line under triples, quadruples or the exact correction; what is usable at a given line is bounded above by the correction and only the floor is negotiable.

Every top stays where it was, and most floors come down. For each four-centre arrangement and each basis size, the window of usable accuracy lines when the correction is assembled from pairs (the thin bar) and when triples are added (the thick one). Every thick bar ends exactly where its thin bar does, because the top of a window is the smallest correction and no assembly changes the correction. The floor is the worst assembly error, and adding triples lowers 16 of the eighteen floors and raises 2, both at three Gaussians a centre — so the next order of the expansion is not an improvement at every basis size, only at most of them.
Fig. 2 Each basis size’s window assembled from pairs (thin) and with triples added (thick). The tops coincide; the floors move.

Of the eighteen floors, sixteen come down. Two go up — with every centre alike and with the heavy centres outside, both at three Gaussians a centre, from 0.0345 to 0.0421 and from 0.112 to 0.124 kilocalories a mole. At those two basis sizes the three-body assembly’s worst error is larger than the pairwise one’s. It is a small effect on windows that are not near the standard line, and it is the first sign that the next order is not a uniform improvement.

The answer stops being one size

The usable stretches, assembled from pairs and from triples. For each four-centre arrangement, the stretches of accuracy line with a usable basis size when each fragment's correction is assembled from pairs, and when the three-body increments are added to it, labelled with the sizes usable there. The covered share rises from 58% to 67% (1-1-1-1), 54% to 75% (1-2-2-1), 60% to 80% (2-1-1-2), and stretches where more than one basis size is usable appear where there were none.
Fig. 3 The stretches of accuracy line with a usable basis size, assembled from pairs and with triples, labelled with the sizes usable in each.

Read across three decades of line, the covered share rises in every arrangement: from 58 to 67 per cent with every centre alike, from 54 to 75 with the heavy centres inside, and from 60 to 80 with them outside. The arrangement still matters — the three shares are not the same, and they spread further apart rather than closer — but the order among them changes, and the arrangement that did worst under pairs, with the heavy centres inside, gains the most: twenty-two points, against twenty with the heavy centres outside and nine with every centre alike.

The larger change is in the shape of the answer. Under pairs, the stretches with an answer almost all had exactly one usable basis size, which is what made the single-size result and its sensitivity to the line the right things to report. With triples, more than one basis size is usable over 23 per cent of the range with every centre alike, 60 per cent with the heavy centres inside and 38 per cent with them outside, where pairs gave 0, 0 and 3.

At the standard line itself: the uniform chain, which had no usable basis size at a kilocalorie a mole under pairs, has one — one Gaussian a centre, over a window from 0.696 to 4.20. The other two arrangements each have two, one and two Gaussians a centre, over windows from 0.575 to 2.03 and from 0.613 to 2.10. The standard line sits more than 40 per cent above the lower edge in every case.

So the fragility that ran through this whole question — one basis size, a window seven tenths of a per cent from vanishing, an answer the arrangement could erase — was a property of truncating the expansion at pairs. One more order, and it is gone.

Nine cells made worse

The floors that rose said the improvement was not uniform, and the cells say how far from uniform it is.

The cells the three-body order makes worse. For each cell of the three four-centre arrangements, the size of the error after the three-body increments are added, divided by the size of the pairwise error. Below one the extra order helps; above one it hurts. It hurts on 5, 0 and 4 of twenty-four cells, and every one of those cells is in an arrangement whose pairwise error changes sign from cell to cell; with the heavy centres inside, where the pairwise assembly over-corrects everywhere, there are none. The usable range widens anyway, because a window is set by its worst cell.
Fig. 4 The size of the error with triples divided by the size of the pairwise error, cell by cell. Above one, the next order made things worse.

With every centre alike the three-body assembly’s error is larger than the pairwise one on five cells of twenty-four. With the heavy centres outside it is larger on four. With the heavy centres inside — the arrangement whose pairwise assembly over-corrects on every cell — it is larger on none. In one heavy-outside cell the ratio is twenty-two, because the pairwise error there had almost vanished and the three-body one had not.

The pattern in where the worse cells sit is exact, and it is the pattern the previous comparison would predict. Every one of the nine is in an arrangement whose pairwise error changes sign from cell to cell. Where the pairwise assembly errs in one direction everywhere, the three-body increments correct that error on every cell; where two contributions of opposite sign compete, adding the next order can overshoot in the other direction. What does not hold up is the tempting explanation that the worse cells are simply cells where fragment errors had happened to cancel: six of the nine do contain fragment errors of both signs, but in most of those the cancellation removed less than a fifth of the fragments’ combined error.

It is the same shape as a finding about basis sets where improving the energy of hydrogen made its mean radius worse, and as the general statement of it: the quantity a procedure is built to improve is the one it is least likely to be wrong about, and nothing follows for the rest. A systematic improvement — a better basis, a higher order — is an improvement in the quantity it is built to improve and in aggregate, and has no obligation to be an improvement in every particular. Here the aggregate is the usable range, which is set by worst cells and improves in every arrangement, and the particulars are individual cells, nine of which get worse.

The price

None of that says whether the three-body assembly is worth computing, and the answer depends on what it costs relative to the calculation it stands in for.

The model of cost used here is deliberately simple and says what it charges. Each fragment is solved once in every set of ghost functions its assembly needs, and each solve is charged the cube of the number of functions, which is what diagonalising dominates as. Nothing is charged for integrals. With n functions a centre and N fragments, the full calculation solves each fragment once in all N·n functions; the pairwise assembly solves each in N − 1 sets of 2n; and the three-body assembly adds, for each fragment, one solve in each of its (N − 1)(N − 2)/2 sets of 3n.

What each order costs, against the full cluster calculation. A model of cost, charging each fragment one diagonalisation for every set of ghost functions it is solved in, at a price cubic in the number of functions and with nothing charged for integrals. Against the full cluster calculation, the pairwise assembly is always cheaper. The three-body assembly costs 1.64 times the full calculation at four fragments and becomes the cheaper of the two only at 11, where the ratio is 0.97. Keeping only the triples whose fragments are consecutive costs 1.008 times the full calculation at four fragments and is the cheaper from 5.
Fig. 5 The cost of each assembly as a multiple of the full calculation, against the number of fragments.

For four fragments the pairwise assembly costs 0.375 times the full calculation, and the three-body assembly costs 1.64 times. The repair is dearer than the thing repaired. It stays dearer as the cluster grows — 1.44 times at six fragments, 1.04 times at ten — and becomes the cheaper of the two only at eleven, where it costs 0.973 times.

That turns the three-body assembly of a four-fragment system from a method into a diagnosis. It says what the pairwise assembly is missing and how the window would look without the missing part; it is not a way of getting the correction more cheaply, because the full calculation is cheaper. The logic is the logic of a correction computed on a small system and added to a large one: a composite scheme pays off only when the expensive part is small compared with what it replaces, and here, until eleven fragments, the “small” part is the expensive one.

The triples a chain would keep

A real assembly does not keep every triple. A trio of fragments with a gap in it spans more distance than any pair it contains, and the obvious economy is to keep only the triples whose three fragments are consecutive, whose number grows in proportion to the length of a chain rather than as its cube. Under the same model of cost, that assembly costs 1.008 times the full calculation at four fragments, 0.645 times at five and about a ninth of it at eleven. It is the version anybody would actually run, and four centres are already enough to test it, because an end fragment of a four-centre chain belongs to one consecutive triple and to two with a gap.

What dropping the triples with a gap in them does. For each four-centre arrangement, the share of three decades of accuracy line with a usable basis size when the correction is assembled from pairs, from pairs and only the triples whose three fragments are consecutive, and from pairs and every triple. With the heavy centres inside or outside, the consecutive triples give the same covered share as every triple — 75% and 80%. With every centre alike they give 52%, less than pairs alone at 58%, because there the triples with a gap carry a median of 42 per cent of each cell's three-body increment, against 16 and 25 in the other two.
Fig. 6 The covered share of the line assembled from pairs, from pairs and consecutive triples, and from pairs and every triple.

With the heavy centres inside or outside, dropping the triples with a gap costs nothing in usable range: the covered share is 75 and 80 per cent, exactly what every triple gives, and the consecutive assembly is worse than pairs on a single cell in each arrangement. The triples with a gap carry a median of 16 and 25 per cent of each cell’s three-body increment in those chains, and what they carry is not where the worst cells are. The answers at the standard line are not quite identical — with the heavy centres outside, one Gaussian a centre drops out of the usable set at a kilocalorie a mole, because one cell’s consecutive-triple error is 1.04 — but the range over which some basis size is usable is untouched.

With every centre alike the result is different, and it is worse than having no triples at all. The consecutive assembly’s covered share is 52 per cent, below the 58 per cent pairs give, and it is worse than pairs on twelve cells of twenty-four. There the triples with a gap carry a median of 42 per cent of the three-body increment. Light centres have diffuse functions, and the help a fragment gets from two sets of ghost functions together does not stay among neighbours — for the same reason that most of an overlap can lie outside the surfaces drawn for it: a function’s reach is set by its tail, and a tail does not respect the spacing of a chain. Keeping the consecutive triples restores part of the three-body order and not the part that balances the rest, and a partial correction of that kind can be further from the answer than no correction.

So the dependence on arrangement that the pairwise comparison found comes back one level up, in a sharper form. Whether the triples a practical assembly would drop can safely be dropped depends on how far the fragments’ functions reach — and the arrangement with nothing chosen about it is the one where they cannot.

What was computed, and from what

Each fragment’s energy was computed alone, in each other fragment’s functions, in each pair of other fragments’ functions and in all of them, for six basis sizes at four separations and three arrangements. Every solve is the one-electron generalised eigenvalue problem in s Gaussians on a line that the basis the other atom lent set up, and the increments and remainders are sums and differences of those energies with nothing fitted.

The three four-centre arrangements, stopped at pairs and at triples. For each four-centre arrangement: the covered share of three decades of accuracy line with pairs and with triples, the share over which more than one basis size is usable, the answer at a kilocalorie a mole in each reading, the worst four-body remainder, and how many of twenty-four cells the three-body order makes worse.
Fig. 7 The three four-centre arrangements under pairs and under triples: covered share, several usable sizes, the answer at a kilocalorie a mole, the worst remainder and the cells made worse.

The checks are these. For three fragments the three-body assembly equals the full correction to a part in 10¹². For four, the remainder is below the line on every cell of every arrangement. Adding triples widens the covered share in every arrangement, leaves every window’s top exactly where it was, and gives an answer at the standard line in every arrangement. It lowers sixteen floors and raises two, both at three Gaussians a centre. The cells it makes worse all lie in the arrangements whose pairwise error has both signs. The pairwise error alone is above the line on one, two and three cells. Keeping only consecutive triples gives the covered share every triple gives with the heavy centres inside or outside, and less than pairs with every centre alike, where the triples with a gap carry by far the largest share of the increment. And under the stated cost model the three-body assembly of four fragments costs more than the full calculation and first costs less at eleven, while the consecutive assembly costs within a per cent of it at four and less from five.

What the cost model leaves out

The crossover at eleven is a property of the model, and the model has two simplifications that pull in opposite directions.

It charges nothing for integrals. In a one-electron problem that is nearly right, because every integral here is a closed form over a pair of Gaussians. In a real calculation the two-electron integrals scale faster than the diagonalisation with the number of functions, which would make the large full calculation relatively dearer and move the crossover down.

And it charges every solve as a dense diagonalisation. That is the right model for a problem this size and not for a large calculation, where iterative solvers and integral screening lower the exponent — which favours the many small solves of an assembly over the one large solve of the whole, and would also bring both crossovers earlier.

Both simplifications favour the assemblies, so eleven and five are best read as ceilings on where each becomes worth running rather than as estimates of it. Neither changes the verdict at four fragments, where the consecutive assembly already costs the same as the full calculation and the complete one costs more.

The rest of the limits are the ones the pairwise comparison already carried: one electron a fragment and no correlation; s functions on a line, equally spaced; charges of one and two; three arrangements; and “usable at every separation” as the test throughout.

Pricing a repair

The transferable point is that an improvement to an approximation has to be priced against the thing being approximated, and not against the approximation.

Against the pairwise assembly, the three-body one is an unambiguous success: the remainder is below the line everywhere, the covered range grows by a sixth to two fifths, and the brittle single-size answer turns into a robust range. Anyone comparing the two alone would adopt it. Against the full calculation it is a loss for every cluster small enough to have been computed here, because the full calculation delivers everything the three-body assembly delivers, exactly, for less. The hierarchy of approximations has a natural direction of improvement and no guarantee that a higher order costs less than the calculation it is converging on.

That is easy to miss because an expansion in orders looks like a sequence of cheap corrections to an expensive target. Whether it is depends on how the cost of each order grows against the cost of the target, and for a cluster of four the answer is that the second correction already costs more than the target does.

Where the pieces come from

The counterpoise correction is Boys and Bernardi’s, and its extension to clusters as a hierarchy of corrections at increasing orders is Valiron and Mayer’s. The one-electron chains, the remainders, the windows and the cost comparison are new here.

The four-centre comparison deserves the credit for the design that made this measurable: three arrangements differing by one centre, a fixed grid of basis sizes and separations, and windows read exactly from two numbers each. None of this needed a new calculation of any kind beyond the triples.

Still open: wider gaps, and the strictness of the test

The obvious open question is the length of the chain. Four centres show that dropping the triples with a gap works for two arrangements and fails for the uniform one, and on four centres a gap is at most one fragment wide. On a chain of six or eight a triple can have a gap of two or three separations, and the question is whether the uniform chain’s gap triples keep carrying two fifths of the increment as the gap widens or fall away with it. That decides whether the consecutive assembly fails for uniform chains in general or only for short ones — the difference between a method with an exception and a method with a warning attached — and it is where the saving begins, since the consecutive assembly is already cheaper than the whole at five fragments.

The nearer question is the one both pairwise essays deferred: the strictness of “usable at every separation”. It is what made single cells load-bearing, and the triples have now shown that most of what a single cell did was a property of stopping at pairs. Whether the weaker test — usable at the separations a calculation would actually use — changes the three-body picture at all, or only the pairwise one, would say how much of the remaining arrangement dependence belongs to the method and how much to the test.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBasis set superposition errorCounterpoiseFragment methodMany-body expansionModel limit