Orbitals

The half that cannot be computed

How many points does each half of a counterpoise correction need to interpolate? The natural expectation is two different numbers. The answer is that the question is not yet askable: the lighter centre's half is a difference of two energies agreeing to six figures, its second differences are seven per cent of its own value, and no interpolation of it means anything. The third thing worth checking — the symmetric-pair check — works perfectly.

Worth reading first: A correction that is two functions · A contraction is a decision made once.

A correction that is two functions split a counterpoise correction into the two halves it is made of and found them unequal — a symmetric pair’s are equal to twelve decimals and an unequal pair’s run from 0.03 per cent to 67 and change places at 6.29 bohr. It ended by naming three things, and here the first is blocked, the second unavailable and the third exactly as good as advertised.

The three were an interpolation, a use for it, and a check. The first was this:

Three computed values of each half and a quadratic through each would restore the derivative that freezing destroys, and the interesting output is how many points each half needs — the heavy centre’s is smooth and the light centre’s is not, so the answer is probably different numbers for the two.

The prediction is right about which half is smooth and wrong about what that means. The lighter centre’s half is not a rough function. It is not a function — the values it takes at neighbouring separations are not related to each other in any way an interpolant can use, and the first thing to do is measure that rather than assume the opposite.

One half is a curve and the other is not. The two halves of a counterpoise correction for an unequal pair, against separation, on a logarithmic axis. The heavier centre's falls smoothly over two decades; the lighter centre's scatters over more than one decade between neighbouring points. It is not a rough function — it is a difference of two energies of order a hartree whose difference is a millionth, and the solver does not have seven figures to spare.
Fig. 1 The two halves against separation. One falls smoothly over two decades; the other scatters over more than one between neighbouring points.

Measuring the roughness before interpolating it

A second difference on a uniform grid separates a smooth function from noise without any fitting: a smooth function’s is a small fraction of its value and shrinks with the spacing, and noise’s does not.

half median second difference worst
the heavier centre’s 0.27% 0.47%
the lighter centre’s 7.1% 588%
How rough each half is, measured. The second difference of each half on a uniform grid, as a fraction of the half's own value. A smooth function's is a small number that shrinks with the spacing; noise's does not. The heavier centre's median is 0.27 per cent and the lighter centre's is 7.1, with a worst of 588. That is the answer to how many points does each half need and it is not a number of points.
Fig. 2 The second difference of each half as a fraction of its own value, point by point.

The worst is worth looking at rather than dismissing as an outlier. At a separation of 3.79 bohr the lighter half comes out at 1.6 × 10⁻⁷, and its neighbours on either side are 3.1 × 10⁻⁶ and 4.1 × 10⁻⁶ — twenty times larger. A physical quantity does not do that; a difference of two numbers agreeing to seven figures does, whenever the seventh figure happens to fall the other way. The point is not that one value is wrong: it is that none of them is right to better than the scatter, and the scatter is the size of the quantity.

Seven per cent at the median and five hundred and eighty-eight at the worst. No interpolation of that quantity says anything, at any order, and the question — how many points does each half need — has no answer for one of the two.

Where the noise comes from

It is not a defect of the model and it is not a defect of the solver. It is a cancellation.

Each half is the difference between an atom’s energy alone and the same atom’s energy with its partner’s basis functions present as ghosts. Both energies are of order a hartree. Their difference is:

separation the heavier half figures needed the lighter half figures needed
1.20 1.1 × 10⁻³ 3 3.9 × 10⁻⁵ 4
2.11 6.1 × 10⁻⁴ 3 5.4 × 10⁻⁶ 5
3.37 2.4 × 10⁻⁴ 4 3.1 × 10⁻⁶ 6
3.79 1.7 × 10⁻⁴ 4 1.6 × 10⁻⁷ 7
Where the noise comes from. Each half is a difference of two energies, both near a hartree. The heavier centre's difference is a thousandth of them and the lighter centre's a millionth — so the second asks for six or seven figures of cancellation from a variational solve, and gets whatever the solver has. Nothing about the calculation warns of it: both differences are computed the same way and only one of them is meaningful.
Fig. 3 The two differences, and how many figures of cancellation each asks for. Both are computed the same way and only one of them survives it.

The lighter centre’s half asks for six or seven figures from a small variational solve that does not have them to spare. The heavier centre’s asks for three or four and gets them.

Nothing in the calculation announces this. Both halves come out of the same routine, in the same units, with the same number of decimal places printed — and one of them is a number and the other is the difference between two numbers that agree.

That is the general hazard with any quantity defined as a difference, and it turns up in at least three places. A shortfall in radius additivity is a difference of two lengths that agree to a tenth. A healing length is an excess over a bulk value that the two ends of a chain are trying to set. Here it is a difference of two energies. In each case the quantity’s own precision is far worse than the precision of what it is made from, and no routine reports that.

What that does to the interpolation

The other half is smooth and can be interpolated, and doing so exposes a second thing worth knowing.

The interpolation stops improving where the noise starts. The worst error of a polynomial through m computed values of the heavier centre's half, against m, as a fraction of the whole correction. It falls from 48 per cent at two points to about 1.5 at six and then stops. The half's own second differences say it is known to 0.27 per cent, and an interpolant through noisy nodes amplifies that — so the plateau is the noise rather than the function.
Fig. 4 The interpolation error of the heavier half against how many points it is given, with the half’s own noise drawn across.

The error falls from 48 per cent at two points to about 1.4 at six — and then stops falling. Six points, seven points, eight points: 1.4, 1.9, 1.7 per cent.

A polynomial through noisy values cannot be more accurate than its values, and it is usually worse: an interpolant amplifies the error at its nodes by a factor that grows with the order. The half is known to 0.27 per cent, the interpolation stalls at about 1.5, and the ratio is the amplification.

So the answer to how many points is six, and the reason is not the shape of the function. Adding a seventh point makes the interpolation worse — which is the clearest possible sign that the limit is the values rather than the polynomial, since a smooth function’s interpolant never gets worse with more Lobatto points.

The check that does work

The third suggestion was the one that costs nothing, and it is exactly as good as it claimed.

A symmetric pair’s two halves are equal by symmetry. Computed here they agree to 7 × 10⁻¹⁶ — the last digit the solver carries.

The check that costs one extra energy. A symmetric pair's two counterpoise halves are equal by symmetry, and computed here they agree to 7.2e-16 — the last digit the solver carries. So any calculation reporting them unequal has a bug, and the test needs one more energy than the correction itself does. It is available to anybody computing a counterpoise correction at all, and it is rarely made because the halves are rarely separated.
Fig. 5 The two halves of a symmetric pair, and the difference a deliberately broken calculation gives.

Leaving out one ghost function — the commonest way to get a counterpoise correction wrong — moves the two halves apart by 6 × 10⁻⁷, a factor of a thousand million above the correct calculation’s agreement.

And it is invisible in the total. The same error moves the whole correction by 6.9 per cent, which is a size an ordinary basis-set effect has and which nothing marks as wrong.

What a broken calculation looks like from two angles. A symmetric pair computed correctly and with one ghost function left out — the commonest way to get a counterpoise correction wrong. The total moves by seven per cent, which is a size an ordinary basis-set effect has and which nothing marks as an error. The share moves from a half to 0.4628, and a half is the right answer by symmetry, so a departure has no other explanation.
Fig. 6 The same broken calculation seen two ways. One number is seven per cent out and could be anything; the other is a symmetry violation.

That is the difference between a quantity with a known exact answer and one without. A total of 6.9 per cent out looks like chemistry; a share of 0.4628 where a half is required looks like a bug, because it is one. The value of a check is not how large a discrepancy it shows but whether the correct answer is known independently, and a symmetry gives one for nothing.

A general lesson about stalled convergence

The interpolation’s plateau is worth generalising, because it is a trap with no warning label.

A convergence that stops improving looks the same for two very different reasons. Either the approximation has reached the complexity of the thing it is approximating — in which case more effort is genuinely wasted — or it has reached the noise in its own inputs, in which case more effort makes things worse and the inputs are what need fixing. Both produce a curve that falls and then flattens.

The way to tell them apart is to measure the inputs’ noise separately, which costs one pass over a fine grid and no fitting at all. Here the measurement says 0.27 per cent and the plateau sits at 1.5 — a factor that is the interpolation’s own amplification — and that ratio is the evidence.

The same shape appears twice more in different clothes. A healing length that depends on the window it was fitted over is a fit reporting its own procedure; an exponent that belongs to its window is the same thing. The common repair is to measure the thing the fit is standing on before trusting the fit, and it is cheap in every case.

What can now be said about the three suggestions

Three statements, in the order they were asked for.

The interpolation is available for one half and not the other, and the asymmetry is the opposite way round from the one that matters. Six points for the heavier centre’s, and the lighter centre’s is not a candidate — which is a stronger version of the prediction and a discouraging one, since the lighter centre’s half is the part a chemist would most like to have, being the correction to the smaller fragment, which is the one a chemist is usually trying to price.

How many points depends on the noise rather than on the function, so six is a statement about this solver — the same shape of statement as the forty-five combinations a spectrum fixes, where what is determined depends on the measurement’s precision and not only on the model. A more accurate one would need fewer points, not more, and the number is not transferable.

And the free check is free and sharp, and it is the one of the three that a reader can use today without any of the calculation above. It should be made whenever a counterpoise correction is computed for a symmetric pair, which is the case chemists check their programs on.

What survives of the unequal division

It is worth being clear that this does not undermine the unequal-division finding, because the lighter half is noise could be read that way.

That finding was about the share, and a share is a ratio of the two halves. Where the lighter half is a hundredth of the total, the share is 0.99 and the noise in the numerator moves it in the third decimal — so the statement the correction is unequally divided is safe, and safest exactly where the lighter half is worst known.

And the crossover it located is at 6.29 bohr, where the two halves are comparable and both are therefore well determined. The original analysis chose to report the range and the crossover rather than the halves themselves, and that choice turns out to have been the robust one.

What this removes is a suggested use rather than an established result: the halves cannot be tabulated and interpolated separately, and the share can. A share is a ratio and a ratio of two badly cancelled numbers is not automatically bad — it is bad only when the numerator is the small one, which is the case here and is why the share is quoted rather than the halves.

Cancellation is not enough to explain it, and what is left is a familiar thesis

Lost to cancellation is the natural diagnosis and it does not survive arithmetic. It is worth doing that arithmetic, because what remains points somewhere more interesting.

Two energies agreeing to six figures, subtracted in double precision, leave a difference carrying about nine or ten significant digits. Nine digits is not noise; it is a smooth function measured to a part in a hundred million, and it would interpolate beautifully. The observed second differences are seven per cent of the value, which is barely one digit. Eight of the nine went somewhere other than the subtraction.

They went into the optimiser. Each energy in this essay is not a closed-form expression evaluated at a separation; it is the endpoint of a search over exponents, and a search stops when its stopping criterion is met. That criterion is on the energy.

Which is exactly the point about variational calculations. An energy converged to a tolerance leaves the wavefunction converged far less well, because the energy is stationary at the answer and everything else is not — a displacement in the exponents that costs a tenth of a microhartree in the energy moves the function by far more than a tenth of a microhartree’s worth. The exponents the optimiser lands on therefore vary from one separation to the next by an amount invisible in the energy and large in anything else, and one of the things they move is the small difference this essay is trying to read.

So the light half is noisy for the same reason a better energy is not a better answer: the quantity being converged is the least sensitive one available, and the quantity being read off is not.

That reframes the repair. More precision in the arithmetic would buy nothing, because the arithmetic is not what is short. What would buy something is holding the exponents fixed across the separations, so that the sequence is a smooth function of one variable rather than a sequence of independently optimised calculations — at the price of a basis that is no longer optimal at any of them.

What is quoted, and what is computed

A basis is a choice and a density is not is the standing statement about what a basis-set correction is correcting.

Nothing is quoted. One electron on a line, two centres with stated charges, a basis of three functions on each, and a variational solve.

The halves are differences of energies from the same solver, taken at forty-one separations on a uniform grid. The second differences are that grid’s, taken directly rather than fitted. The interpolants use Chebyshev–Lobatto nodes, which include the ends of the interval — the first-kind points do not, and using them made the error at the ends an extrapolation and stalled the convergence at a per cent for a reason that had nothing to do with the function.

The broken calculation leaves out one of the six ghost functions and is otherwise identical, so the comparison is between two runs of one routine.

What this cannot say

This is a one-electron model on a line. A real counterpoise correction is a many-electron quantity in three dimensions, and the cancellation is what carries over rather than the numbers: a real correction is millihartrees against total energies of tens or hundreds, which is a worse ratio than anything here.

The noise is this solver’s. A calculation carrying more digits would push the lighter centre’s half further before it disappeared into the cancellation, and nothing here says how many digits would be enough. What is transferable is the diagnosis: measure the second differences before interpolating, because a stalled convergence looks the same whether the function is complicated or the values are noisy.

A second difference is a crude smoothness measure. It is one of several, it is sensitive to the grid spacing, and a function with genuine structure on the scale of the spacing would be reported as noisy. What makes the diagnosis safe here is the ratio between the two halves — twenty-six to one on the same grid, from the same routine, at the same separations — rather than either number on its own.

And a symmetric pair is one test. It catches an asymmetric error — a missing ghost function on one centre — and is blind to a symmetric one. A calculation that left out the same function on both centres would pass it and be equally wrong.

What was checked

The lighter centre’s half is far rougher than the heavier centre’s, by a factor of more than twenty in the median second difference, and rough in absolute terms — two claims rather than one, because the first could hold with both halves smooth.

The heavier centre’s is smooth enough to interpolate, at under a per cent, which is what makes six points a statement about anything.

The interpolation stalls at the level of the half’s own noise rather than converging, checked as an inequality between the interpolation’s plateau and the measured noise.

And the symmetric check is exact. A correct calculation’s halves agree to the solver’s last digit; a broken one’s differ by a factor of a thousand million more; and the broken share departs from a half, which has no explanation but an error. The last is the refusal — a total that moves by seven per cent explains nothing, and the check is written on the share rather than on the total for exactly that reason.

Still open: computing the half directly, and three fragments

The obvious open question is the precision. The lighter centre’s half is lost to cancellation, and cancellation is a problem with a known family of answers: compute the difference directly rather than as a difference. The counterpoise half is a first-order energy change under adding basis functions, and a perturbation expression for it would involve the coupling between the two centres’ functions rather than two total energies — which is the same difference of approach a bond order taken from the eigenvectors has over a difference of energies, and it is worth the same.

The nearer question is the three-fragment case. An interaction energy between three fragments is assembled from pairwise counterpoise corrections, each of which is now known to be unequally divided — so the assembly double-counts one fragment’s share and undercounts another’s by amounts that do not cancel. Computing the three-body correction directly and comparing it against the assembled pairwise one is the first place the division has a consequence for a number anybody quotes, and it needs no more precision than the heavier halves already have — which is the one encouraging thing to say about the arithmetic.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBasisBasis set superpositionClosed formConvergenceLeast-squaresModel limitNormalisationReference stateVariational principle