What a spectrum settles

The correction that was invented

A standard error model puts the zero-point error in a rotational constant at a few tenths of a per cent, shared between the three moments by invented weights, with the parent and its deuterated form differing by five. Computed from a force field it is 1.88 per cent, one of the three shares is negative, and the mismatch is 35 — so the cancellation the substitution method rests on is worth a factor of 2.5 and not thirteen.

Worth reading first: The coordinate an isotope reports · The force field is not in the spectrum.

The coordinate an isotope reports took Kraitchman’s substitution method apart and found that its errors come from two competing effects: an amplification, because the equations divide by the change in a planar moment and that change is small; and a cancellation, because the zero-point error in the parent and in the substituted species are nearly the same and mostly subtract out.

Both were computed against an error model, and the model was invented. It said so: a systematic inflation of each moment by a few tenths of a per cent, shared between the three by weights of 1, 0.62 and 0.81, with the substituted species inflated by the same fractions times 1.05 — a five per cent mismatch. Everything it concluded, including a cancellation worth a factor of thirteen, rests on those numbers. That is the shape of finding an estimate that can be wrong by two approaches from the other side: there the invented number turned out not to matter, and here it does.

They are computable. Harmonic force fields exist for these molecules, and a zero-point correction to a moment of inertia is a second derivative and an amplitude.

Every one of them was optimistic.

The factor of thirteen was generous. An invented error model against the one computed from a force field. Its invented mismatch of five per cent made the correlation between the parent and the substituted species worth a factor of 12; the computed mismatch of 35 per cent makes it worth 2.5. The substitution structure went from being 2.5 times worse than a direct fit to 26, and at the computed size of the correction its worst coordinate is out by 10.7 per cent.
Fig. 1 The invented error model against the computed one, on the three quantities the error model reported. None of the three moves in the direction that helps.

The correction, computed

A measured moment of inertia is an average over the zero-point motion, so to second order

IIe=12k2IQk2Qk2,\langle I \rangle - I_e = \tfrac{1}{2} \sum_k \frac{\partial^2 I}{\partial Q_k^2} \left\langle Q_k^2 \right\rangle,

and both factors are already here. The amplitude Qk2=/2ωk\langle Q_k^2 \rangle = \hbar/2\omega_k is computed for every normal coordinate; the second derivative comes from displacing the molecule along the mode and recomputing the moments.

For the four molecules with force fields:

correction to each moment
water −0.46%, +1.88%, +1.07%
ammonia +1.47%, +1.47%, +1.94%
methane +1.68%, +1.68%, +1.68%
boron trifluoride +0.16%, +0.16%, +0.12%

A per cent or two, not a few tenths. A force field fitted to frequencies is what makes the computation possible at all, and it is the same six molecules that limit it. Water’s largest is five times the modelled estimate, and only boron trifluoride — whose atoms are heavy and whose amplitudes are therefore small — comes in under a fifth of a per cent.

And one of water’s three is negative. The zero-point motion moves mass towards one principal axis and away from another, so one moment shrinks while the others grow. No model of the form multiply every moment by one plus a positive fraction can produce that, and the invented model was exactly that form.

The zero-point correction, computed rather than assumed. How much each molecule's moments of inertia are inflated by its own zero-point motion, from its harmonic force field: ½ Σ (∂²I/∂Q²)⟨Q²⟩ with the computed zero-point amplitudes. Water's three come out at -0.46%, 1.88%, 1.07% — one of them negative, which no uniform inflation can produce. A symmetric top's equal moments are averaged, because a displacement splits them and a second difference of either alone measures the splitting rather than the average.
Fig. 2 The computed corrections, with each molecule’s three moments drawn either side of zero. Water’s first bar goes the other way from the other two.

The degeneracy that had to be handled

A symmetric top’s two equal moments are not analytic functions of a displacement, and taking a second difference of either one alone produces a number about the wrong thing.

The reason is that a distortion splits them: one goes up and the other down by nearly as much. Methane’s three moments, taken one at a time, give corrections of −58, +2 and +61 per cent — enormous, opposite in sign, and about the splitting rather than about the average. Averaged over the degenerate set they give +1.68 per cent, which is the physical answer and is of the same size as water’s.

So every degenerate set is averaged, and the tolerance for calling two moments equal is relative and generous. That is not fastidiousness: ammonia’s two equal moments differ in the seventh figure, which is the geometry’s own rounding, and at a tight tolerance the pair goes unrecognised and its corrections come back at ±40 per cent.

Both halves are checked — the raw corrections must be enormous and the averaged ones must be small — because a routine that quietly averaged everything would fail as surely as one that averaged nothing.

The mismatch, which was the important one

The correction’s size affects the amplification. Its correlation between the parent and the substituted species is what the whole method rests on, and that is the number the model invented most freely.

Water’s fractional corrections against monodeuterated water’s:

moment H₂O HDO they differ by
a −0.46% −0.76% 65%
b +1.88% +1.49% −21%
c +1.07% +0.85% −20%

Twenty to sixty-five per cent, against an assumed five. A parameter fitted to nothing is a parameter that can be wrong by any amount, which is the general form of what a curve’s own information content measures. The typical disagreement is 35 per cent, seven times what the model supposed.

That is the finding, and it goes straight into the method’s arithmetic. The cancellation is what survives when the two errors subtract, and it survives in proportion to how badly they match.

The factor of thirteen

Feeding the computed weights and mismatch back into the substitution analysis, on the same molecule and at the same range of errors:

invented computed
what the correlation buys 12× 2.5×
worse than a structure fitted directly 2.5× 26×
worst coordinate error, at the computed size 1.9% 10.7%

The factor of thirteen was generous by nearly a factor of five. The correlation between the parent and the substituted species is real and it buys a factor of two and a half, which is worth having and is not the argument the method is usually defended with.

The second row is the one that changes the verdict. Under the invented model, a substitution structure was two and a half times worse than reading lengths straight off the rotational constants — bad, but the same order. Under the computed one it is twenty-six times worse, which is a different claim about the method entirely.

And the third row is what a user would meet: at the computed size of the correction, the worst coordinate of a formaldehyde substitution structure is out by ten per cent. That is not a structural refinement; it is a structural error.

What a correlated error is worth, and what it is not enough for. The largest relative error in a recovered coordinate, at four fractional errors in the moments. The first column has the substituted species inheriting all but a twentieth of the parent's error, which is what two isotopologues of one molecule do; the second has the two errors independent. The ratio between them — about 13.2 — is the whole of the cancellation the method is famous for. The last column is what the same error does to a length read straight off a moment, and the substitution structure is a factor of 2.5 worse than it, because the difference the equations act on is a small fraction of the moments.
Fig. 3 The original picture of the cancellation, which is not changed here — only the number that goes into it. What correlation buys is a real effect measured against a mismatch, and the mismatch was the invented part.

What the method is actually doing

None of this says the substitution method is wrong, and it is worth being precise about what it does say.

The method is not used without a correction. Nobody in the field takes ground-state constants, runs Kraitchman and publishes the result as an equilibrium structure; the vibrational correction is computed and applied, and the whole apparatus of mass-dependent methods exists because everybody knows it has to be.

What is established here is the size of what has to be applied, and it is larger than assumed on both counts. A correction of one to two per cent in a moment, with a correlation between isotopologues that buys a factor of two and a half rather than thirteen, is a correction that has to be computed properly rather than absorbed.

And it says which molecules are safe — which is more than three numbers is not a structure could say, since that difficulty is a shortage of data and this one is a shortage of accuracy. Boron trifluoride’s correction is a sixth of water’s, for the obvious reason: its atoms are heavy, its amplitudes are small, and the whole effect scales with the amplitude squared. A substitution structure on a molecule of heavy atoms is a much better proposition than one on a hydride, and the difference is an order of magnitude rather than a nuance.

A bond's length, and how much of it is uncertain. One bond and one angle from each of four molecules, with the zero-point spread of each beside its value. Every bond here is uncertain by about seven per cent of its own length, and every angle by eight degrees or more.
Fig. 4 Where the amplitudes come from, applied to the quantities a chemist recognises: one bond and one angle from each of four molecules, with the zero-point spread of each beside its value. Everything above is that spread squared, applied to a moment of inertia instead.

Why the negative correction is the sharpest part

Of the three findings the negative moment is the one that could not have been guessed, and it is worth saying what it rules out.

An error model of the form every moment is inflated by ε times a weight, and every weight is positive is not a bad guess about the size. It is a structurally wrong guess: it can never produce a shrinking moment, so it can never produce the cancellation pattern that a real molecule has, however its weights are tuned.

And the pattern matters more than the size. Kraitchman’s equations are driven by differences of planar moments, so what propagates into a coordinate is the difference between what happens to two moments — and two corrections of opposite sign make a much larger difference than two of the same sign and different magnitude. A model with all-positive weights therefore understates the error twice over: once in the size and once in the arrangement.

The physical reason is worth having too, because it is not subtle. The bending mode of water moves both hydrogens towards the same side, which brings mass closer to one principal axis while moving it away from the other two. Every molecule with a bend has one moment that can shrink, and every hydride’s bend has a large amplitude.

So the sign is a property of the mode rather than of the molecule, and the check to make on any error model of this kind is whether it can produce one.

What is quoted, and what is computed

The frequencies each field is fitted to are quoted, and they are measurements. Everything else is computed: the force fields by least squares against those frequencies, the normal coordinates from the fields, the amplitudes from the frequencies, the second derivatives by displacement, and the corrections by summing.

The correction is computed on water and applied to formaldehyde, and that is stated rather than hidden. Formaldehyde has no force field in this collection — the fields here are six molecules — so what is transferred is the shape of the error model rather than a number belonging to formaldehyde. A reader should take the ratios seriously and the third decimal place not at all.

The substitution analysis is used unchanged. Only its two inputs are replaced, which is what makes the comparison a comparison.

What this cannot say

A harmonic field cannot give the whole correction. The real vibration–rotation interaction constants have a cubic contribution this model has no access to, and it is not small — for a hydride it is comparable with the harmonic part. So the computed correction is a lower bound on the size, which strengthens the finding, and is not a substitute for a proper calculation.

Second order in the amplitude. The expansion truncates at Q2\langle Q^2 \rangle, which is the leading term and is the one an anharmonic treatment corrects. The truncation is why the numbers should be read as a size and a sign rather than as a correction to apply.

Four molecules is not a survey. They are the four with fitted fields available, and those fields were fitted for other reasons; three of them are hydrides and one is not, which is exactly the wrong balance for a claim about how the correction scales with mass.

Nothing here is a measured correction. Every number is a model’s own estimate of its own error, which is the only kind available without an equilibrium structure to compare against — and a molecule with a known equilibrium structure would settle it in one line.

And a mismatch of 35 per cent is water’s. A heavier molecule’s isotopologues have more nearly equal amplitudes, so its mismatch is smaller and its cancellation better — which is the same conclusion the size of the correction gives, arriving by a second route.

Every sign, lost. Formaldehyde in its own principal axes. Open circles are the atoms where they are; filled ones are where Kraitchman's equations put them, from the change in the three moments when each atom in turn is made heavier. The two agree to 7.6e-8 ångström — the equations are an identity for a rigid structure — but they return the square of each coordinate, so the two hydrogens at b = ±0.9348 both come back at +0.9348 and land on the same point.
Fig. 5 The coordinates the correction is applied to, and the reason it has to be. A substitution structure is a set of squared coordinates, so every sign has to be supplied from outside the measurement — and a coordinate near zero is the case where the square is smaller than the correction being added to it.
The coordinate that comes back imaginary. The square of the out-of-plane coordinate returned by Kraitchman's equations for each of formaldehyde's four atoms, against a fractional error in the moments. Every atom lies in the plane, so every one of these is exactly zero for a rigid structure — and a fractional error of 0.00001 already sends them negative, so the method returns the square root of a negative number. That is Costain's case, and it is the reason a substitution structure cannot locate an atom near a principal plane.
Fig. 6 Costain’s rule against the thing it was standing in for. The rule is an error estimate fitted to experience — 0.15 divided by the coordinate, in ångströms — and the quantity it estimates is now computable. Where the coordinate is large the two agree well enough that nobody would have questioned the rule; where it is small they part company, and the rule is the more optimistic of the two.

What was checked

The correction is a per cent or two rather than a few tenths, checked as a bound rather than as a value, because it differs between molecules by a factor of ten.

One of the three moments shrinks while the others grow, which is the qualitative statement no uniform inflation can make.

The mismatch is far more than five per cent, checked against the invented number rather than against a round figure.

Methane’s raw corrections are enormous and its averaged ones are small — both, so a routine that averaged everything or nothing would fail.

An asymmetric top has no degenerate set and nothing is averaged, which is the other half of the same check.

The invented mismatch made the cancellation worth several times what it is, and the substitution structure loses to a direct fit by far more than it appeared to — the two conclusions, each checked as a ratio between the two models so that a change to either would be caught.

And one of the computed weights is negative, which the invented set could not be.

How an invented number survives

A weight invented rather than computed is a specific kind of defect, and it is worth naming how such a thing persists — because the answer is not that anybody was careless.

An invented parameter survives when nothing it produces could contradict it. These weights shared a correction between three moments; any set of weights summing correctly produces a correction of about the right size; and the resulting structure was compared against structures obtained the same way, which used the same weights.

So the number was never tested. It was consistent with every comparison it entered, and consistency with a comparison that uses the same assumption is not evidence about the assumption.

The repair is the one performed here and it has a general form. Compute the quantity from something the assumption had no part in — here, a force field fitted to vibrational frequencies, which knows nothing about rotational constants — and compare. The answer came out 1.88 per cent against a few tenths, with one share of the wrong sign, and none of that could have emerged from within the original scheme.

The test to apply to any such number is one sentence long: what measurement would have come out differently if this were wrong? A parameter with no answer to that question is not a parameter that has been verified; it is one that has never been asked.

Why the invented numbers were both too small

Two invented numbers, both optimistic in the same direction, is a pattern rather than a coincidence, and the reason is worth stating.

An error model written to make a method look workable is written by somebody who expects the method to work. The error model was written to analyse the substitution method’s errors, not to attack it, and the natural thing to assume about a systematic error is that it is small and that it cancels — because those are the assumptions under which the method is worth analysing at all.

Neither assumption was checked against anything, and both were available to check. The force fields existed, the amplitudes were already computed for a different purpose, and the second derivative of a moment of inertia is three lines of arithmetic. What stopped it was not difficulty; it was that the numbers were plausible and nothing depended on them being right — until the conclusion did.

The habit that follows is narrow and cheap: when a model needs a number nobody has, compute the number before computing the conclusion, or state the conclusion as a function of it. The error model did the second, honestly and in so many words, which is why this correction could be made at all — and an estimate that can be wrong by two is the same discipline applied where the answer turned out not to matter.

Still open: the cubic term, a mass formula, and formaldehyde

The obvious open question is the anharmonic half. The correction computed here is second order in the amplitude and harmonic in the field, and a real vibration–rotation constant has a cubic term of comparable size — so the honest next step is a cubic field, which is not attempted here for the same reason a quick Gaussian SCF is not: a wrong one produces plausible numbers. What is available without it is the sign: the harmonic term computed here and the cubic term are known to have opposite signs for a stretching coordinate, so the true correction is smaller than this one and the direction of the error is known even where its size is not.

The nearer question is the one the table’s last row keeps raising. Boron trifluoride’s correction is a sixth of water’s, methane’s is comparable with water’s, and the difference is the amplitude — which goes as the reciprocal square root of the mass and of the frequency. Writing the correction as a function of those two, and checking it against the four molecules computed here, would turn heavy molecules are safer into an expression a spectroscopist could evaluate before deciding whether a substitution structure is worth attempting.

And there is a third thing left open. The corrections computed here are for the four molecules with force fields, and the force fields here are six molecules because fitting one needs measured frequencies for the molecule and at least one isotopologue. Formaldehyde has both in the literature, so the field done without here is a straightforward fit away — and it is the molecule the whole substitution analysis was about.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

  • The axis that goes the other way — both name convention, degeneracy, harmonic approximation, isotope substitution, model limit, moment of inertia, normal mode, vibrational modes, zero-point energy
  • How much of a band is a bond stretch — both name convention, degeneracy, force constant, harmonic approximation, least-squares, model limit, normal mode, vibrational modes
  • The forty-five that are fixed — both name convention, degeneracy, force constant, harmonic approximation, least-squares, model limit, normal mode, vibrational modes
  • More coordinates than motions — both name convention, degeneracy, force constant, least-squares, model limit, normal mode, vibrational modes
  • Ten directions no frequency can see — both name convention, degeneracy, force constant, least-squares, model limit, normal mode, vibrational modes
  • A label that prices nothing — both name convention, degeneracy, model limit, normal mode, vibrational modes

Named objects

A dashed tag is an object no other essay names yet.

ConventionDegeneracyForce constantHarmonic approximationIsotope substitutionLeast-squaresModel limitMoment of inertiaNormal modeRotational constantVibrational modesZero-point energy