When the molecule does not stop

A count rather than an average

Two couplings quoted from one distribution sit a hundred and seventy decades apart. Neither is a summary of it. The quantity that decides how much of a spectrum near a box level is resonant pairs is a count of pairs above a threshold — and it has a closed form, which is a density times a reach times the logarithm of ten.

Worth reading first: A band that is a hundred and seventy decades of nothing · One defect is a level, many are a band.

Two couplings between defect runs in a disordered chain make the point. The coupling at the closest separation a chain can contain is 0.0877β for runs of four; the coupling at the typical separation is 10⁻¹⁹ for runs of four and 10⁻¹⁷⁵ for runs of eight. The conclusion drawn from them — that a width computed from the second is not small but below anything that could be called a quantity — stands.

What does not stand is the idea that two numbers describe what is between them. They are both taken from one distribution, and the distribution is the thing that decides how much of a spectrum near a box level is resonant pairs and how much is isolated levels. That question is not an average. It is a count.

The count, for five run lengths. The fraction of runs whose nearest neighbour of the same length is coupled more strongly than a threshold, against how many decades below the closest possible coupling that threshold sits. Every curve is a straight line over this range, because the fraction is small and the geometric tail is linear in the separation there. The slope is what the next figure is about: it is the whole content of the distribution, and it is a product of two numbers.
Fig. 1 The fraction of runs whose nearest neighbour of the same length is coupled more strongly than a threshold, for five run lengths.

The count has a closed form

A site is low with probability xx independently, so a run of exactly LL low sites needs LL lows bounded by two highs and occurs at a rate ρ=xL(1x)2\rho = x^L(1-x)^2 per site. Consecutive runs of the same length are therefore separated by a geometric number of sites, and the probability that a run’s nearest same-length neighbour is within ss of it is

P(gs)=1(1ρ)sP(g \le s) = 1 - (1-\rho)^s

The coupling falls exponentially in the separation, t(g)=t0eg/ξt(g) = t_0 e^{-g/\xi}, so a threshold τ\tau is a separation s(τ)=ξln(t0/τ)s(\tau) = \xi\ln(t_0/\tau), and the fraction of runs coupled more strongly than τ\tau is the geometric tail evaluated there.

Both halves are measured rather than assumed.

How many pairs are coupled more strongly than a stated amount. Runs of 6 low sites at half concentration. For each threshold — stated as decades below the coupling of the closest pair a chain can contain — the separation at which the coupling falls to it, the fraction of runs whose neighbour is closer than that from the closed form, and the fraction generated chains actually have. The closed form is a geometric tail and it is measured rather than assumed: the two agree within 1.4 percentage points across seven decades.
Fig. 2 For each threshold: the separation it corresponds to, the fraction the closed form predicts, and the fraction generated chains actually have.

The closed form and three chains of two hundred thousand sites agree within 1.4 percentage points at every threshold from two decades below the closest coupling to fourteen — a range over which the separation runs from 9.5 sites to 64.9 and the fraction from 3.7 per cent to 22.4. The run density itself agrees with xL(1x)2x^L(1-x)^2 to better than a per cent.

The fit that supplies ξ\xi and t0t_0 is the same one used to reach those two numbers, so nothing here rests on a second model of the coupling. What is new is the second factor: the separations, counted in chains rather than summarised by their mean.

That agreement is the licence for everything that follows. Every quantity below is computed from ρ\rho, ξ\xi and t0t_0 without generating a chain, and the chains are what say the computation describes them.

Where the two quoted numbers actually sit

With the distribution in hand, the two couplings can be located in it.

Where the two quoted numbers sit in the distribution. The usual comparison quotes the coupling at the closest separation a chain contains and the coupling at the typical separation, and finds them a hundred and seventy decades apart. Neither is a summary. The closest is the extreme of the distribution — the fraction of runs at or above it is 0.125 per cent — and the typical is a coupling almost no pair has, because the separation distribution is geometric and its mean sits far out in a tail that most of its weight is nowhere near.
Fig. 3 The closest pair, the mean separation and the median separation, with the coupling at each and the share of runs at least that strongly coupled.

The closest pair is the extreme. Only 0.125 per cent of runs of six have a neighbour one site away, so the coupling quoted as the band’s ceiling is a coupling one run in eight hundred has.

The typical separation is worse, and it is worse in the way a mean always is for a geometric distribution. Its mean is 250 sites; its median is 177. So a coupling quoted at the mean separation is smaller than the coupling 95.8 per cent of pairs have. It is not a typical coupling; it is a coupling almost nothing has, sitting far out in a tail that most of the distribution’s weight is nowhere near.

That is not a criticism of the arithmetic behind the two numbers, which is right, and it is not a criticism of the conclusion, which the numbers here confirm: the couplings at any separation past a few tens of sites are below anything a spectrum could resolve. It is a statement about what “typical” can be made to mean when a distribution is skewed by an exponential. A mean separation put through an exponential is not the mean of anything, and the hundred and seventy decades between them is a distance between two points of a distribution rather than a width of it.

There is a third quantity worth naming while the distribution is on the page, because it is the one a spectroscopist would ask for and it had not been computed. The median coupling — the coupling half the pairs exceed — is 1.57×10401.57\times10^{-40} for runs of six. That is between the two quoted and is no more measurable than either, so the correction being made here is not that the conclusion was wrong. It is that the conclusion followed from a fact about the whole distribution and was argued from two points of it, and the two points happened to bracket the answer.

How much of a band is a pair

Now the question the count was built for. A pair is resonant when its coupling exceeds the mean spacing of the levels around it — the criterion used to locate the onset of a band, applied to every pair rather than to the closest one. The spacing falls as 1/N1/N, so the answer moves with the chain length.

How much of a defect band is actually a pair. The share of runs sitting in a pair whose coupling exceeds the mean level spacing — the usual criterion for a band, applied to every pair rather than to the closest one. It rises with the chain length because the spacing falls, nearly logarithmically over this range: runs of 6 go from 1.26 per cent at a thousand sites to 16.1 at a million million. Most of a defect band is isolated levels at every length anything could have.
Fig. 4 The share of runs sitting in a resonant pair, against the number of sites in the chain, for five run lengths.

For runs of six it is 1.26 per cent at a thousand sites, 4.76 per cent at a hundred thousand, 9.78 per cent at a hundred million, and 16.1 per cent at a million million. For runs of eight it is 5.19 per cent at a million million. For runs of four — the most numerous — it reaches 41.4 per cent there, and that is the largest number in the whole table.

Most of a defect band is isolated levels at every chain length anything could have. The computed onset is the length at which the closest pair becomes resonant, which is where a band can be said to begin; it is nowhere near the length at which most of the states are in one.

And over this range the rise is nearly logarithmic. Every curve is close to a straight line against the logarithm of the chain length, over nine decades. It is not exactly one: each decade adds the linear amount times the share of runs not yet resonant, so the curves bend over and the share saturates — slowly at this concentration, where the slope of runs of six falls only from 1.74 points a decade to 1.57 across the range shown. Half of the runs of six are in resonant pairs at a chain of about 104110^{41} sites; a straight line through the range would have said 103110^{31}.

Two numbers multiplied

The near-straight lines are what make the whole thing predictable, because their initial slope has a closed form.

For small ρs\rho s the geometric tail is 1(1ρ)sρs1 - (1-\rho)^s \approx \rho s, and the separation grows as ξln(t0/spacing)\xi\ln(t_0/\text{spacing}) with the spacing falling as 1/N1/N. So the share rises by ρξln10\rho\xi\ln 10 per decade of chain length — a density times a reach times a constant, with nothing fitted in it.

The whole distribution is two numbers multiplied. The share of runs in a resonant pair rises by a fixed amount per decade of chain length, and that amount has a closed form: the run density times the coupling's decay length times the logarithm of ten. Measured against predicted for five run lengths, over a factor of ten thousand in the answer. The agreement improves as the density falls, which is the linearisation of the geometric tail being what it says it is.
Fig. 5 The predicted slope beside the measured one, for five run lengths, across a factor of ten thousand in the answer.

Predicted 5.172, 3.093, 1.802, 1.027 and 0.575 percentage points per decade; measured 4.518, 2.878, 1.735, 1.007 and 0.569. The ratio runs 0.874, 0.930, 0.963, 0.981, 0.990 — approaching one from below as the density falls, which is exactly what the linearisation says it should do, since ρs\rho s is smallest where ρ\rho is.

So the whole content of the distribution, for this purpose, is two numbers multiplied. That is worth having for a reason beyond economy: it says what a change to the model would do. Doubling the concentration multiplies ρ\rho by 2L2^L and the share by the same factor; a softer barrier lengthens ξ\xi and multiplies the share by the same ratio. Neither requires a chain.

The two factors run against each other

The two numbers do not move together as the runs get longer, and the direction of the second is the opposite of the natural guess.

The decay length of a pair's coupling, against the length of the runs. The second of the two numbers, measured for five run lengths. It is a straight line in the run length with a slope of 0.2801 sites per site — so a longer defect couples further, and the two factors in the product run against each other: the density falls by a factor of two per extra site and the reach rises by about a seventh. The density wins, which is why the longer runs are the more isolated ones.
Fig. 6 The decay length of a pair’s coupling, against the number of low sites in a run.

The decay length rises from 1.4376 sites for runs of four to 2.5563 for runs of eight — a straight line with a slope of 0.2801 sites per site. A longer defect couples further, which follows from its state sitting closer to the top of the barrier and therefore decaying more slowly through it — the same relation between a level’s height and its reach that a single defect level is built on.

The density falls by a factor of two per extra low site, at half concentration. So the product ρξ\rho\xi falls by a factor of about 1.8 per extra site rather than 2, and the longer runs are the more isolated ones because they are rarer, not because their coupling is shorter-ranged. The earlier reading attributed the isolation to the coupling; the coupling is the factor working the other way.

The slope of 0.2801 is itself worth a note. Over the five lengths measured it is straight to within a hundredth of a site, and a straight line in LL for a decay length through a barrier is what a barrier of fixed height gives when the state’s energy rises linearly with the box level — which it does not, exactly, since 2cos(π/(L+1))2\cos(\pi/(L+1)) is not linear in LL. Over four to eight it is nearly so, and that is the whole of the explanation available here.

What was computed, and how

The couplings are exact. A chain of two hundred sites is built with two runs of LL low sites separated by a stated number of high ones, the tridiagonal matrix is diagonalised, the two eigenvalues nearest the isolated box level 2cos(π/(L+1))2\cos(\pi/(L+1)) are found, and their difference is the splitting. That is repeated for separations from one to twenty, and ξ\xi and t0t_0 come from a least-squares fit on logarithms over separations of two to fourteen — the short end excluded because one site apart is not in the asymptotic regime, the long end because the splitting there is at the floor of a double.

The separations are counted rather than modelled. Three chains of two hundred thousand sites are generated with the same site-energy generator the rest of this collection uses, every run of exactly LL is found, and the gaps between consecutive ones are collected — 2,354 pairs for runs of six, 9,387 for runs of four, 582 for runs of eight.

The level spacing is the one used for the band’s onset, unchanged: the gap states are those of runs from three to nine, they span the energy between the box levels of those two lengths, and their mean spacing is that range divided by their number.

Two refusals hold the formula honest at its ends. A threshold far below every coupling in the distribution must catch all of the runs — the fraction must go to one, not to something near it — and a threshold above the closest pair’s coupling must catch almost none. A formula that returned a number outside the unit interval anywhere would be an extrapolation with a percentage sign on it.

The chains are checked against the closed form for the density as well as for the tail, because a run census that miscounted boundaries would produce a density wrong by a factor and a tail that still looked geometric.

Where the model stops

Everything here treats a pair as two runs and nothing else. A real chain has a third run somewhere, and when three runs are mutually within a few decay lengths the splitting is not a two-level problem. At the densities here that is rare — the probability of two neighbours both within twenty sites is a few parts in ten thousand for runs of six — but it is not zero, and it is the correction that grows as the concentration rises. The same three-body case is what makes an arrangement problem hard rather than merely large.

The exponential law is fitted over separations of two to fourteen and used out to sixty-five. That extrapolation is doing real work in the tail thresholds, and it is the weakest step here. What licenses it is that the couplings are computed rather than assumed at every separation up to twenty, and the fit’s residuals over that range are what a straight line in a logarithm looks like; what it does not license is any statement about a separation of hundreds, where the coupling is below the precision of the diagonalisation that would measure it.

Nothing here is a spectrum. A share of runs in a resonant pair is not a share of spectral weight, because a resonant pair contributes two levels where two isolated runs also contribute two. What the share decides is how many of the levels near a box level are split rather than degenerate, which is what a measurement of a defect band’s width is looking at, and that is the quantity the two couplings were meant to bound.

And the resonance criterion is a comparison of a coupling with a mean spacing, which is the standard criterion and is a rule of thumb rather than a theorem. A pair whose coupling is a tenth of the spacing is not sharply different from one whose coupling is ten times it.

The generalisation

A distribution reached through an exponential has no typical value. The mean separation here is 250 sites and the median is 177, which is a modest difference; put both through eg/ξe^{-g/\xi} and the coupling at the mean is a factor of 101610^{16} below the coupling at the median. Every quantity of the form the typical value of a function of a random variable is exposed to this, and the repair is never a better average — it is a count above a threshold, which is a question the distribution can answer.

The same shape turns up on a rate ratio whose ceiling comes from a Boltzmann factor, where the same exponential turns a modest spread in one quantity into a large one in another, and on a decay length fitted over a window, where the quantity being averaged is itself not constant.

And a logarithmic approach is an approach that does not arrive. The share of resonant pairs rises by a fixed number of points per decade, which reads like progress towards a limit and is not: nine decades of chain length take runs of six from one per cent to sixteen, and the next nine would take them to thirty-one. Any statement of the form in a large enough system has to be checked against how large, and the answer here is a number no system has.

Who found it, and when

The geometric distribution of gaps between rare independent events is elementary. Its use for defect states in a disordered chain is the standard picture behind Mott’s argument for variable-range hopping, where the same competition appears — a coupling falling exponentially in a distance whose distribution has a long tail — and where the resolution is the same one: optimise a product rather than average either factor.

What is done here is smaller and more specific. The two quoted couplings are placed in their own distribution, and the count that replaces them is given a closed form and checked against generated chains. That the answer is ρξln10\rho\xi\ln 10 per decade is not a discovery about disordered chains; it is the arithmetic of a geometric tail, written down for disordered chains, so that what follows has a formula rather than a pair of extremes.

Still open: the same count on a square net

The obvious open question is the dimension. Everything here rests on the separation distribution of runs along a line, which is geometric because a chain is one-dimensional. On a square net of low and high sites the typical separation between defects of a given kind goes as the reciprocal square root of their density, and the same competition between a falling coupling and a growing count has a different balance — the count of neighbours within a distance grows as the square of it rather than linearly. Those nets are already built here, and the resonant share there should rise faster than logarithmically.

The nearer question is the concentration. Everything above is at x=12x = \tfrac12, where ρ=2L/4\rho = 2^{-L}/4 and the two factors happen to be within a factor of two of each other for runs of six. Sweeping xx moves ρ\rho by orders of magnitude while leaving ξ\xi alone, so the product ρξ\rho\xi can be scanned across four decades with one number changing — which would test the closed form where it is least comfortable, at densities high enough that ρs\rho s is no longer small and the linearisation fails.

What links here

Computed from the collection rather than written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays naming at least two of the same things, that neither author linked.

Named objects

A dashed tag is an object no other essay names yet.

ApproximationBands in a solidClosed formConvergenceDefect stateDisorderExact diagonalisationLocalisationModel limitReference stateThermodynamic limitTight-binding models