set.seed(542)
repetitions <- 10000
sample_sizes <- c(20, 200)
ma1_coefficient <- 0.7
tau_squared <- (1 + ma1_coefficient)^2
# Generate one path with n observations; n is the sample size T in R.
simulate_standardized_mean <- function(n) {
z <- rexp(n + 1) - 1
x <- z[2:(n + 1)] + ma1_coefficient * z[1:n]
sqrt(n) * mean(x) / sqrt(tau_squared)
}
# Each replication contributes one standardized mean to the sampling distribution.
standardized_means <- lapply(
sample_sizes,
function(n) replicate(repetitions, simulate_standardized_mean(n))
)
# Compare the simulated sampling densities with the standard normal density.
par(mfrow = c(1, 2), mar = c(3.5, 4, 3, 1), mgp = c(2.2, 0.7, 0), bg = "gray92")
for (j in seq_along(sample_sizes)) {
values <- standardized_means[[j]]
plot(stats::density(values), lwd = 2,
xlim = c(-4, 4), ylim = c(0, 0.5),
xlab = "Standardized sample mean", main = paste("T =", sample_sizes[j]))
curve(dnorm(x), add = TRUE, col = "#00539B", lwd = 2)
abline(v = 0, lty = 3)
}Workbook 7 — From Dependence to a Central Limit Theorem
STA 542 · Introduction to Time Series Analysis
\[\def\E{\mathbb{E}} \def\P{\mathbb{P}} \def\R{\mathbb{R}} \def\Z{\mathbb{Z}} \def\N{\mathbb{N}} \def\cH{\mathcal{H}} \def\Var{\operatorname{Var}} \def\Cov{\operatorname{Cov}} \def\Corr{\operatorname{Corr}}\]
A ratio-consistent estimate of the sample-mean variance makes a standard error computable from one path. The resulting feasible Chebyshev interval uses a multiplier of about \(4.47\) for an asymptotic coverage bound of at least 95%. In an i.i.d. setting, the CLT gives a multiplier of \(1.96\) for the same coverage. Under what additional conditions can we expect a Gaussian limit for our time series average? The question concerns the sampling distribution of the mean relative to its uncertainty scale.
A Central Limit Theorem Under Finite Dependence
What changes if sufficiently separated groups of observations depend on entirely separate innovations?
Example 7.2 (A moving average with non-Gaussian innovations). Consider independent exponential random variables with rate one, and form the process \[ E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1), \qquad Z_t:=E_t-1, \qquad X_t:=Z_t+0.7Z_{t-1},\qquad t\in\Z. \] The innovations \(Z_t\) are right-skewed, with mean zero and variance one. Applying the same rule to every shifted pair of i.i.d. innovations makes \((X_t)\) strictly stationary. Its covariance calculation gives \[ \begin{aligned} \gamma(0)&=1+0.7^2=1.49,\\ \gamma(1)&=0.7,\qquad \gamma(h)=0\quad(|h|>1),\\ \tau^2&=1.49+2(0.7)=2.89. \end{aligned} \]
For any integer \(s\), observations through time \(s\) use innovations with indices at most \(s\). Observations from time \(s+2\) onward use innovations with indices at least \(s+1\). These two groups of observations are independent because the innovation groups are disjoint. Neighboring observations still share an innovation. \(\Diamond\)
The independence property in Example 7.2 extends to processes with a longer gap between the two groups.
Definition 7.3 (\(m\)-dependence). Let \(m\) be a nonnegative integer. A process \((X_t)\) is \(m\)-dependent if, for every integer \(s\), the collections \[ (\ldots,X_{s-1},X_s) \quad\text{and}\quad (X_{s+m+1},X_{s+m+2},\ldots) \] are independent.
An \(m\)-dependent weakly stationary process has \(\gamma(h)=0\) for \(|h|>m\). The converse need not hold: zero covariance is weaker than independence.
Theorem 7.4 (Central limit theorem for \(m\)-dependent processes). Suppose \((X_t)\) is strictly stationary and \(m\)-dependent, \(\E[X_t]=\mu\), \(\E[(X_t-\mu)^2]<\infty\), and \[ \tau^2=\sum_{h=-m}^{m}\gamma(h)>0. \] Then \[ \frac{\sqrt T(\bar X_T-\mu)}{\tau} \xrightarrow{d}N(0,1). \]
Example 7.2 is \(1\)-dependent and has \(\tau=\sqrt{2.89}=1.7\). The theorem therefore gives \[ \frac{\sqrt T\,\bar X_T}{1.7}\xrightarrow{d}N(0,1). \] This describes the distribution of the standardized mean across independent realizations of a long path. The innovations themselves remain right-skewed.
The reason is concrete. Divide the path into long blocks and discard the \(m\) observations separating adjacent blocks. The retained block sums are independent. If the blocks grow while remaining short relative to \(T\), the discarded observations are negligible and an independent-sum CLT applies.
First suppose \(|X_t-\mu|\le M\). Choose block length \(b_T=\lfloor T^{1/3}\rfloor\). Partition the observations into retained blocks of length \(b_T\) separated by gaps of length \(m\). By \(m\)-dependence, the retained block sums are independent.
The total number of discarded observations is \(O(T/b_T)\), and the variance of their centered sum is also \(O(T/b_T)\). After division by \(\sqrt T\), this part has variance \(O(1/b_T)\to0\). Each retained block is bounded by \(Mb_T=o(\sqrt T)\), so no block can dominate the normalized total. The independent triangular-array CLT applies to the retained sums, whose variance divided by \(T\) converges to \(\tau^2\).
For unbounded variables with a finite second moment, a standard truncation argument first replaces \(X_t-\mu\) by a bounded version, applies the preceding argument, and then lets the truncation level increase. We take that technical step on faith.
A Central Limit Theorem for Causal Linear Processes
A stationary AR(1) with positive innovation variance and \(0<|\phi|<1\) has nonzero covariances at arbitrarily large lags, so it is not \(m\)-dependent for any finite \(m\). Its causal representation is an infinite weighted sum of past innovations. We can also obtain a CLT for such sums when the absolute values of their coefficients have a finite sum.
Write \(\psi_j\) for the coefficient multiplying the innovation \(j\) time steps in the past. For the AR(1), \(\psi_j=\phi^j\). The following result allows a general sequence of these coefficients.
Theorem 7.5 (Causal linear-process CLT). Suppose \[ X_t=\mu+\sum_{j=0}^{\infty}\psi_jZ_{t-j}, \] where \((Z_t)\) is i.i.d. with mean zero and variance \(\sigma^2<\infty\), and \[ \sum_{j=0}^{\infty}|\psi_j|<\infty. \] Then \[ \sqrt T(\bar X_T-\mu) \xrightarrow{d}N(0,\tau^2), \qquad \tau^2=\sigma^2 \left(\sum_{j=0}^{\infty}\psi_j\right)^2. \]
We take Theorem 7.5 on faith. Its approximation argument first keeps only finitely many terms of the infinite weighted sum, giving a finite moving average. Absolute summability of the coefficients controls the contribution of the omitted innovations, allowing the finite-dependence result to extend to the infinite sum.
The same condition gives \[ \sum_{h\in\Z}\gamma(h) =\sigma^2\left(\sum_{j=0}^{\infty}\psi_j\right)^2. \] Thus the variance in the CLT is the long-run variance from Definition 6.1. If it is zero, the stated limit is a point mass at zero. A standard normal studentization requires the additional condition \(\tau^2>0\).
For the AR(1), \(\sum_{j=0}^{\infty}|\phi|^j<\infty\) and \(\sum_{j=0}^{\infty}\phi^j=1/(1-\phi)\). Theorem 7.5 therefore gives a Gaussian limit with variance \(\sigma^2/(1-\phi)^2\), agreeing with the covariance-sum calculation for the AR(1).
Oracle Gaussian Intervals
Suppose a CLT holds with \(\tau^2>0\). If the population value \(\tau\) were known, then \[ \frac{\sqrt T(\bar X_T-\mu)}{\tau} \xrightarrow{d}N(0,1). \]
Let \(z_p\) denote the \(p\)-quantile of \(N(0,1)\).
Proposition 7.6 (Oracle Gaussian interval). Under the preceding CLT, for \(0<\alpha<1\), the interval \[ C_T^{\mathrm{Gauss}} :=\left[ \bar X_T-z_{1-\alpha/2}\frac{\tau}{\sqrt T}, \ \bar X_T+z_{1-\alpha/2}\frac{\tau}{\sqrt T} \right] \] satisfies \[ \P\{\mu\in C_T^{\mathrm{Gauss}}\}\longrightarrow1-\alpha. \]
Proof. The coverage event is equivalent to \[ \left| \frac{\sqrt T(\bar X_T-\mu)}{\tau} \right|\le z_{1-\alpha/2}. \] The CLT and continuity of the standard normal CDF give the result. \(\square\)
If the exact variance is known, Chebyshev’s inequality gives the finite-sample interval \[ \bar X_T\ \pm\ \sqrt{\frac{\Var(\bar X_T)}{\alpha}}. \] For coverage of at least 95%, its multiplier is \(1/\sqrt{0.05}\approx4.47\). The oracle Gaussian interval uses the asymptotic multiplier \(1.96\). The smaller multiplier requires a CLT. Estimating the variance alone does not justify it.
The Moving-Average CLT in Simulation
We return to the moving average in Example 7.2, for which \(X_t=Z_t+0.7Z_{t-1}\) and \(Z_t=E_t-1\) with i.i.d. rate-one exponential \(E_t\). Its mean is zero and its long-run variance is \(2.89\). To examine the CLT, we simulate many independent paths of a given length and record one standardized sample mean from each path. We then repeat the experiment with longer paths.
| T | mean | sd | skewness | q025 | q975 | Gaussian_coverage | Chebyshev_coverage |
|---|---|---|---|---|---|---|---|
| 20 | -0.004 | 0.998 | 0.444 | -1.772 | 2.136 | 0.954 | 1 |
| 200 | -0.023 | 1.003 | 0.127 | -1.951 | 2.013 | 0.948 | 1 |
At \(T=20\), the sampling distribution remains visibly right-skewed. At \(T=200\), its shape is much closer to standard normal. The CLT describes this change in shape; Proposition 6.2 describes convergence of its variance. These are distinct conclusions.
From an Oracle CLT to a Feasible Interval
The preceding CLTs supply a limiting Gaussian shape. Known-range estimation or Bartlett consistency supplies an estimated scale that converges to the required positive value. Slutsky’s theorem combines these two conclusions; it does not require the estimate and the sample mean to be independent.
Proposition 7.7 (Studentized inference for the mean). Suppose \[ \sqrt T(\bar X_T-\mu)\xrightarrow{d}N(0,\tau^2), \qquad \tau^2>0, \] and \(\widehat\tau_T^2\to\tau^2\) in probability, with \(\widehat\tau_T^2\ge0\), and let \(0<\alpha<1\). Set \(\tau:=\sqrt{\tau^2}\), \(\widehat\tau_T:=\sqrt{\widehat\tau_T^2}\), and define \(S_T\) by \[ S_T:= \begin{cases} \sqrt T(\bar X_T-\mu)/\widehat\tau_T, &\widehat\tau_T>0,\\ 0,&\widehat\tau_T=0. \end{cases} \] Then \[ S_T \xrightarrow{d}N(0,1). \] If \(z_p\) denotes the \(p\) quantile of \(N(0,1)\), the interval \[ C_T^{\mathrm{feasible}} :=\left[ \bar X_T-z_{1-\alpha/2}\frac{\widehat\tau_T}{\sqrt T}, \ \bar X_T+z_{1-\alpha/2}\frac{\widehat\tau_T}{\sqrt T} \right] \] satisfies \[ \P\{\mu\in C_T^{\mathrm{feasible}}\} \longrightarrow1-\alpha. \]
Proof. The CLT gives \(\sqrt T(\bar X_T-\mu)/\tau\to N(0,1)\) in distribution. Consistency and \(\tau^2>0\) give \(\widehat\tau_T/\tau\to1\) in probability and \(\P(\widehat\tau_T=0)\to0\). Slutsky’s theorem gives the studentized limit; the arbitrary value on the vanishing zero event has no effect. Outside that zero event, coverage is equivalent to \(|S_T|\le z_{1-\alpha/2}\). The symmetric difference between these two events is therefore contained in \(\{\widehat\tau_T=0\}\), whose probability tends to zero. The studentized limit now gives the coverage statement. \(\square\)
The interval is feasible because every quantity in it can be computed from the observed path. It is not exact: both the Gaussian shape and the estimated standard error are asymptotic approximations.
The conclusion belongs to the sample mean. A new estimator changes the centered statistic and usually its normalization, so it needs its own CLT. Consistency of that estimator, or consistency of a variance estimate, does not supply the missing sampling shape.
The uncertainty-estimation workbook already gave the feasible Chebyshev interval \(\bar X_T\pm\sqrt{\widehat v_T/\alpha}\) under ratio consistency. When \(\widehat v_T=\widehat\tau_T^2/T\), the interval in Proposition 7.7 uses the same standard error with multiplier \(z_{1-\alpha/2}\). At 95%, the two multipliers are approximately \(4.47\) and \(1.96\). The CLT is the additional reason for using the smaller one.
Exercises
Exercise 7.1 (Seeing the CLT in a non-Gaussian MA(2)). Let \[ E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1), \qquad Z_t:=E_t-1, \qquad X_t:=\mu+Z_t+0.5Z_{t-1}+0.25Z_{t-2}. \]
Method and report: Complete parts (a)–(c) by hand. Use R for part (d), with the stated seed and simulation sizes. Report a table and interpret the change in sampling shape.
Show that \((X_t)\) is strictly stationary and \(2\)-dependent. Derive \(\gamma(0)\), \(\gamma(1)\), \(\gamma(2)\), and \(\tau^2\).
State the CLT for \(\bar X_T\) under this law. Explain why neither Gaussian innovations nor an estimated ACVF is needed for this conclusion.
For \(T>2\), derive the exact value of \(T\Var(\bar X_T)\). Write the finite-sample Chebyshev interval with coverage at least 95% and the oracle asymptotic 95% Gaussian interval for \(\mu\).
Set \(\mu=0\). With seed
542, simulate 5,000 paths at each \(T\in\{25,250\}\). For \[ S_T:=\frac{\sqrt T\,\bar X_T}{\tau}, \] first state the standard-normal benchmarks for its mean, standard deviation, skewness, and central 95% range. Then report the simulated values and the coverage of the two intervals in a table. Plot the density of the 5,000 standardized sample means with the standard normal density overlaid at each sample size. Each replication must generate an independent path.Interpret which features approach their limiting values and why the Gaussian interval is not exact at either sample size. Explain why this experiment studies the distribution of sample means rather than the distribution of individual observations, and why the simulation illustrates the CLT without proving it.
# Student starter code
set.seed(542)
sample_sizes <- c(25, 250)
repetitions <- 5000
tau_squared <- 3.0625 # Population value from part (a), held fixed across paths.
# Generate one full path and return its standardized sample mean.
simulate_standardized_mean <- function(n) {
z <- rexp(n + 2) - 1
x <- z[3:(n + 2)] + 0.5 * z[2:(n + 1)] + 0.25 * z[1:n]
sqrt(n) * mean(x) / sqrt(tau_squared)
}Exercise 7.2 (Two multipliers for one estimated standard error). Consider again \[ X_t=\mu+Z_t+0.5Z_{t-1}+0.25Z_{t-2}, \qquad Z_t=E_t-1, \qquad E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1). \] Exercise 6.2 used the positive-part estimate \[ \widehat q_T:=\max\left\{ \widehat\gamma_T(0) +2\left(1-\frac1T\right)\widehat\gamma_T(1) +2\left(1-\frac2T\right)\widehat\gamma_T(2),\ 0\right\}, \qquad \widehat v_{2,T}:=\frac{\widehat q_T}{T}, \] where each sample autocovariance uses divisor \(T\). The population quantities are \(\tau^2=49/16\) and \(q_T:=T\Var(\bar X_T)=49/16-9/(4T)\) for \(T>2\).
Method and report. Complete parts (a)–(c) by hand. Use R for part (d). Report sampling summaries and interval coverage with Monte Carlo standard errors, then interpret the two approximations in part (e).
Explain why \(\widehat q_T\xrightarrow{P}\tau^2\) and \(\widehat v_{2,T}/\Var(\bar X_T)\xrightarrow{P}1\) under this law. Identify which result concerns covariance learning and which concerns the limiting scale. What fails in this argument for an MA(2) whose coefficient sum \(1+\theta_1+\theta_2\) is zero?
State the CLT and use Proposition 7.7 to justify the interval \[ C_T^{G}:=\bar X_T\ \pm\ z_{0.975}\sqrt{\widehat v_{2,T}}. \] Write its asymptotic coverage event. Explain why covariance range two, without the independent-innovation construction, would not supply this CLT.
Compare \(C_T^G\) with the feasible Chebyshev interval \[ C_T^{C}:=\bar X_T\ \pm\ \sqrt{\widehat v_{2,T}/0.05}. \] State the coverage conclusion available for each. Derive the ratio of their full widths whenever \(\widehat v_{2,T}>0\). Which extra assumption permits the smaller multiplier?
Set \(\mu=0\). Using seed
542, simulate 5,000 independent paths at \(T=100\), followed by 5,000 at \(T=500\), resetting the seed only once before the two experiments. For each path, compute \(\widehat q_T\) and \[ S_T:= \begin{cases} \sqrt T\bar X_T/\sqrt{\widehat q_T},&\widehat q_T>0,\\ 0,&\widehat q_T=0. \end{cases} \] Report the mean, standard deviation, and central 95% range of \(S_T\), and the proportion of zero estimates. Plot its sampling density with the standard normal density overlaid. On the same paths, report coverage and mean full width for \(C_T^G\), \(C_T^C\), and the oracle Gaussian comparison \(\bar X_T\pm z_{0.975}\sqrt{q_T/T}\). For every coverage fraction \(\hat p\), report the estimated Monte Carlo standard error \(\sqrt{\hat p(1-\hat p)/5000}\). A zero estimated variance gives a zero-width interval; retain it when checking coverage.Interpret the difference between oracle and feasible Gaussian coverage. Explain why neither Gaussian interval is exact for these innovations, why the estimated Chebyshev interval lacks the oracle finite-sample guarantee, and why the value assigned to \(S_T\) at zero cannot be used to count interval coverage.
From a Stationary Model to an Observed Record
We can estimate the uncertainty of a stationary process mean and, under a separate CLT, justify Gaussian calibration. Both arguments assume one common mean and covariance structure across the record. Trends, seasonal patterns, or changing spread can challenge that premise. The filtering workbook examines adjustments and the structural assumptions under which their output could be modeled as stationary.
Sources
The dependent central limit theorems follow the treatments in A. W. van der Vaart, Time Series (Sections 4.1–4.2), and P. J. Brockwell and R. A. Davis, Introduction to Time Series and Forecasting. Slutsky’s theorem supplies the studentization step once the CLT and scale consistency have been established.