
---
title: "Workbook 7 — From Dependence to a Central Limit Theorem"
subtitle: "STA 542 · Introduction to Time Series Analysis"
format:
  html:
    embed-resources: true
    toc: true
    toc-depth: 2
    number-sections: false
    code-copy: true
    fig-width: 9
    fig-height: 3.4
    fig-dpi: 150
execute:
  warning: false
  message: false
editor:
  markdown:
    wrap: 72
---

::: hidden
$$\def\E{\mathbb{E}}
\def\P{\mathbb{P}}
\def\R{\mathbb{R}}
\def\Z{\mathbb{Z}}
\def\N{\mathbb{N}}
\def\cH{\mathcal{H}}
\def\Var{\operatorname{Var}}
\def\Cov{\operatorname{Cov}}
\def\Corr{\operatorname{Corr}}$$
:::

```{=html}
<style>
.exercise { border-left: 3px solid #00539B; padding: 0.7em 1.1em; background: #f6f8fa; margin: 1.4em 0; border-radius: 0 4px 4px 0; }
.exercise p:first-child { margin-top: 0; }
figure figcaption { color: #555; font-size: 0.9em; }
</style>
```

```{r}
#| label: setup
#| include: false
set.seed(542)
```

A ratio-consistent estimate of the sample-mean variance makes a standard
error computable from one path. The resulting feasible Chebyshev interval
uses a multiplier of about $4.47$ for an asymptotic coverage bound of at
least 95%. In an i.i.d. setting, the CLT gives a multiplier of $1.96$ for the same coverage. Under what additional conditions can we expect a Gaussian limit for our time series average? The question concerns the sampling distribution of the mean relative to its uncertainty scale.
<!-- 
## Scale Is Not Shape

Proposition 6.2 identifies the scale of the sample mean. It does not say that
the standardized sample mean is approximately Gaussian. A process can have
an absolutely summable ACVF and still retain non-Gaussian information that
averaging does not remove.

**Example 7.1 (A finite long-run variance without a Gaussian limit).** Let
$\Lambda$ equal $1$ or $2$, each with probability $1/2$, and let $(R_t)$ be
i.i.d. with
$$
\P(R_t=1)=\P(R_t=-1)=\frac12,
$$
independently of $\Lambda$. Define $X_t=\Lambda R_t$. Then
$$
\gamma(0)=\E[\Lambda^2]=\frac52,\qquad
\gamma(h)=0\quad(h\ne0),
$$
so $\tau^2=5/2$. Conditional on $\Lambda=\lambda$, the ordinary i.i.d. CLT
gives
$$
\sqrt T\,\bar X_T\xrightarrow{d}N(0,\lambda^2).
$$
Averaging over the two possible values of $\Lambda$ gives the unconditional
limit
$$
\frac12N(0,1)+\frac12N(0,4),
$$
not $N(0,5/2)$. The correct variance scale does not supply a Gaussian
sampling shape. $\Diamond$ -->

## A Central Limit Theorem Under Finite Dependence

What changes if sufficiently separated groups of
observations depend on entirely separate innovations?

**Example 7.2 (A moving average with non-Gaussian innovations).** Consider
independent exponential random variables with rate one, and form the process
$$
E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1),
\qquad Z_t:=E_t-1,
\qquad X_t:=Z_t+0.7Z_{t-1},\qquad t\in\Z.
$$
The innovations $Z_t$ are right-skewed, with mean zero and variance one.
Applying the same rule to every shifted pair of i.i.d. innovations makes
$(X_t)$ strictly stationary. Its covariance calculation gives
$$
\begin{aligned}
\gamma(0)&=1+0.7^2=1.49,\\
\gamma(1)&=0.7,\qquad \gamma(h)=0\quad(|h|>1),\\
\tau^2&=1.49+2(0.7)=2.89.
\end{aligned}
$$

For any integer $s$, observations through time $s$ use innovations with
indices at most $s$. Observations from time $s+2$ onward use innovations
with indices at least $s+1$. These two groups of observations are independent
because the innovation groups are disjoint. Neighboring observations still
share an innovation. $\Diamond$

The independence property in Example 7.2 extends to processes with a longer
gap between the two groups.

**Definition 7.3 ($m$-dependence).** Let $m$ be a nonnegative integer.
A process $(X_t)$ is
*$m$-dependent* if, for every integer $s$, the collections
$$
(\ldots,X_{s-1},X_s)
\quad\text{and}\quad
(X_{s+m+1},X_{s+m+2},\ldots)
$$
are independent.

An $m$-dependent weakly stationary process has $\gamma(h)=0$ for
$|h|>m$. The converse need not hold: zero covariance is weaker than
independence.

**Theorem 7.4 (Central limit theorem for $m$-dependent processes).** Suppose
$(X_t)$ is strictly stationary and $m$-dependent,
$\E[X_t]=\mu$, $\E[(X_t-\mu)^2]<\infty$, and
$$
\tau^2=\sum_{h=-m}^{m}\gamma(h)>0.
$$
Then
$$
\frac{\sqrt T(\bar X_T-\mu)}{\tau}
\xrightarrow{d}N(0,1).
$$

Example 7.2 is $1$-dependent and has $\tau=\sqrt{2.89}=1.7$. The theorem
therefore gives
$$
\frac{\sqrt T\,\bar X_T}{1.7}\xrightarrow{d}N(0,1).
$$
This describes the distribution of the standardized mean across independent
realizations of a long path. The innovations themselves remain right-skewed.

The reason is concrete. Divide the path into long blocks and discard the
$m$ observations separating adjacent blocks. The retained block sums are
independent. If the blocks grow while remaining short relative to $T$, the
discarded observations are negligible and an independent-sum CLT applies.

::: {.callout-note collapse="true"}
## More Details ♠: The Blocking Argument

First suppose $|X_t-\mu|\le M$. Choose block length
$b_T=\lfloor T^{1/3}\rfloor$. Partition the observations into retained
blocks of length $b_T$ separated by gaps of length $m$. By
$m$-dependence, the retained block sums are independent.

The total number of discarded observations is $O(T/b_T)$, and the variance
of their centered sum is also $O(T/b_T)$. After division by $\sqrt T$, this
part has variance $O(1/b_T)\to0$. Each retained block is bounded by
$Mb_T=o(\sqrt T)$, so no block can dominate the normalized total. The
independent triangular-array CLT applies to the retained sums, whose variance
divided by $T$ converges to $\tau^2$.

For unbounded variables with a finite second moment, a standard truncation
argument first replaces $X_t-\mu$ by a bounded version, applies the preceding
argument, and then lets the truncation level increase. We take that technical
step on faith.
:::

## A Central Limit Theorem for Causal Linear Processes

A stationary AR(1) with positive innovation variance and $0<|\phi|<1$ has
nonzero covariances at arbitrarily large lags, so it is not $m$-dependent for
any finite $m$. Its causal representation is an infinite weighted sum of
past innovations. We can also
obtain a CLT for such sums when the absolute values of their coefficients
have a finite sum.

Write $\psi_j$ for the coefficient multiplying the innovation $j$ time steps
in the past. For the AR(1), $\psi_j=\phi^j$. The following result allows a
general sequence of these coefficients.

**Theorem 7.5 (Causal linear-process CLT).** Suppose
$$
X_t=\mu+\sum_{j=0}^{\infty}\psi_jZ_{t-j},
$$
where $(Z_t)$ is i.i.d. with mean zero and variance
$\sigma^2<\infty$, and
$$
\sum_{j=0}^{\infty}|\psi_j|<\infty.
$$
Then
$$
\sqrt T(\bar X_T-\mu)
\xrightarrow{d}N(0,\tau^2),
\qquad
\tau^2=\sigma^2
\left(\sum_{j=0}^{\infty}\psi_j\right)^2.
$$

We take Theorem 7.5 on faith. Its approximation argument first keeps only
finitely many terms of the infinite weighted sum, giving a finite moving
average. Absolute summability of the coefficients controls the contribution
of the omitted innovations, allowing the finite-dependence result to extend
to the infinite sum.

The same condition gives
$$
\sum_{h\in\Z}\gamma(h)
=\sigma^2\left(\sum_{j=0}^{\infty}\psi_j\right)^2.
$$
Thus the variance in the CLT is the long-run variance from Definition 6.1.
If it is zero, the stated limit is a point mass at zero. A standard normal
studentization requires the additional condition $\tau^2>0$.

For the AR(1), $\sum_{j=0}^{\infty}|\phi|^j<\infty$ and
$\sum_{j=0}^{\infty}\phi^j=1/(1-\phi)$. Theorem 7.5 therefore gives a
Gaussian limit with variance $\sigma^2/(1-\phi)^2$, agreeing with the
covariance-sum calculation for the AR(1).

## Oracle Gaussian Intervals

Suppose a CLT holds with $\tau^2>0$. If the population value $\tau$ were
known, then
$$
\frac{\sqrt T(\bar X_T-\mu)}{\tau}
\xrightarrow{d}N(0,1).
$$

Let $z_p$ denote the $p$-quantile of $N(0,1)$.

**Proposition 7.6 (Oracle Gaussian interval).** Under the preceding CLT,
for $0<\alpha<1$, the interval
$$
C_T^{\mathrm{Gauss}}
:=\left[
\bar X_T-z_{1-\alpha/2}\frac{\tau}{\sqrt T},
\ \bar X_T+z_{1-\alpha/2}\frac{\tau}{\sqrt T}
\right]
$$
satisfies
$$
\P\{\mu\in C_T^{\mathrm{Gauss}}\}\longrightarrow1-\alpha.
$$

*Proof.* The coverage event is equivalent to
$$
\left|
\frac{\sqrt T(\bar X_T-\mu)}{\tau}
\right|\le z_{1-\alpha/2}.
$$
The CLT and continuity of the standard normal CDF give the result.
$\square$

If the exact variance is known, Chebyshev's inequality gives the
finite-sample interval
$$
\bar X_T\ \pm\
\sqrt{\frac{\Var(\bar X_T)}{\alpha}}.
$$
For coverage of at least 95%, its multiplier is
$1/\sqrt{0.05}\approx4.47$. The oracle Gaussian interval uses the
asymptotic multiplier $1.96$. The smaller multiplier requires a CLT.
Estimating the variance alone does not justify it.

## The Moving-Average CLT in Simulation

We return to the moving average in Example 7.2, for which
$X_t=Z_t+0.7Z_{t-1}$ and $Z_t=E_t-1$ with i.i.d. rate-one exponential
$E_t$. Its mean is zero and its long-run variance is $2.89$. To examine
the CLT, we simulate many independent paths of a given length and record
one standardized sample mean from each path. We then repeat the experiment
with longer paths.

```{r}
#| label: fig-wb7-nongaussian-clt
#| fig-cap: "Sampling distributions of the standardized mean for a non-Gaussian MA(1). The standard normal density is overlaid in blue. Simulated data."
set.seed(542)
repetitions <- 10000
sample_sizes <- c(20, 200)
ma1_coefficient <- 0.7
tau_squared <- (1 + ma1_coefficient)^2

# Generate one path with n observations; n is the sample size T in R.
simulate_standardized_mean <- function(n) {
  z <- rexp(n + 1) - 1
  x <- z[2:(n + 1)] + ma1_coefficient * z[1:n]
  sqrt(n) * mean(x) / sqrt(tau_squared)
}

# Each replication contributes one standardized mean to the sampling distribution.
standardized_means <- lapply(
  sample_sizes,
  function(n) replicate(repetitions, simulate_standardized_mean(n))
)

# Compare the simulated sampling densities with the standard normal density.
par(mfrow = c(1, 2), mar = c(3.5, 4, 3, 1), mgp = c(2.2, 0.7, 0), bg = "gray92")
for (j in seq_along(sample_sizes)) {
  values <- standardized_means[[j]]
  plot(stats::density(values), lwd = 2,
       xlim = c(-4, 4), ylim = c(0, 0.5),
       xlab = "Standardized sample mean", main = paste("T =", sample_sizes[j]))
  curve(dnorm(x), add = TRUE, col = "#00539B", lwd = 2)
  abline(v = 0, lty = 3)
}
```

```{r}
#| label: tbl-wb7-nongaussian-clt
#| echo: false
summarize_standardized_means <- function(values, n) {
  exact_scaled_variance <- tau_squared - 2 * ma1_coefficient / n
  data.frame(
    T = n,
    mean = mean(values),
    sd = sd(values),
    skewness = mean((values - mean(values))^3) / sd(values)^3,
    q025 = unname(quantile(values, 0.025)),
    q975 = unname(quantile(values, 0.975)),
    Gaussian_coverage = mean(abs(values) <= 1.96),
    Chebyshev_coverage = mean(
      abs(values) <= sqrt(exact_scaled_variance / (0.05 * tau_squared))
    )
  )
}

simulation_summary <- do.call(
  rbind,
  Map(summarize_standardized_means, standardized_means, sample_sizes)
)
knitr::kable(simulation_summary, digits = 3)
```

At $T=20$, the sampling distribution remains visibly right-skewed. At
$T=200$, its shape is much closer to standard normal. The CLT describes this
change in shape; Proposition 6.2 describes convergence of its variance. These
are distinct conclusions.

## From an Oracle CLT to a Feasible Interval

The preceding CLTs supply a limiting Gaussian shape. Known-range
estimation or Bartlett consistency supplies an estimated scale that converges
to the required positive value. Slutsky's theorem combines these two
conclusions; it does not require the estimate and the sample mean to be
independent.

**Proposition 7.7 (Studentized inference for the mean).** Suppose
$$
\sqrt T(\bar X_T-\mu)\xrightarrow{d}N(0,\tau^2),
\qquad \tau^2>0,
$$
and $\widehat\tau_T^2\to\tau^2$ in probability, with
$\widehat\tau_T^2\ge0$, and let $0<\alpha<1$. Set
$\tau:=\sqrt{\tau^2}$,
$\widehat\tau_T:=\sqrt{\widehat\tau_T^2}$, and define $S_T$ by
$$
S_T:=
\begin{cases}
\sqrt T(\bar X_T-\mu)/\widehat\tau_T,
&\widehat\tau_T>0,\\
0,&\widehat\tau_T=0.
\end{cases}
$$
Then
$$
S_T
\xrightarrow{d}N(0,1).
$$
If $z_p$ denotes the $p$ quantile of $N(0,1)$, the interval
$$
C_T^{\mathrm{feasible}}
:=\left[
\bar X_T-z_{1-\alpha/2}\frac{\widehat\tau_T}{\sqrt T},
\ \bar X_T+z_{1-\alpha/2}\frac{\widehat\tau_T}{\sqrt T}
\right]
$$
satisfies
$$
\P\{\mu\in C_T^{\mathrm{feasible}}\}
\longrightarrow1-\alpha.
$$

*Proof.* The CLT gives
$\sqrt T(\bar X_T-\mu)/\tau\to N(0,1)$ in distribution. Consistency and
$\tau^2>0$ give $\widehat\tau_T/\tau\to1$ in probability and
$\P(\widehat\tau_T=0)\to0$. Slutsky's theorem gives the studentized limit;
the arbitrary value on the vanishing zero event has no effect. Outside that
zero event, coverage is equivalent to
$|S_T|\le z_{1-\alpha/2}$. The symmetric difference between these two events
is therefore contained in $\{\widehat\tau_T=0\}$, whose probability tends to
zero. The studentized limit now gives the coverage statement. $\square$

The interval is feasible because every quantity in it can be computed from
the observed path. It is not exact: both the Gaussian shape and the estimated
standard error are asymptotic approximations.

The conclusion belongs to the sample mean. A new estimator changes the
centered statistic and usually its normalization, so it needs its own CLT.
Consistency of that estimator, or consistency of a variance estimate, does
not supply the missing sampling shape.

The uncertainty-estimation workbook already gave the feasible Chebyshev
interval $\bar X_T\pm\sqrt{\widehat v_T/\alpha}$ under ratio consistency.
When $\widehat v_T=\widehat\tau_T^2/T$, the interval in Proposition 7.7
uses the same standard error with multiplier $z_{1-\alpha/2}$.
At 95%, the two multipliers are approximately $4.47$ and $1.96$.
The CLT is the additional reason for using the smaller one.

## Exercises

::: {.exercise}
**Exercise 7.1 (Seeing the CLT in a non-Gaussian MA(2)).** Let
$$
E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1),
\qquad Z_t:=E_t-1,
\qquad
X_t:=\mu+Z_t+0.5Z_{t-1}+0.25Z_{t-2}.
$$

*Method and report:* Complete parts (a)--(c) by hand. Use R for part (d),
with the stated seed and simulation sizes. Report a table and interpret the
change in sampling shape.

(a) Show that $(X_t)$ is strictly stationary and $2$-dependent. Derive
$\gamma(0)$, $\gamma(1)$, $\gamma(2)$, and $\tau^2$.

(b) State the CLT for $\bar X_T$ under this law. Explain why neither
Gaussian innovations nor an estimated ACVF is needed for this conclusion.

(c) For $T>2$, derive the exact value of $T\Var(\bar X_T)$. Write the
finite-sample Chebyshev interval with coverage at least 95% and the oracle
asymptotic 95% Gaussian interval for $\mu$.

(d) Set $\mu=0$. With seed `542`, simulate 5,000 paths at each
$T\in\{25,250\}$. For
$$
S_T:=\frac{\sqrt T\,\bar X_T}{\tau},
$$
first state the standard-normal benchmarks for its mean, standard deviation,
skewness, and central 95% range. Then report the simulated values and the
coverage of the two intervals in a table. Plot the density of the 5,000
standardized sample means with the standard normal density overlaid at each
sample size. Each replication must generate an independent path.

(e) Interpret which features approach their limiting values and why the
Gaussian interval is not exact at either sample size. Explain why this
experiment studies the distribution of sample means rather than the
distribution of individual observations, and why the simulation illustrates
the CLT without proving it.

```{r}
#| eval: false
# Student starter code
set.seed(542)
sample_sizes <- c(25, 250)
repetitions <- 5000
tau_squared <- 3.0625  # Population value from part (a), held fixed across paths.

# Generate one full path and return its standardized sample mean.
simulate_standardized_mean <- function(n) {
  z <- rexp(n + 2) - 1
  x <- z[3:(n + 2)] + 0.5 * z[2:(n + 1)] + 0.25 * z[1:n]
  sqrt(n) * mean(x) / sqrt(tau_squared)
}
```
:::

::: {.exercise}
**Exercise 7.2 (Two multipliers for one estimated standard error).** Consider
again
$$
X_t=\mu+Z_t+0.5Z_{t-1}+0.25Z_{t-2},
\qquad Z_t=E_t-1,
\qquad E_t\overset{\text{iid}}{\sim}\operatorname{Exponential}(1).
$$
Exercise 6.2 used the positive-part estimate
$$
\widehat q_T:=\max\left\{
\widehat\gamma_T(0)
+2\left(1-\frac1T\right)\widehat\gamma_T(1)
+2\left(1-\frac2T\right)\widehat\gamma_T(2),\ 0\right\},
\qquad
\widehat v_{2,T}:=\frac{\widehat q_T}{T},
$$
where each sample autocovariance uses divisor $T$. The population quantities
are $\tau^2=49/16$ and
$q_T:=T\Var(\bar X_T)=49/16-9/(4T)$ for $T>2$.

*Method and report.* Complete parts (a)--(c) by hand. Use R for part (d).
Report sampling summaries and interval coverage with Monte Carlo standard
errors, then interpret the two approximations in part (e).

(a) Explain why $\widehat q_T\xrightarrow{P}\tau^2$ and
$\widehat v_{2,T}/\Var(\bar X_T)\xrightarrow{P}1$ under this law.
Identify which result concerns covariance learning and which concerns the
limiting scale. What fails in this argument for an MA(2) whose coefficient
sum $1+\theta_1+\theta_2$ is zero?

(b) State the CLT and use Proposition 7.7 to justify the interval
$$
C_T^{G}:=\bar X_T\ \pm\ z_{0.975}\sqrt{\widehat v_{2,T}}.
$$
Write its asymptotic coverage event. Explain why covariance range two,
without the independent-innovation construction, would not supply this CLT.

(c) Compare $C_T^G$ with the feasible Chebyshev interval
$$
C_T^{C}:=\bar X_T\ \pm\ \sqrt{\widehat v_{2,T}/0.05}.
$$
State the coverage conclusion available for each. Derive the ratio of their
full widths whenever $\widehat v_{2,T}>0$. Which extra assumption permits
the smaller multiplier?

(d) Set $\mu=0$. Using seed `542`, simulate 5,000 independent paths at
$T=100$, followed by 5,000 at $T=500$, resetting the seed only once before
the two experiments. For each path, compute $\widehat q_T$ and
$$
S_T:=
\begin{cases}
\sqrt T\bar X_T/\sqrt{\widehat q_T},&\widehat q_T>0,\\
0,&\widehat q_T=0.
\end{cases}
$$
Report the mean, standard deviation, and central 95% range of $S_T$, and
the proportion of zero estimates. Plot its sampling density with the
standard normal density overlaid. On the same paths, report coverage and
mean full width for $C_T^G$, $C_T^C$, and the oracle Gaussian comparison
$\bar X_T\pm z_{0.975}\sqrt{q_T/T}$. For every coverage fraction $\hat p$,
report the estimated Monte Carlo standard error
$\sqrt{\hat p(1-\hat p)/5000}$.
A zero estimated variance gives a zero-width interval; retain it when
checking coverage.

(e) Interpret the difference between oracle and feasible Gaussian coverage.
Explain why neither Gaussian interval is exact for these innovations,
why the estimated Chebyshev interval lacks the oracle finite-sample
guarantee, and why the value assigned to $S_T$ at zero cannot be used to
count interval coverage.
:::

## From a Stationary Model to an Observed Record

We can estimate the uncertainty of a stationary process mean and, under a
separate CLT, justify Gaussian calibration. Both arguments assume one
common mean and covariance structure across the record. Trends, seasonal
patterns, or changing spread can challenge that premise. The filtering
workbook examines adjustments and the structural assumptions under which
their output could be modeled as stationary.

## Sources

The dependent central limit theorems follow the treatments in A. W. van der
Vaart, *Time Series* (Sections 4.1--4.2), and P. J. Brockwell and R. A. Davis,
*Introduction to Time Series and Forecasting*. Slutsky's theorem supplies the
studentization step once the CLT and scale consistency have been established.
