
---
title: "Workbook 6 — Estimating Uncertainty from One Path"
subtitle: "STA 542 · Introduction to Time Series Analysis"
format:
  html:
    embed-resources: true
    toc: true
    toc-depth: 2
    number-sections: false
    code-copy: true
    fig-width: 9
    fig-height: 3.4
    fig-dpi: 150
execute:
  warning: false
  message: false
editor:
  markdown:
    wrap: 72
---

::: hidden
$$\def\E{\mathbb{E}}
\def\P{\mathbb{P}}
\def\R{\mathbb{R}}
\def\Z{\mathbb{Z}}
\def\N{\mathbb{N}}
\def\cH{\mathcal{H}}
\def\Var{\operatorname{Var}}
\def\Cov{\operatorname{Cov}}
\def\Corr{\operatorname{Corr}}$$
:::

```{=html}
<style>
.exercise { border-left: 3px solid #00539B; padding: 0.7em 1.1em; background: #f6f8fa; margin: 1.4em 0; border-radius: 0 4px 4px 0; }
.exercise p:first-child { margin-top: 0; }
figure figcaption { color: #555; font-size: 0.9em; }
</style>
```

```{r}
#| label: setup
#| include: false
set.seed(542)
```

Workbook 5 ended with a difficulty. At each fixed lag, one path may teach us
the corresponding autocovariance. The variance of the sample mean, however,
contains a sum whose number of lags grows with the sample size. Before trying
to estimate that sum, we ask what population quantity it approaches and how one path can estimate it.

## The Long-Run Variance

For a weakly stationary process with mean $\mu$ and ACVF $\gamma$,
Workbook 5 established
$$
T\Var(\bar X_T)
=\gamma(0)+2\sum_{h=1}^{T-1}
\left(1-\frac hT\right)\gamma(h).
$$
For independent observations, multiplication by $T$ gives the fixed value
$\gamma(0)$. Under dependence, we ask when the weighted covariance sum on the
right also approaches a finite limit.

**Definition 6.1 (Long-run variance).** If
$$
\sum_{h\in\Z}|\gamma(h)|<\infty,
$$
the *long-run variance* of $(X_t)$ is
$$
\tau^2:=\sum_{h\in\Z}\gamma(h)
=\gamma(0)+2\sum_{h=1}^{\infty}\gamma(h).
$$
When $\tau^2>0$, we write $\tau:=\sqrt{\tau^2}$ for its positive square
root.

**Proposition 6.2 (The limiting variance scale).** Under the condition in
Definition 6.1,
$$
T\Var(\bar X_T)\longrightarrow\tau^2.
$$
In particular, $\tau^2\ge0$.

*Proof.* At each fixed lag, the weight $1-h/T$ approaches one. To control the
whole sum, we separate a fixed number of lags from the remaining tail.
Absolute summability lets us make that tail small before increasing $T$.

Let $\varepsilon>0$. Choose an integer $K$ large enough that
$\sum_{h>K}|\gamma(h)|<\varepsilon/4$, and then keep $K$ fixed. For $T>K$,
subtracting the definition of $\tau^2$ from the exact variance identity gives
$$
T\Var(\bar X_T)-\tau^2
=-\frac2T\sum_{h=1}^{T-1}h\gamma(h)
  -2\sum_{h=T}^{\infty}\gamma(h).
$$
Split the first sum at $K$. Since $h/T\le1$ for every included lag,
$$
\begin{aligned}
\left|T\Var(\bar X_T)-\tau^2\right|
&\le\frac2T\sum_{h=1}^{K}h|\gamma(h)|
 +2\sum_{h=K+1}^{T-1}\frac hT|\gamma(h)|
 +2\sum_{h=T}^{\infty}|\gamma(h)|\\
&\le\underbrace{\frac2T\sum_{h=1}^{K}h|\gamma(h)|}_{\text{finitely many lags}}
 +\underbrace{2\sum_{h>K}|\gamma(h)|}_{\text{remaining tail}}.
\end{aligned}
$$

The tail term is less than $\varepsilon/2$ by our choice of $K$. The finite
sum in the first term is now a fixed number, so that term is also less than
$\varepsilon/2$ for all sufficiently large $T$. The total error is then less
than $\varepsilon$, proving the limit. Since $T\Var(\bar X_T)\ge0$ for
every $T$, its limit $\tau^2$ is nonnegative. $\square$

When $\tau^2>0$,
$$
\Var(\bar X_T)\approx\frac{\tau^2}{T}.
$$
The marginal variance $\gamma(0)$ is replaced by the sum of all
autocovariances. This is the first effect of dependence on inference for the
mean.

A zero long-run variance is possible. For
$X_t=Z_t-Z_{t-1}$ with i.i.d. centered innovations of variance
$\sigma^2>0$, the only nonzero covariances are
$\gamma(0)=2\sigma^2$ and $\gamma(1)=-\sigma^2$. Hence
$\tau^2=\gamma(0)+2\gamma(1)=0$.

Summation cancels every interior innovation:
$$
\bar X_T=\frac{Z_T-Z_0}{T},\qquad
\Var(\bar X_T)=\frac{2\sigma^2}{T^2}.
$$
Thus $T\Var(\bar X_T)\to0$, as Proposition 6.2 requires. The exact
variance is still positive for each $T$, but it has order $T^{-2}$.
The estimation results below use $\tau^2>0$ when converting a consistent
long-run-variance estimate into an estimate of the mean's variance.

## Accuracy on the Right Scale

Write
$$
v_T:=\Var(\bar X_T).
$$
Under absolute covariance summability, $v_T\to0$. The useless estimator
$\widetilde v_T:=0$ therefore satisfies
$|\widetilde v_T-v_T|\to0$. Absolute consistency does not ensure that an
estimated standard error has the right size.

**Definition 6.3 (Ratio-consistent variance estimate).** Suppose $v_T>0$ for
all sufficiently large $T$. A nonnegative estimator $\widehat v_T$ is
*ratio-consistent* for $v_T$ if
$$
\frac{\widehat v_T}{v_T}\xrightarrow{P}1.
$$

Ratio consistency gives
$\sqrt{\widehat v_T}/\sqrt{v_T}\to1$ in probability. Thus the estimated
standard error has the correct size relative to the unknown oracle standard
error. The zero estimator fails this comparison completely.

Under absolute covariance summability, Proposition 6.2 gives
$$
q_T:=T\Var(\bar X_T)\longrightarrow
\tau^2=\sum_{h\in\Z}\gamma(h).
$$
If $\tau^2>0$ and a nonnegative estimator satisfies
$\widehat\tau_T^2\to\tau^2$ in probability, then
$$
\widehat v_T:=\frac{\widehat\tau_T^2}{T},
\qquad
\frac{\widehat v_T}{v_T}
=\frac{\widehat\tau_T^2}{q_T}
\xrightarrow{P}1.
$$
We can therefore estimate the stable covariance sum rather than the vanishing
variance directly.

## Why Not Use Every Lag?

Workbook 5 defined the divisor-$T$ sample ACVF by
$$
\widehat\gamma_T(h)
:=\frac1T\sum_{t=1}^{T-h}
(X_t-\bar X_T)(X_{t+h}-\bar X_T),
\qquad 0\le h<T.
$$
The most literal estimate of $q_T$ substitutes every available sample
autocovariance into its exact formula:
$$
\widehat q_T^{\mathrm{all}}
:=\widehat\gamma_T(0)
+2\sum_{h=1}^{T-1}
\left(1-\frac hT\right)\widehat\gamma_T(h).
$$

**Example 6.4 (Every available lag is too many).** Suppose
$$
X_t\overset{\mathrm{iid}}{\sim}N(\mu,\sigma^2),
\qquad \sigma^2>0.
$$
The target is $q_T=\sigma^2$. Every fixed-lag sample autocovariance is
consistent, yet
$$
\E[\widehat q_T^{\mathrm{all}}]
=\sigma^2\frac{T^2-1}{3T^2}
\longrightarrow\frac{\sigma^2}{3}.
$$
The small bias at each positive lag accumulates over order $T$ estimated
lags. The calculation uses only independence and a finite variance; normality
merely gives the simulation below a fully specified law.

```{r}
#| label: fig-wb6-all-lag-failure
#| code-fold: true
#| code-summary: "R code"
#| fig-height: 4.0
#| fig-cap: "Monte Carlo means and central 80% ranges for two estimates of the scaled sample-mean variance under i.i.d. N(0,1) sampling. The dashed line is the true value 1; the dotted line is the all-lag expectation limit 1/3. Based on 2,500 simulated paths at each sample size. Simulated data."
estimate_iid_scales <- function(n) {
  x <- rnorm(n)
  e <- x - mean(x)
  c(
    `lag zero` = mean(e^2),
    `all lags` = 2 / n^2 * sum(cumsum(e)[seq_len(n - 1)]^2)
  )
}

set.seed(542)
all_lag_sizes <- c(50, 200, 800)
all_lag_repetitions <- 2500
all_lag_results <- lapply(all_lag_sizes, function(n) {
  t(replicate(all_lag_repetitions, estimate_iid_scales(n)))
})

all_lag_summary <- do.call(rbind, lapply(seq_along(all_lag_sizes), function(i) {
  values <- all_lag_results[[i]]
  data.frame(
    n = all_lag_sizes[i],
    method = colnames(values),
    center = colMeans(values),
    lower = apply(values, 2, quantile, probs = 0.10),
    upper = apply(values, 2, quantile, probs = 0.90)
  )
}))

method_colors <- c("lag zero" = "#0072B2", "all lags" = "#D55E00")
method_offsets <- c("lag zero" = -0.07, "all lags" = 0.07)
par(mfrow = c(1, 1), mar = c(4, 4, 2.5, 1), bg = "gray92")
plot(1:3, rep(1, 3), type = "n", ylim = c(0, 1.25), xaxt = "n",
     xlab = "Sample size T", ylab = "Estimated scaled variance",
     main = "Using every lag fails under independence")
axis(1, at = 1:3, labels = all_lag_sizes)
abline(h = 1, lty = 2)
abline(h = 1 / 3, lty = 3)
for (method in names(method_colors)) {
  rows <- all_lag_summary$method == method
  x_positions <- 1:3 + method_offsets[method]
  arrows(x_positions, all_lag_summary$lower[rows],
         x_positions, all_lag_summary$upper[rows],
         angle = 90, code = 3, length = 0.04,
         col = method_colors[method])
  points(x_positions, all_lag_summary$center[rows], pch = 16,
         col = method_colors[method])
}
legend("right", names(method_colors), col = method_colors,
       pch = 16, bty = "n")
```

As $T$ grows, the lag-zero estimate tightens around the true value $1$. The
all-lag estimate remains centered near $1/3$ and does not tighten around the
target. More sample autocovariances have produced a worse variance estimate.

::: {.callout-note collapse="true"}
## More Details ♠: An Elementary Proof of the All-Lag Failure

Let $e_t:=X_t-\bar X_T$. For i.i.d. observations with finite variance,
$$
\E[e_se_t]
=\sigma^2\left\{\mathbf 1(s=t)-\frac1T\right\}.
$$
Consequently,
$$
\E[\widehat\gamma_T(0)]
=\sigma^2\left(1-\frac1T\right),
\qquad
\E[\widehat\gamma_T(h)]
=-\sigma^2\frac{T-h}{T^2},
\quad h\ge1.
$$

Substitution into the all-lag estimator gives
$$
\begin{aligned}
\E[\widehat q_T^{\mathrm{all}}]
&=\sigma^2\left(1-\frac1T\right)
-\frac{2\sigma^2}{T}
\sum_{h=1}^{T-1}\left(1-\frac hT\right)^2\\
&=\sigma^2\frac{T^2-1}{3T^2}
\longrightarrow\frac{\sigma^2}{3}.
\end{aligned}
$$
Each fixed-lag bias vanishes, but there are order $T$ such biases.

The same estimator has the identity
$$
\widehat q_T^{\mathrm{all}}
=\frac{2}{T^2}\sum_{k=1}^{T-1}
\left(\sum_{t=1}^k e_t\right)^2
\ge0.
$$
It follows by expanding both sides and using
$\sum_{t=1}^T e_t=0$.

Markov's inequality now gives
$$
\P\left(\widehat q_T^{\mathrm{all}}>\frac{\sigma^2}{2}\right)
\le
\frac{2\E[\widehat q_T^{\mathrm{all}}]}{\sigma^2}
\longrightarrow\frac23.
$$
If $\widehat q_T^{\mathrm{all}}$ converged in probability to $\sigma^2$, the
probability on the left would converge to one. The estimator is genuinely
inconsistent; no additional limit theorem is needed.
:::

## A Known Covariance Range

Suppose a known integer $m\ge0$ satisfies
$$
\gamma(h)=0\qquad(|h|>m).
$$
This is a covariance cutoff, not an independence claim. Only the fixed lags
$0,\ldots,m$ enter the exact sample-mean variance, so Workbook 5's fixed-lag
learning result can now do all the required estimation.

For $T>m$, define
$$
\widehat q_{m,T}
:=\widehat\gamma_T(0)
+2\sum_{h=1}^{m}
\left(1-\frac hT\right)\widehat\gamma_T(h),
\qquad
\widehat v_{m,T}:=\frac{\max(\widehat q_{m,T},0)}{T}.
$$
The positive part makes the variance estimate nonnegative. It has no
asymptotic effect when the long-run variance is positive.

**Proposition 6.5 (Known-range variance estimation).** Suppose $(X_t)$ is
weakly stationary, its ACVF is learnable through the known lag $m$, and
$\gamma(h)=0$ for $|h|>m$. If
$$
\tau^2=\gamma(0)+2\sum_{h=1}^{m}\gamma(h)>0,
$$
then
$$
\frac{\widehat v_{m,T}}{\Var(\bar X_T)}\xrightarrow{P}1.
$$

*Proof.* ACVF learnability through $m$ gives
$\widehat\gamma_T(h)\to\gamma(h)$ in probability for each
$h=0,\ldots,m$. Because the sum has a fixed number of terms,
$$
\widehat q_{m,T}\xrightarrow{P}\tau^2.
$$
The covariance cutoff also gives
$q_T=T\Var(\bar X_T)\to\tau^2$. Positivity of the limit makes the positive
part asymptotically irrelevant, and therefore
$$
\frac{\widehat v_{m,T}}{\Var(\bar X_T)}
=\frac{\max(\widehat q_{m,T},0)}{q_T}
\xrightarrow{P}1.
$$
$\square$

Proposition 6.5 is complete only because $m$ is fixed and known. A sample ACF
cannot certify that a population covariance is exactly zero beyond a selected
lag. Real series therefore return us to the growing-sum problem.

## A Real Series Has No Known Cutoff

**Example 6.6 (DJIA returns and squared returns).** The object
`astsa::djia` contains 2,518 daily closing observations of the Dow Jones
Industrial Average, in index points, from April 20, 2006, through April 20,
2016. Let $c_t$ denote an observed closing level and define the observed log
return
$$
r_t:=\log c_t-\log c_{t-1}.
$$
We subtract the sample mean from the returns before comparing the sample ACF
of the adjusted returns with the sample ACF of their squares.

The notation `astsa::djia` accesses the data without attaching the package.
The package must still be installed once; if needed, run
`install.packages("astsa")` before rendering the workbook.

```{r}
#| label: fig-wb6-djia-acfs
#| code-fold: true
#| code-summary: "R code"
#| fig-height: 3.8
#| fig-cap: "Sample ACFs through lag 10 for DJIA daily drift-adjusted log returns and their squares, April 20, 2006–April 20, 2016. The squared returns have larger positive sample ACF values at every displayed lag. Real data: `astsa::djia`."
djia_data <- astsa::djia
djia_closes <- as.numeric(djia_data[, "Close"])
djia_returns <- diff(log(djia_closes))
djia_returns <- djia_returns - mean(djia_returns)

djia_return_acf <- as.numeric(
  stats::acf(djia_returns, lag.max = 10, plot = FALSE)$acf
)[-1]
djia_squared_acf <- as.numeric(
  stats::acf(djia_returns^2, lag.max = 10, plot = FALSE)$acf
)[-1]

plot_acf_values <- function(values, title) {
  plot(1:10, values, type = "h", lwd = 2, ylim = c(-0.45, 0.45),
       xlab = "Lag (trading days)", ylab = "Sample ACF", main = title)
  points(1:10, values, pch = 16)
  abline(h = 0)
}

par(mfrow = c(1, 2), mar = c(4, 4, 2.5, 1), bg = "white")
plot_acf_values(djia_return_acf, "Drift-adjusted returns")
plot_acf_values(djia_squared_acf, "Squared returns")
```

The signed-return sample ACF values range from
$`r sprintf("%.3f", min(djia_return_acf))`$ to
$`r sprintf("%.3f", max(djia_return_acf))`$. They describe the relative lag
pattern of signed returns; the corresponding sample autocovariances enter an
estimated covariance sum for the mean return. The squared-return sample ACF
values are positive at every displayed lag and range from
$`r sprintf("%.3f", min(djia_squared_acf))`$ to
$`r sprintf("%.3f", max(djia_squared_acf))`$. They concern a different
feature: persistence in the magnitude of returns.

The squared-return plot suggests volatility persistence, but those
correlations do not enter the long-run variance of the signed-return mean
directly. At a fixed lag $h$, the relevant product functions include
$g_h(x_0,\ldots,x_h)=x_0x_h$ for returns and
$g_h^{(2)}(x_0,\ldots,x_h)=x_0^2x_h^2$ for squared returns. A finite plot of
their realized sample averages cannot establish $g_h$-ergodicity, convergence
to population correlations, or a population covariance cutoff. Nor does it
identify a volatility model. The plots pose a question about dependence; they
do not supply the assumptions needed to learn it.

## Bartlett Regularization

The known-range estimator keeps a fixed number of lags. The all-lag estimator
keeps too many. For an unknown covariance range, we instead let the number of
included lags grow slowly and downweight the least reliable ones.

**Definition 6.7 (Bartlett estimator).** Choose an integer bandwidth
$0\le H<T$. The *Bartlett estimator* of the long-run variance is
$$
\widehat\tau_{H,T}^2
:=\widehat\gamma_T(0)
+2\sum_{h=1}^{H}
\left(1-\frac{h}{H+1}\right)\widehat\gamma_T(h).
$$
With the divisor-$T$ sample ACVF, these triangular weights make the estimate
nonnegative.

Two similar weights now have different jobs. The factor $1-h/T$ in the exact
variance formula counts the fraction of observed pairs at lag $h$. The factor
$1-h/(H+1)$ in Definition 6.7 is a deliberate taper chosen by the analyst. At
$H=T-1$, the two weights coincide and Definition 6.7 reproduces the failed
all-lag estimator from Example 6.4.

If $H$ stays fixed while the ACVF has a nonzero tail, the estimator omits that
tail forever. If $H$ grows too quickly, the accumulated estimation error in
the included sample autocovariances need not vanish. Consistency requires both
a controllable population tail and collective accuracy over the retained
lags.

**Proposition 6.8 (Bartlett consistency under collective accuracy).** Suppose
$(X_t)$ is weakly stationary,
$$
\sum_{h\in\Z}|\gamma(h)|<\infty,
$$
and let $(H_T)$ be a deterministic sequence of integer bandwidths satisfying
$H_T\to\infty$ and $H_T<T$. If
$$
\sum_{h=0}^{H_T}
|\widehat\gamma_T(h)-\gamma(h)|
\xrightarrow{P}0,
$$
then
$$
\widehat\tau_{H_T,T}^2\xrightarrow{P}\tau^2.
$$
If also $\tau^2>0$, then
$\widehat\tau_{H_T,T}^2/T$ is ratio-consistent for
$\Var(\bar X_T)$.

*Proof.* Put $w_{h,T}:=1-h/(H_T+1)$. The triangle inequality gives
$$
\begin{aligned}
|\widehat\tau_{H_T,T}^2-\tau^2|
\le{}&
|\widehat\gamma_T(0)-\gamma(0)|
+2\sum_{h=1}^{H_T}w_{h,T}
|\widehat\gamma_T(h)-\gamma(h)|\\
&+2\sum_{h=1}^{H_T}(1-w_{h,T})|\gamma(h)|
+2\sum_{h>H_T}|\gamma(h)|.
\end{aligned}
$$
The collective-accuracy condition sends the first line to zero in
probability. Absolute summability sends the omitted tail to zero.

For the taper term, fix an integer $K$. Once $H_T>K$,
$$
\begin{aligned}
2\sum_{h=1}^{H_T}(1-w_{h,T})|\gamma(h)|
&=2\sum_{h=1}^{H_T}\frac{h}{H_T+1}|\gamma(h)|\\
&\le\frac{2}{H_T+1}\sum_{h=1}^{K}h|\gamma(h)|
  +2\sum_{h>K}|\gamma(h)|.
\end{aligned}
$$
First choose $K$ to make the tail arbitrarily small, then increase $T$
to make the fixed finite sum divided by $H_T+1$ small. The taper term
therefore tends to zero. This proves long-run-variance consistency. The
ratio conclusion follows from Proposition 6.2 and $\tau^2>0$. $\square$

The collective condition is stronger than consistency at every fixed lag.
For example, it follows if
$$
\sup_{0\le h\le H_T}
\E[|\widehat\gamma_T(h)-\gamma(h)|]
=O(T^{-1/2})
$$
and $H_T=o(\sqrt T)$. This calculation is only a sufficient illustration;
different process classes permit different bandwidth ranges.

::: {.callout-note collapse="true"}
## More Details ♠: A Classical Short-Memory Route

Suppose the fixed process law has the causal linear representation
$$
X_t=\mu+\sum_{j=0}^{\infty}\psi_jZ_{t-j},
$$
where $(Z_t)$ is i.i.d. with mean zero, variance $\sigma^2>0$, and finite
fourth moment, and
$$
\sum_{j=0}^{\infty}|\psi_j|<\infty.
$$
For this short-memory class, the standard Bartlett result gives
$$
\widehat\tau_{H_T,T}^2\xrightarrow{P}\tau^2
$$
when $H_T\to\infty$ and $H_T/T\to0$. We take this model-specific result on
faith. Its proof controls the weighted collection of sample autocovariances
jointly; fixed-lag $g$-ergodicity alone does not provide that control.

The classical proof controls the weighted covariance sum directly. It need
not verify the stronger sum of absolute errors in Proposition 6.8, which is
why its bandwidth condition can be broader than the illustrative
$H_T=o(\sqrt T)$ rule.

Stable causal ARMA processes with i.i.d. innovations having a finite fourth
moment belong to this class. A causal representation by itself does not
establish Bartlett consistency for every nonlinear process,
so we do not treat the result as a generic consequence of causality.
:::

## A Conservative Interval with an Estimated Variance

Workbook 5's interval used the known variance $v_T=\Var(\bar X_T)$.
We now replace it by a nonnegative estimate. The replacement makes the
threshold random and dependent on the same path as the mean, so the
finite-sample Chebyshev calculation cannot simply be reused unchanged.
Ratio consistency supplies the additional argument.

**Proposition 6.9 (A feasible conservative interval).** Fix a process law
under which $\E[\bar X_T]=\mu$ and
$0<v_T:=\Var(\bar X_T)<\infty$ for all sufficiently large $T$.
Suppose $\widehat v_T\ge0$ and
$\widehat v_T/v_T\xrightarrow{P}1$. For a fixed $0<\alpha<1$, define
$$
C_T^{\mathrm{cons}}
:=\left[
\bar X_T-\sqrt{\frac{\widehat v_T}{\alpha}},
\ \bar X_T+\sqrt{\frac{\widehat v_T}{\alpha}}
\right].
$$
Then
$$
\liminf_{T\to\infty}\P\{\mu\in C_T^{\mathrm{cons}}\}\ge1-\alpha.
$$
No independence between $\bar X_T$ and $\widehat v_T$ is required.

*Proof.* Fix $0<\varepsilon<1$ and let
$$
A_T:=\{\widehat v_T\ge(1-\varepsilon)v_T\}.
$$
Ratio consistency gives $\P(A_T^c)\to0$. On $A_T$, failure of coverage
implies
$$
|\bar X_T-\mu|>
\sqrt{\frac{(1-\varepsilon)v_T}{\alpha}}.
$$

An event inclusion, followed by Chebyshev's inequality, gives
$$
\begin{aligned}
\P\{\mu\notin C_T^{\mathrm{cons}}\}
&\le\P(A_T^c)
 +\P\left\{|\bar X_T-\mu|>
        \sqrt{\frac{(1-\varepsilon)v_T}{\alpha}}\right\}\\
&\le\P(A_T^c)+\frac{\alpha}{1-\varepsilon}.
\end{aligned}
$$
Taking the limit superior yields a noncoverage bound
$\alpha/(1-\varepsilon)$. Letting $\varepsilon\downarrow0$ gives
$\limsup_T\P\{\mu\notin C_T^{\mathrm{cons}}\}\le\alpha$, which is
the stated coverage conclusion. $\square$

For a weakly stationary process, $\E[\bar X_T]=\mu$ automatically.
Proposition 6.5 supplies a suitable $\widehat v_T$ when the covariance
range is known. Under Proposition 6.8 with $\tau^2>0$, we may instead use
$\widehat v_T=\widehat\tau_{H_T,T}^2/T$. Both routes produce a computable
interval from one path.

The conclusion is a limiting lower bound at each fixed process law. It
neither promises coverage of at least $1-\alpha$ at every finite sample
size nor says that coverage converges to exactly $1-\alpha$. No Gaussian
approximation was used. A more precise account of the sampling shape may
permit a shorter interval; that is a separate question.

## Exercises

::: {.exercise}
**Exercise 6.1 (Dependence and effective sample size).** Let
$$
X_t=\mu+Z_t+\theta Z_{t-1},
$$
where $(Z_t)$ is i.i.d. with mean zero and variance $\sigma^2>0$.
When $\gamma(0)>0$ and $\tau^2>0$, define the *effective sample size
for the mean* by
$$
T_{\mathrm{eff}}:=T\frac{\gamma(0)}{\tau^2}.
$$
It is the number of independent observations with marginal variance
$\gamma(0)$ that would match the dependent mean's approximate variance
$\tau^2/T$. This definition concerns the sample mean, not every target.

*Method and report.* Derive the quantities by hand. Evaluate the requested
ratios and explain in complete sentences what the comparison measures.

(a) Derive $\gamma(0)$, $\gamma(1)$, and $\tau^2$.

(b) For $\theta\ne-1$, derive $T_{\mathrm{eff}}/T$. Evaluate it at
$\theta=0.8$ and $\theta=-0.5$, then interpret both values.

(c) Explain what changes at $\theta=-1$. Derive $\bar X_T-\mu$ directly and
its exact variance.

(d) Compare the MA(1) law above at $\theta=1/2$ and $\sigma^2=4/5$ with
a second law for $(X_t)$: the stationary causal AR(1)
$$
X_t-\mu=\frac25(X_{t-1}-\mu)+Z_t,
\qquad \E[Z_t]=0,\qquad \Var(Z_t)=\frac{21}{25},
$$
whose innovations are again i.i.d. Show that the two laws have the same
marginal variance and lag-one correlation. Calculate their long-run
variances and effective sample sizes. Explain why agreement at lag one does
not give equal uncertainty about the mean.
:::

::: {.exercise}
**Exercise 6.2 (Estimating uncertainty for an MA(2)).** Let
$$
E_t\overset{\mathrm{iid}}{\sim}\operatorname{Exponential}(1),
\qquad Z_t:=\sigma(E_t-1),\qquad \sigma>0,
$$
and define, for $t\in\Z$,
$$
X_t:=\mu+Z_t+\theta_1Z_{t-1}+\theta_2Z_{t-2}.
$$
Treat $\mu$, $\theta_1$, $\theta_2$, and $\sigma^2$ as
fixed features of the law. The innovations have mean zero and variance
$\sigma^2$. The feasible rule uses the observed path and the known
covariance range two.

*Method and report.* Complete parts (a)--(c) by hand. Use R for part (d), with
the stated seed and simulation sizes. Report the table and figure, then
interpret the variance ratios and coverage in complete sentences.

(a) Show that its covariance vanishes beyond lag two. Derive $\gamma(0)$,
$\gamma(1)$, $\gamma(2)$, $\tau^2$, and the exact value of
$q_T=T\Var(\bar X_T)$ for $T>2$.

(b) Explain why the three required sample autocovariances are consistent.
Write the positive-part estimator $\widehat v_{2,T}$ from Proposition 6.5 and
prove that it is ratio-consistent when $\tau^2>0$.

(c) Write the oracle 95% Chebyshev interval using the exact variance and
the feasible interval using $\widehat v_{2,T}$. State the coverage
guarantee for each, and identify the result that permits the estimated
variance to replace the known one.

(d) Set $\mu=0$, $\theta_1=1/2$, $\theta_2=1/4$, and $\sigma^2=1$.
Using seed `542`, simulate 5,000 paths at each
$T\in\{100,500\}$. For each $T$, report the mean, standard deviation, and
central 95% range of
$$
\frac{\widehat v_{2,T}}{\Var(\bar X_T)}.
$$
Also report the coverage of the oracle and feasible Chebyshev intervals,
with the estimated Monte Carlo standard error
$\sqrt{\hat p(1-\hat p)/5000}$ for each coverage fraction $\hat p$. Plot the
variance ratios at the two sample sizes. Use the generating parameters
only to calculate the true variance and score the oracle comparison.

(e) Interpret how the variance ratio and feasible coverage change with $T$.
Explain why the oracle interval has a finite-sample lower coverage bound,
while the feasible interval has only a limiting lower bound. Must either
coverage converge to exactly 95%?

```{r}
#| eval: false
# Student starter code
set.seed(542)
theta_1 <- 0.5
theta_2 <- 0.25
sample_acvf <- function(x, h) {
  n <- length(x)
  e <- x - mean(x)
  sum(e[seq_len(n - h)] * e[(h + 1):n]) / n
}
simulate_ma2_path <- function(n) {
  z <- rexp(n + 2) - 1
  z[3:(n + 2)] + theta_1 * z[2:(n + 1)] + theta_2 * z[1:n]
}
sample_sizes <- c(100, 500)
repetitions <- 5000

# For each path, compute the three sample autocovariances, the positive-part
# variance estimate, the variance ratio, and the two coverage indicators.
```
:::

::: {.exercise}
**Exercise 6.3 (Bartlett sensitivity for an AR(1)).** Let $(X_t)$ be the
stationary causal solution of
$$
X_t=\phi X_{t-1}+Z_t,
\qquad
Z_t\overset{\mathrm{iid}}{\sim}N(0,1-\phi^2),
\qquad \phi=0.8.
$$
Thus $\Var(X_t)=1$.

*Method and report.* Complete parts (a) and (b) by hand. Use R for part (c),
with the stated design. Report the table and figure, then interpret the
bandwidth comparison without proposing a universal rule.

(a) Derive the ACVF and show that $\tau^2=9$.

(b) If the population autocovariances were inserted into the Bartlett formula
at a fixed bandwidth $H$, its target would be
$$
\tau_H^2
:=1+2\sum_{h=1}^{H}
\left(1-\frac{h}{H+1}\right)0.8^h.
$$
Explain why $\tau_H^2<9$ at every fixed $H$ and why $H$ must grow to remove
this population bias.

(c) Using seed `542`, simulate 3,000 paths of length $T=300$. For
$H\in\{0,2,5,10,20,40,80,299\}$, compute the Bartlett estimate. Report its
population taper target $\tau_H^2$, Monte Carlo mean, Monte Carlo standard
deviation, and root mean squared error relative to $9$. Plot the simulated
estimates against $H$ and mark the true value.

(d) Identify the bandwidth with the smallest root mean squared error in this
simulation. Explain the bias--variability tradeoff visible in the table.

(e) Explain why the result does not provide a universal bandwidth rule. State
the roles of absolute covariance summability, a growing bandwidth,
collective sample-ACVF accuracy, and positive long-run variance in applying
Proposition 6.9. Is a CLT needed for its conservative interval?

```{r}
#| eval: false
# Student starter code
set.seed(542)
n <- 300
phi <- 0.8
bandwidths <- c(0, 2, 5, 10, 20, 40, 80, n - 1)

simulate_stationary_ar1 <- function(n, phi) {
  x <- numeric(n)
  x[1] <- rnorm(1)
  for (t in 2:n) {
    x[t] <- phi * x[t - 1] + rnorm(1, sd = sqrt(1 - phi^2))
  }
  x
}

# Write a function that computes the Bartlett estimate at one bandwidth,
# then use replicate() to apply it to 3,000 independent paths.
```
:::

## What a Variance Estimate Does Not Supply

Workbook 5 supplied the fixed covariance features that one path can learn.
A known cutoff or suitable Bartlett bandwidth lets us combine those
features into a ratio-consistent estimate of the sample-mean variance.
Chebyshev then gives an asymptotically conservative interval. None of
these arguments identifies the shape of the sampling distribution.
Workbook 7 asks which additional dependence conditions supply a central
limit theorem and permit a Gaussian interval.

## Sources

- van der Vaart, A. W. (2010). *Time Series*, Section 4.1, for the
  sample-mean variance and long-run variance.
- Newey, W. K., and West, K. D. (1987). “A Simple, Positive Semi-Definite,
  Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.”
  *Econometrica* 55, 703--708.
- Brockwell, P. J., and Davis, R. A. (1991). *Time Series: Theory and
  Methods*, 2nd ed. Springer.
- The `astsa::djia` help file supplies the DJIA data provenance and period.
