Workbook 6 — Estimating Uncertainty from One Path

STA 542 · Introduction to Time Series Analysis

Workbook 5 ended with a difficulty. At each fixed lag, one path may teach us the corresponding autocovariance. The variance of the sample mean, however, contains a sum whose number of lags grows with the sample size. Before trying to estimate that sum, we ask what population quantity it approaches and how one path can estimate it.

The Long-Run Variance

For a weakly stationary process with mean \(\mu\) and ACVF \(\gamma\), Workbook 5 established \[ T\Var(\bar X_T) =\gamma(0)+2\sum_{h=1}^{T-1} \left(1-\frac hT\right)\gamma(h). \] For independent observations, multiplication by \(T\) gives the fixed value \(\gamma(0)\). Under dependence, we ask when the weighted covariance sum on the right also approaches a finite limit.

Definition 6.1 (Long-run variance). If \[ \sum_{h\in\Z}|\gamma(h)|<\infty, \] the long-run variance of \((X_t)\) is \[ \tau^2:=\sum_{h\in\Z}\gamma(h) =\gamma(0)+2\sum_{h=1}^{\infty}\gamma(h). \] When \(\tau^2>0\), we write \(\tau:=\sqrt{\tau^2}\) for its positive square root.

Proposition 6.2 (The limiting variance scale). Under the condition in Definition 6.1, \[ T\Var(\bar X_T)\longrightarrow\tau^2. \] In particular, \(\tau^2\ge0\).

Proof. At each fixed lag, the weight \(1-h/T\) approaches one. To control the whole sum, we separate a fixed number of lags from the remaining tail. Absolute summability lets us make that tail small before increasing \(T\).

Let \(\varepsilon>0\). Choose an integer \(K\) large enough that \(\sum_{h>K}|\gamma(h)|<\varepsilon/4\), and then keep \(K\) fixed. For \(T>K\), subtracting the definition of \(\tau^2\) from the exact variance identity gives \[ T\Var(\bar X_T)-\tau^2 =-\frac2T\sum_{h=1}^{T-1}h\gamma(h) -2\sum_{h=T}^{\infty}\gamma(h). \] Split the first sum at \(K\). Since \(h/T\le1\) for every included lag, \[ \begin{aligned} \left|T\Var(\bar X_T)-\tau^2\right| &\le\frac2T\sum_{h=1}^{K}h|\gamma(h)| +2\sum_{h=K+1}^{T-1}\frac hT|\gamma(h)| +2\sum_{h=T}^{\infty}|\gamma(h)|\\ &\le\underbrace{\frac2T\sum_{h=1}^{K}h|\gamma(h)|}_{\text{finitely many lags}} +\underbrace{2\sum_{h>K}|\gamma(h)|}_{\text{remaining tail}}. \end{aligned} \]

The tail term is less than \(\varepsilon/2\) by our choice of \(K\). The finite sum in the first term is now a fixed number, so that term is also less than \(\varepsilon/2\) for all sufficiently large \(T\). The total error is then less than \(\varepsilon\), proving the limit. Since \(T\Var(\bar X_T)\ge0\) for every \(T\), its limit \(\tau^2\) is nonnegative. \(\square\)

When \(\tau^2>0\), \[ \Var(\bar X_T)\approx\frac{\tau^2}{T}. \] The marginal variance \(\gamma(0)\) is replaced by the sum of all autocovariances. This is the first effect of dependence on inference for the mean.

A zero long-run variance is possible. For \(X_t=Z_t-Z_{t-1}\) with i.i.d. centered innovations of variance \(\sigma^2>0\), the only nonzero covariances are \(\gamma(0)=2\sigma^2\) and \(\gamma(1)=-\sigma^2\). Hence \(\tau^2=\gamma(0)+2\gamma(1)=0\).

Summation cancels every interior innovation: \[ \bar X_T=\frac{Z_T-Z_0}{T},\qquad \Var(\bar X_T)=\frac{2\sigma^2}{T^2}. \] Thus \(T\Var(\bar X_T)\to0\), as Proposition 6.2 requires. The exact variance is still positive for each \(T\), but it has order \(T^{-2}\). The estimation results below use \(\tau^2>0\) when converting a consistent long-run-variance estimate into an estimate of the mean’s variance.

Accuracy on the Right Scale

Write \[ v_T:=\Var(\bar X_T). \] Under absolute covariance summability, \(v_T\to0\). The useless estimator \(\widetilde v_T:=0\) therefore satisfies \(|\widetilde v_T-v_T|\to0\). Absolute consistency does not ensure that an estimated standard error has the right size.

Definition 6.3 (Ratio-consistent variance estimate). Suppose \(v_T>0\) for all sufficiently large \(T\). A nonnegative estimator \(\widehat v_T\) is ratio-consistent for \(v_T\) if \[ \frac{\widehat v_T}{v_T}\xrightarrow{P}1. \]

Ratio consistency gives \(\sqrt{\widehat v_T}/\sqrt{v_T}\to1\) in probability. Thus the estimated standard error has the correct size relative to the unknown oracle standard error. The zero estimator fails this comparison completely.

Under absolute covariance summability, Proposition 6.2 gives \[ q_T:=T\Var(\bar X_T)\longrightarrow \tau^2=\sum_{h\in\Z}\gamma(h). \] If \(\tau^2>0\) and a nonnegative estimator satisfies \(\widehat\tau_T^2\to\tau^2\) in probability, then \[ \widehat v_T:=\frac{\widehat\tau_T^2}{T}, \qquad \frac{\widehat v_T}{v_T} =\frac{\widehat\tau_T^2}{q_T} \xrightarrow{P}1. \] We can therefore estimate the stable covariance sum rather than the vanishing variance directly.

Why Not Use Every Lag?

Workbook 5 defined the divisor-\(T\) sample ACVF by \[ \widehat\gamma_T(h) :=\frac1T\sum_{t=1}^{T-h} (X_t-\bar X_T)(X_{t+h}-\bar X_T), \qquad 0\le h<T. \] The most literal estimate of \(q_T\) substitutes every available sample autocovariance into its exact formula: \[ \widehat q_T^{\mathrm{all}} :=\widehat\gamma_T(0) +2\sum_{h=1}^{T-1} \left(1-\frac hT\right)\widehat\gamma_T(h). \]

Example 6.4 (Every available lag is too many). Suppose \[ X_t\overset{\mathrm{iid}}{\sim}N(\mu,\sigma^2), \qquad \sigma^2>0. \] The target is \(q_T=\sigma^2\). Every fixed-lag sample autocovariance is consistent, yet \[ \E[\widehat q_T^{\mathrm{all}}] =\sigma^2\frac{T^2-1}{3T^2} \longrightarrow\frac{\sigma^2}{3}. \] The small bias at each positive lag accumulates over order \(T\) estimated lags. The calculation uses only independence and a finite variance; normality merely gives the simulation below a fully specified law.

R code
estimate_iid_scales <- function(n) {
  x <- rnorm(n)
  e <- x - mean(x)
  c(
    `lag zero` = mean(e^2),
    `all lags` = 2 / n^2 * sum(cumsum(e)[seq_len(n - 1)]^2)
  )
}

set.seed(542)
all_lag_sizes <- c(50, 200, 800)
all_lag_repetitions <- 2500
all_lag_results <- lapply(all_lag_sizes, function(n) {
  t(replicate(all_lag_repetitions, estimate_iid_scales(n)))
})

all_lag_summary <- do.call(rbind, lapply(seq_along(all_lag_sizes), function(i) {
  values <- all_lag_results[[i]]
  data.frame(
    n = all_lag_sizes[i],
    method = colnames(values),
    center = colMeans(values),
    lower = apply(values, 2, quantile, probs = 0.10),
    upper = apply(values, 2, quantile, probs = 0.90)
  )
}))

method_colors <- c("lag zero" = "#0072B2", "all lags" = "#D55E00")
method_offsets <- c("lag zero" = -0.07, "all lags" = 0.07)
par(mfrow = c(1, 1), mar = c(4, 4, 2.5, 1), bg = "gray92")
plot(1:3, rep(1, 3), type = "n", ylim = c(0, 1.25), xaxt = "n",
     xlab = "Sample size T", ylab = "Estimated scaled variance",
     main = "Using every lag fails under independence")
axis(1, at = 1:3, labels = all_lag_sizes)
abline(h = 1, lty = 2)
abline(h = 1 / 3, lty = 3)
for (method in names(method_colors)) {
  rows <- all_lag_summary$method == method
  x_positions <- 1:3 + method_offsets[method]
  arrows(x_positions, all_lag_summary$lower[rows],
         x_positions, all_lag_summary$upper[rows],
         angle = 90, code = 3, length = 0.04,
         col = method_colors[method])
  points(x_positions, all_lag_summary$center[rows], pch = 16,
         col = method_colors[method])
}
legend("right", names(method_colors), col = method_colors,
       pch = 16, bty = "n")
Figure 1: Monte Carlo means and central 80% ranges for two estimates of the scaled sample-mean variance under i.i.d. N(0,1) sampling. The dashed line is the true value 1; the dotted line is the all-lag expectation limit 1/3. Based on 2,500 simulated paths at each sample size. Simulated data.

As \(T\) grows, the lag-zero estimate tightens around the true value \(1\). The all-lag estimate remains centered near \(1/3\) and does not tighten around the target. More sample autocovariances have produced a worse variance estimate.

Let \(e_t:=X_t-\bar X_T\). For i.i.d. observations with finite variance, \[ \E[e_se_t] =\sigma^2\left\{\mathbf 1(s=t)-\frac1T\right\}. \] Consequently, \[ \E[\widehat\gamma_T(0)] =\sigma^2\left(1-\frac1T\right), \qquad \E[\widehat\gamma_T(h)] =-\sigma^2\frac{T-h}{T^2}, \quad h\ge1. \]

Substitution into the all-lag estimator gives \[ \begin{aligned} \E[\widehat q_T^{\mathrm{all}}] &=\sigma^2\left(1-\frac1T\right) -\frac{2\sigma^2}{T} \sum_{h=1}^{T-1}\left(1-\frac hT\right)^2\\ &=\sigma^2\frac{T^2-1}{3T^2} \longrightarrow\frac{\sigma^2}{3}. \end{aligned} \] Each fixed-lag bias vanishes, but there are order \(T\) such biases.

The same estimator has the identity \[ \widehat q_T^{\mathrm{all}} =\frac{2}{T^2}\sum_{k=1}^{T-1} \left(\sum_{t=1}^k e_t\right)^2 \ge0. \] It follows by expanding both sides and using \(\sum_{t=1}^T e_t=0\).

Markov’s inequality now gives \[ \P\left(\widehat q_T^{\mathrm{all}}>\frac{\sigma^2}{2}\right) \le \frac{2\E[\widehat q_T^{\mathrm{all}}]}{\sigma^2} \longrightarrow\frac23. \] If \(\widehat q_T^{\mathrm{all}}\) converged in probability to \(\sigma^2\), the probability on the left would converge to one. The estimator is genuinely inconsistent; no additional limit theorem is needed.

A Known Covariance Range

Suppose a known integer \(m\ge0\) satisfies \[ \gamma(h)=0\qquad(|h|>m). \] This is a covariance cutoff, not an independence claim. Only the fixed lags \(0,\ldots,m\) enter the exact sample-mean variance, so Workbook 5’s fixed-lag learning result can now do all the required estimation.

For \(T>m\), define \[ \widehat q_{m,T} :=\widehat\gamma_T(0) +2\sum_{h=1}^{m} \left(1-\frac hT\right)\widehat\gamma_T(h), \qquad \widehat v_{m,T}:=\frac{\max(\widehat q_{m,T},0)}{T}. \] The positive part makes the variance estimate nonnegative. It has no asymptotic effect when the long-run variance is positive.

Proposition 6.5 (Known-range variance estimation). Suppose \((X_t)\) is weakly stationary, its ACVF is learnable through the known lag \(m\), and \(\gamma(h)=0\) for \(|h|>m\). If \[ \tau^2=\gamma(0)+2\sum_{h=1}^{m}\gamma(h)>0, \] then \[ \frac{\widehat v_{m,T}}{\Var(\bar X_T)}\xrightarrow{P}1. \]

Proof. ACVF learnability through \(m\) gives \(\widehat\gamma_T(h)\to\gamma(h)\) in probability for each \(h=0,\ldots,m\). Because the sum has a fixed number of terms, \[ \widehat q_{m,T}\xrightarrow{P}\tau^2. \] The covariance cutoff also gives \(q_T=T\Var(\bar X_T)\to\tau^2\). Positivity of the limit makes the positive part asymptotically irrelevant, and therefore \[ \frac{\widehat v_{m,T}}{\Var(\bar X_T)} =\frac{\max(\widehat q_{m,T},0)}{q_T} \xrightarrow{P}1. \] \(\square\)

Proposition 6.5 is complete only because \(m\) is fixed and known. A sample ACF cannot certify that a population covariance is exactly zero beyond a selected lag. Real series therefore return us to the growing-sum problem.

A Real Series Has No Known Cutoff

Example 6.6 (DJIA returns and squared returns). The object astsa::djia contains 2,518 daily closing observations of the Dow Jones Industrial Average, in index points, from April 20, 2006, through April 20, 2016. Let \(c_t\) denote an observed closing level and define the observed log return \[ r_t:=\log c_t-\log c_{t-1}. \] We subtract the sample mean from the returns before comparing the sample ACF of the adjusted returns with the sample ACF of their squares.

The notation astsa::djia accesses the data without attaching the package. The package must still be installed once; if needed, run install.packages("astsa") before rendering the workbook.

R code
djia_data <- astsa::djia
djia_closes <- as.numeric(djia_data[, "Close"])
djia_returns <- diff(log(djia_closes))
djia_returns <- djia_returns - mean(djia_returns)

djia_return_acf <- as.numeric(
  stats::acf(djia_returns, lag.max = 10, plot = FALSE)$acf
)[-1]
djia_squared_acf <- as.numeric(
  stats::acf(djia_returns^2, lag.max = 10, plot = FALSE)$acf
)[-1]

plot_acf_values <- function(values, title) {
  plot(1:10, values, type = "h", lwd = 2, ylim = c(-0.45, 0.45),
       xlab = "Lag (trading days)", ylab = "Sample ACF", main = title)
  points(1:10, values, pch = 16)
  abline(h = 0)
}

par(mfrow = c(1, 2), mar = c(4, 4, 2.5, 1), bg = "white")
plot_acf_values(djia_return_acf, "Drift-adjusted returns")
plot_acf_values(djia_squared_acf, "Squared returns")
Figure 2: Sample ACFs through lag 10 for DJIA daily drift-adjusted log returns and their squares, April 20, 2006–April 20, 2016. The squared returns have larger positive sample ACF values at every displayed lag. Real data: astsa::djia.

The signed-return sample ACF values range from \(-0.101\) to \(0.053\). They describe the relative lag pattern of signed returns; the corresponding sample autocovariances enter an estimated covariance sum for the mean return. The squared-return sample ACF values are positive at every displayed lag and range from \(0.190\) to \(0.414\). They concern a different feature: persistence in the magnitude of returns.

The squared-return plot suggests volatility persistence, but those correlations do not enter the long-run variance of the signed-return mean directly. At a fixed lag \(h\), the relevant product functions include \(g_h(x_0,\ldots,x_h)=x_0x_h\) for returns and \(g_h^{(2)}(x_0,\ldots,x_h)=x_0^2x_h^2\) for squared returns. A finite plot of their realized sample averages cannot establish \(g_h\)-ergodicity, convergence to population correlations, or a population covariance cutoff. Nor does it identify a volatility model. The plots pose a question about dependence; they do not supply the assumptions needed to learn it.

Bartlett Regularization

The known-range estimator keeps a fixed number of lags. The all-lag estimator keeps too many. For an unknown covariance range, we instead let the number of included lags grow slowly and downweight the least reliable ones.

Definition 6.7 (Bartlett estimator). Choose an integer bandwidth \(0\le H<T\). The Bartlett estimator of the long-run variance is \[ \widehat\tau_{H,T}^2 :=\widehat\gamma_T(0) +2\sum_{h=1}^{H} \left(1-\frac{h}{H+1}\right)\widehat\gamma_T(h). \] With the divisor-\(T\) sample ACVF, these triangular weights make the estimate nonnegative.

Two similar weights now have different jobs. The factor \(1-h/T\) in the exact variance formula counts the fraction of observed pairs at lag \(h\). The factor \(1-h/(H+1)\) in Definition 6.7 is a deliberate taper chosen by the analyst. At \(H=T-1\), the two weights coincide and Definition 6.7 reproduces the failed all-lag estimator from Example 6.4.

If \(H\) stays fixed while the ACVF has a nonzero tail, the estimator omits that tail forever. If \(H\) grows too quickly, the accumulated estimation error in the included sample autocovariances need not vanish. Consistency requires both a controllable population tail and collective accuracy over the retained lags.

Proposition 6.8 (Bartlett consistency under collective accuracy). Suppose \((X_t)\) is weakly stationary, \[ \sum_{h\in\Z}|\gamma(h)|<\infty, \] and let \((H_T)\) be a deterministic sequence of integer bandwidths satisfying \(H_T\to\infty\) and \(H_T<T\). If \[ \sum_{h=0}^{H_T} |\widehat\gamma_T(h)-\gamma(h)| \xrightarrow{P}0, \] then \[ \widehat\tau_{H_T,T}^2\xrightarrow{P}\tau^2. \] If also \(\tau^2>0\), then \(\widehat\tau_{H_T,T}^2/T\) is ratio-consistent for \(\Var(\bar X_T)\).

Proof. Put \(w_{h,T}:=1-h/(H_T+1)\). The triangle inequality gives \[ \begin{aligned} |\widehat\tau_{H_T,T}^2-\tau^2| \le{}& |\widehat\gamma_T(0)-\gamma(0)| +2\sum_{h=1}^{H_T}w_{h,T} |\widehat\gamma_T(h)-\gamma(h)|\\ &+2\sum_{h=1}^{H_T}(1-w_{h,T})|\gamma(h)| +2\sum_{h>H_T}|\gamma(h)|. \end{aligned} \] The collective-accuracy condition sends the first line to zero in probability. Absolute summability sends the omitted tail to zero.

For the taper term, fix an integer \(K\). Once \(H_T>K\), \[ \begin{aligned} 2\sum_{h=1}^{H_T}(1-w_{h,T})|\gamma(h)| &=2\sum_{h=1}^{H_T}\frac{h}{H_T+1}|\gamma(h)|\\ &\le\frac{2}{H_T+1}\sum_{h=1}^{K}h|\gamma(h)| +2\sum_{h>K}|\gamma(h)|. \end{aligned} \] First choose \(K\) to make the tail arbitrarily small, then increase \(T\) to make the fixed finite sum divided by \(H_T+1\) small. The taper term therefore tends to zero. This proves long-run-variance consistency. The ratio conclusion follows from Proposition 6.2 and \(\tau^2>0\). \(\square\)

The collective condition is stronger than consistency at every fixed lag. For example, it follows if \[ \sup_{0\le h\le H_T} \E[|\widehat\gamma_T(h)-\gamma(h)|] =O(T^{-1/2}) \] and \(H_T=o(\sqrt T)\). This calculation is only a sufficient illustration; different process classes permit different bandwidth ranges.

Suppose the fixed process law has the causal linear representation \[ X_t=\mu+\sum_{j=0}^{\infty}\psi_jZ_{t-j}, \] where \((Z_t)\) is i.i.d. with mean zero, variance \(\sigma^2>0\), and finite fourth moment, and \[ \sum_{j=0}^{\infty}|\psi_j|<\infty. \] For this short-memory class, the standard Bartlett result gives \[ \widehat\tau_{H_T,T}^2\xrightarrow{P}\tau^2 \] when \(H_T\to\infty\) and \(H_T/T\to0\). We take this model-specific result on faith. Its proof controls the weighted collection of sample autocovariances jointly; fixed-lag \(g\)-ergodicity alone does not provide that control.

The classical proof controls the weighted covariance sum directly. It need not verify the stronger sum of absolute errors in Proposition 6.8, which is why its bandwidth condition can be broader than the illustrative \(H_T=o(\sqrt T)\) rule.

Stable causal ARMA processes with i.i.d. innovations having a finite fourth moment belong to this class. A causal representation by itself does not establish Bartlett consistency for every nonlinear process, so we do not treat the result as a generic consequence of causality.

A Conservative Interval with an Estimated Variance

Workbook 5’s interval used the known variance \(v_T=\Var(\bar X_T)\). We now replace it by a nonnegative estimate. The replacement makes the threshold random and dependent on the same path as the mean, so the finite-sample Chebyshev calculation cannot simply be reused unchanged. Ratio consistency supplies the additional argument.

Proposition 6.9 (A feasible conservative interval). Fix a process law under which \(\E[\bar X_T]=\mu\) and \(0<v_T:=\Var(\bar X_T)<\infty\) for all sufficiently large \(T\). Suppose \(\widehat v_T\ge0\) and \(\widehat v_T/v_T\xrightarrow{P}1\). For a fixed \(0<\alpha<1\), define \[ C_T^{\mathrm{cons}} :=\left[ \bar X_T-\sqrt{\frac{\widehat v_T}{\alpha}}, \ \bar X_T+\sqrt{\frac{\widehat v_T}{\alpha}} \right]. \] Then \[ \liminf_{T\to\infty}\P\{\mu\in C_T^{\mathrm{cons}}\}\ge1-\alpha. \] No independence between \(\bar X_T\) and \(\widehat v_T\) is required.

Proof. Fix \(0<\varepsilon<1\) and let \[ A_T:=\{\widehat v_T\ge(1-\varepsilon)v_T\}. \] Ratio consistency gives \(\P(A_T^c)\to0\). On \(A_T\), failure of coverage implies \[ |\bar X_T-\mu|> \sqrt{\frac{(1-\varepsilon)v_T}{\alpha}}. \]

An event inclusion, followed by Chebyshev’s inequality, gives \[ \begin{aligned} \P\{\mu\notin C_T^{\mathrm{cons}}\} &\le\P(A_T^c) +\P\left\{|\bar X_T-\mu|> \sqrt{\frac{(1-\varepsilon)v_T}{\alpha}}\right\}\\ &\le\P(A_T^c)+\frac{\alpha}{1-\varepsilon}. \end{aligned} \] Taking the limit superior yields a noncoverage bound \(\alpha/(1-\varepsilon)\). Letting \(\varepsilon\downarrow0\) gives \(\limsup_T\P\{\mu\notin C_T^{\mathrm{cons}}\}\le\alpha\), which is the stated coverage conclusion. \(\square\)

For a weakly stationary process, \(\E[\bar X_T]=\mu\) automatically. Proposition 6.5 supplies a suitable \(\widehat v_T\) when the covariance range is known. Under Proposition 6.8 with \(\tau^2>0\), we may instead use \(\widehat v_T=\widehat\tau_{H_T,T}^2/T\). Both routes produce a computable interval from one path.

The conclusion is a limiting lower bound at each fixed process law. It neither promises coverage of at least \(1-\alpha\) at every finite sample size nor says that coverage converges to exactly \(1-\alpha\). No Gaussian approximation was used. A more precise account of the sampling shape may permit a shorter interval; that is a separate question.

Exercises

Exercise 6.1 (Dependence and effective sample size). Let \[ X_t=\mu+Z_t+\theta Z_{t-1}, \] where \((Z_t)\) is i.i.d. with mean zero and variance \(\sigma^2>0\). When \(\gamma(0)>0\) and \(\tau^2>0\), define the effective sample size for the mean by \[ T_{\mathrm{eff}}:=T\frac{\gamma(0)}{\tau^2}. \] It is the number of independent observations with marginal variance \(\gamma(0)\) that would match the dependent mean’s approximate variance \(\tau^2/T\). This definition concerns the sample mean, not every target.

Method and report. Derive the quantities by hand. Evaluate the requested ratios and explain in complete sentences what the comparison measures.

  1. Derive \(\gamma(0)\), \(\gamma(1)\), and \(\tau^2\).

  2. For \(\theta\ne-1\), derive \(T_{\mathrm{eff}}/T\). Evaluate it at \(\theta=0.8\) and \(\theta=-0.5\), then interpret both values.

  3. Explain what changes at \(\theta=-1\). Derive \(\bar X_T-\mu\) directly and its exact variance.

  4. Compare the MA(1) law above at \(\theta=1/2\) and \(\sigma^2=4/5\) with a second law for \((X_t)\): the stationary causal AR(1) \[ X_t-\mu=\frac25(X_{t-1}-\mu)+Z_t, \qquad \E[Z_t]=0,\qquad \Var(Z_t)=\frac{21}{25}, \] whose innovations are again i.i.d. Show that the two laws have the same marginal variance and lag-one correlation. Calculate their long-run variances and effective sample sizes. Explain why agreement at lag one does not give equal uncertainty about the mean.

Exercise 6.2 (Estimating uncertainty for an MA(2)). Let \[ E_t\overset{\mathrm{iid}}{\sim}\operatorname{Exponential}(1), \qquad Z_t:=\sigma(E_t-1),\qquad \sigma>0, \] and define, for \(t\in\Z\), \[ X_t:=\mu+Z_t+\theta_1Z_{t-1}+\theta_2Z_{t-2}. \] Treat \(\mu\), \(\theta_1\), \(\theta_2\), and \(\sigma^2\) as fixed features of the law. The innovations have mean zero and variance \(\sigma^2\). The feasible rule uses the observed path and the known covariance range two.

Method and report. Complete parts (a)–(c) by hand. Use R for part (d), with the stated seed and simulation sizes. Report the table and figure, then interpret the variance ratios and coverage in complete sentences.

  1. Show that its covariance vanishes beyond lag two. Derive \(\gamma(0)\), \(\gamma(1)\), \(\gamma(2)\), \(\tau^2\), and the exact value of \(q_T=T\Var(\bar X_T)\) for \(T>2\).

  2. Explain why the three required sample autocovariances are consistent. Write the positive-part estimator \(\widehat v_{2,T}\) from Proposition 6.5 and prove that it is ratio-consistent when \(\tau^2>0\).

  3. Write the oracle 95% Chebyshev interval using the exact variance and the feasible interval using \(\widehat v_{2,T}\). State the coverage guarantee for each, and identify the result that permits the estimated variance to replace the known one.

  4. Set \(\mu=0\), \(\theta_1=1/2\), \(\theta_2=1/4\), and \(\sigma^2=1\). Using seed 542, simulate 5,000 paths at each \(T\in\{100,500\}\). For each \(T\), report the mean, standard deviation, and central 95% range of \[ \frac{\widehat v_{2,T}}{\Var(\bar X_T)}. \] Also report the coverage of the oracle and feasible Chebyshev intervals, with the estimated Monte Carlo standard error \(\sqrt{\hat p(1-\hat p)/5000}\) for each coverage fraction \(\hat p\). Plot the variance ratios at the two sample sizes. Use the generating parameters only to calculate the true variance and score the oracle comparison.

  5. Interpret how the variance ratio and feasible coverage change with \(T\). Explain why the oracle interval has a finite-sample lower coverage bound, while the feasible interval has only a limiting lower bound. Must either coverage converge to exactly 95%?

# Student starter code
set.seed(542)
theta_1 <- 0.5
theta_2 <- 0.25
sample_acvf <- function(x, h) {
  n <- length(x)
  e <- x - mean(x)
  sum(e[seq_len(n - h)] * e[(h + 1):n]) / n
}
simulate_ma2_path <- function(n) {
  z <- rexp(n + 2) - 1
  z[3:(n + 2)] + theta_1 * z[2:(n + 1)] + theta_2 * z[1:n]
}
sample_sizes <- c(100, 500)
repetitions <- 5000

# For each path, compute the three sample autocovariances, the positive-part
# variance estimate, the variance ratio, and the two coverage indicators.

Exercise 6.3 (Bartlett sensitivity for an AR(1)). Let \((X_t)\) be the stationary causal solution of \[ X_t=\phi X_{t-1}+Z_t, \qquad Z_t\overset{\mathrm{iid}}{\sim}N(0,1-\phi^2), \qquad \phi=0.8. \] Thus \(\Var(X_t)=1\).

Method and report. Complete parts (a) and (b) by hand. Use R for part (c), with the stated design. Report the table and figure, then interpret the bandwidth comparison without proposing a universal rule.

  1. Derive the ACVF and show that \(\tau^2=9\).

  2. If the population autocovariances were inserted into the Bartlett formula at a fixed bandwidth \(H\), its target would be \[ \tau_H^2 :=1+2\sum_{h=1}^{H} \left(1-\frac{h}{H+1}\right)0.8^h. \] Explain why \(\tau_H^2<9\) at every fixed \(H\) and why \(H\) must grow to remove this population bias.

  3. Using seed 542, simulate 3,000 paths of length \(T=300\). For \(H\in\{0,2,5,10,20,40,80,299\}\), compute the Bartlett estimate. Report its population taper target \(\tau_H^2\), Monte Carlo mean, Monte Carlo standard deviation, and root mean squared error relative to \(9\). Plot the simulated estimates against \(H\) and mark the true value.

  4. Identify the bandwidth with the smallest root mean squared error in this simulation. Explain the bias–variability tradeoff visible in the table.

  5. Explain why the result does not provide a universal bandwidth rule. State the roles of absolute covariance summability, a growing bandwidth, collective sample-ACVF accuracy, and positive long-run variance in applying Proposition 6.9. Is a CLT needed for its conservative interval?

# Student starter code
set.seed(542)
n <- 300
phi <- 0.8
bandwidths <- c(0, 2, 5, 10, 20, 40, 80, n - 1)

simulate_stationary_ar1 <- function(n, phi) {
  x <- numeric(n)
  x[1] <- rnorm(1)
  for (t in 2:n) {
    x[t] <- phi * x[t - 1] + rnorm(1, sd = sqrt(1 - phi^2))
  }
  x
}

# Write a function that computes the Bartlett estimate at one bandwidth,
# then use replicate() to apply it to 3,000 independent paths.

What a Variance Estimate Does Not Supply

Workbook 5 supplied the fixed covariance features that one path can learn. A known cutoff or suitable Bartlett bandwidth lets us combine those features into a ratio-consistent estimate of the sample-mean variance. Chebyshev then gives an asymptotically conservative interval. None of these arguments identifies the shape of the sampling distribution. Workbook 7 asks which additional dependence conditions supply a central limit theorem and permit a Gaussian interval.

Sources

  • van der Vaart, A. W. (2010). Time Series, Section 4.1, for the sample-mean variance and long-run variance.
  • Newey, W. K., and West, K. D. (1987). “A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.” Econometrica 55, 703–708.
  • Brockwell, P. J., and Davis, R. A. (1991). Time Series: Theory and Methods, 2nd ed. Springer.
  • The astsa::djia help file supplies the DJIA data provenance and period.