Workbook 4 — Bayesian Forecasting Under Parameter Uncertainty

STA 542 · Introduction to Time Series Analysis

Workbook 3 assumed that one process law and all of its parameters were known. We now have a family of possible laws. A prior assigns weights to those laws, and forecasting carries the resulting uncertainty about the law into the future.

A Prior Completes the Probability Model

Let \(\Theta\) be the parameter space. We write the statistical model as \[ \{\P_\theta:\theta\in\Theta\}. \] For each fixed \(\theta\), \(\P_\theta\) is the complete joint law of the observed window \(W\) and a future target \(F\). We write \(\E_\theta\) and \(\Var_\theta\) for expectation and variance under that law. The subscript selects one candidate law; it is not a conditioning statement. In Example 4.3, for instance, \(\P_\delta\) will denote the random-walk law with the drift held fixed at \(\delta\).

Uppercase \(W\) and \(F\) denote random quantities. Lowercase \(w\) denotes the observed value of \(W\), and \(f\) denotes a possible value of \(F\).

A prior density \(\pi\) assigns weights to the candidate parameter values. The model family and prior together induce a hierarchical law, denoted by \(\P^\pi\), for the parameter, the observed window, and the future target. Under this law, one value \(\theta\) is drawn from \(\pi\), and the pair \((W,F)\) is then generated according to \(\P_\theta\). The same value of \(\theta\) governs the observed window and the future. For any event \(A\) involving only \(W\) and \(F\), \[ \P^\pi(A) = \int_\Theta \P_\theta(A)\pi(\theta)\,d\theta. \] We write \(\E^\pi\), \(\Var^\pi\), and \(\Cov^\pi\) for moments under this law. The superscript records the prior used to form the hierarchical law. Bare \(\P\) retains its generic meaning from the earlier workbooks; it is not shorthand for \(\P^\pi\).

Suppose that \((W,F)\) has joint density \(p_\theta(w,f)\) under \(\P_\theta\). Write \(p_\theta(w)\) for the marginal density of the observed window and \(p_\theta(f\mid w)\) for the fixed-law predictive density. Once \(w\) is observed, \(p_\theta(w)\) viewed as a function of \(\theta\) is the likelihood. It compares the support that the observed window gives to the candidate laws; it is not itself a probability density on \(\Theta\).

Definition 4.1 (Posterior predictive law). Assume \[ 0 < \int_\Theta p_{\theta'}(w)\pi(\theta')\,d\theta' < \infty. \] After observing \(W=w\), Bayes’ rule gives the posterior density \[ \pi(\theta\mid w) = \frac{p_\theta(w)\pi(\theta)} {\int_\Theta p_{\theta'}(w)\pi(\theta')\,d\theta'}, \] and the conditional density of \(F\) under \(\P^\pi\) is the posterior predictive density \[ p^\pi(f\mid w) = \int_\Theta p_\theta(f\mid w)\pi(\theta\mid w)\,d\theta. \] For a discrete parameter space, replace the prior density and the integrals over \(\Theta\) by prior masses and sums.

For each fixed \(\theta\), \(p_\theta(f\mid w)\) is the known-law forecast from Workbook 3. Posterior prediction averages these forecasts using the parameter values still plausible after observing \(w\). The observed path changes the weights assigned to the candidate laws, but the forecast does not replace the unknown parameter by one fitted value.

The notation keeps three operations separate:

  • a subscript \(\theta\) holds the parameter fixed and selects \(\P_\theta\);
  • a superscript \(\pi\) marks the hierarchical law induced by the prior; and
  • \(\mid W=w\) conditions on the observed random window.

Thus we write \(\P_\theta(A)\) for a fixed-law probability; we do not express the choice of a law with a conditioning bar. When prior averaging and conditioning on the observed window are both intended, we write \(\P^\pi(A\mid W=w)\).

Two Sources of Forecast Uncertainty

Proposition 4.2 (Posterior predictive mean and variance). Suppose \(F\) is square-integrable under \(\P^\pi\). For each fixed \(\theta\in\Theta\), write \[ m_\theta(W):=\E_\theta[F\mid W], \qquad v_\theta(W):=\Var_\theta(F\mid W), \] and define \[ \overline m(W) := \int_\Theta m_\theta(W)\pi(\theta\mid W)\,d\theta. \] Then \[ \E^\pi[F\mid W]=\overline m(W) \] and \[ \Var^\pi(F\mid W) = \int_\Theta v_\theta(W)\pi(\theta\mid W)\,d\theta + \int_\Theta \{m_\theta(W)-\overline m(W)\}^2 \pi(\theta\mid W)\,d\theta. \]

Proof. The mean identity follows by integrating the fixed-parameter conditional mean against the posterior density. For the variance, write \[ F-\overline m(W) = \{F-m_\theta(W)\} + \{m_\theta(W)-\overline m(W)\}. \] The remaining conditional-expectation calculation is worked through below.

Fix an observed value \(W=w\). Because \(\overline m(w)=\E^\pi[F\mid W=w]\), the posterior predictive variance is \[ \Var^\pi(F\mid W=w) = \E^\pi\!\left[ \{F-\overline m(w)\}^2 \mid W=w \right]. \]

First hold \(\theta\) fixed. Squaring the decomposition gives \[ \begin{aligned} \E_\theta\!\left[ \{F-\overline m(w)\}^2 \mid W=w \right] ={}& \E_\theta\!\left[ \{F-m_\theta(w)\}^2 \mid W=w \right]\\ &+2\{m_\theta(w)-\overline m(w)\} \E_\theta\!\left[ F-m_\theta(w) \mid W=w \right]\\ &+\{m_\theta(w)-\overline m(w)\}^2. \end{aligned} \] The first expectation is \(v_\theta(w)\). The expectation in the cross-product is \[ \E_\theta[F-m_\theta(w)\mid W=w] =m_\theta(w)-m_\theta(w) =0. \] The final squared term is constant once \(w\) and \(\theta\) are fixed. Therefore \[ \E_\theta\!\left[ \{F-\overline m(w)\}^2 \mid W=w \right] = v_\theta(w) + \{m_\theta(w)-\overline m(w)\}^2. \]

Now average this fixed-\(\theta\) identity using the posterior density \(\pi(\theta\mid w)\). The posterior predictive mixture gives \[ \begin{aligned} \Var^\pi(F\mid W=w) ={}& \int_\Theta v_\theta(w)\pi(\theta\mid w)\,d\theta\\ &+ \int_\Theta \{m_\theta(w)-\overline m(w)\}^2 \pi(\theta\mid w)\,d\theta. \end{aligned} \] Thus integrating over \(F\) at fixed \(\theta\) produces the two terms above, and integrating over the posterior combines them into the posterior predictive variance. Since the calculation holds for every observed \(w\), it proves the stated identity as a function of \(W\).

\(\square\)

The first variance term averages future variation at fixed parameter values. The second records variation in the parameter-dependent conditional mean. A plug-in forecast generally omits the second term.

A Random Walk with Unknown Drift

Example 4.3 (Unknown drift). For each \(\delta\in\R\), let \(\P_\delta\) be the law under which \[ X_t=X_{t-1}+\delta+Z_t, \qquad Z_t\overset{\mathrm{iid}}{\sim}N(0,\sigma^2), \] where \(X_0\) is observed and \(\sigma^2\) is known. Place the prior \[ \delta\sim N(m_0,s_0^2) \] on the drift and let \(\P^\pi\) denote the resulting hierarchical law. The drift draw is independent of the innovations used to construct the path. Under \(\P_\delta\), the process increments satisfy \[ \Delta X_t:=X_t-X_{t-1} \overset{\mathrm{iid}}{\sim}N(\delta,\sigma^2). \]

For observations \(x_0,\ldots,x_T\), write \[ X_{0:T}:=(X_0,\ldots,X_T), \qquad x_{0:T}:=(x_0,\ldots,x_T), \] and define the observed increments and their average by \[ \Delta x_t:=x_t-x_{t-1}, \qquad \overline{\Delta x}_T := \frac1T\sum_{t=1}^T\Delta x_t. \]

The increment likelihood and prior have kernels \[ p_\delta(\Delta x_1,\ldots,\Delta x_T) \propto \exp\!\left\{ -\frac1{2\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2 \right\} \] and \[ \pi(\delta) \propto \exp\!\left\{-\frac{(\delta-m_0)^2}{2s_0^2}\right\}. \] Their product is therefore proportional to \[ \exp\!\left[ -\frac12\left\{ \frac{(\delta-m_0)^2}{s_0^2} + \frac1{\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2 \right\} \right]. \]

Collecting the terms that involve \(\delta\) gives \[ \frac{(\delta-m_0)^2}{s_0^2} + \frac1{\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2 = A\delta^2-2B\delta+C, \] where \(C\) does not depend on \(\delta\) and \[ A = \frac1{s_0^2}+\frac{T}{\sigma^2}, \qquad B = \frac{m_0}{s_0^2} + \frac{T\overline{\Delta x}_T}{\sigma^2}. \] Now \[ A\delta^2-2B\delta+C = A\left(\delta-\frac BA\right)^2 + \left(C-\frac{B^2}{A}\right). \] The final parenthesis does not depend on \(\delta\), so it is absorbed into the normalizing constant. Hence, under \(\P^\pi\), the posterior distribution is \[ \delta\mid X_0=x_0,\ldots,X_T=x_T \sim N(m_T,s_T^2), \] where \[ s_T^2 = \left(\frac1{s_0^2}+\frac{T}{\sigma^2}\right)^{-1}, \qquad m_T = s_T^2\left( \frac{m_0}{s_0^2} +\frac{T\overline{\Delta x}_T}{\sigma^2} \right). \]

The precision of a normal distribution is the reciprocal of its variance. Under \(\P_\delta\), \[ \overline{\Delta X}_T := \frac1T\sum_{t=1}^T\Delta X_t \sim N\!\left(\delta,\frac{\sigma^2}{T}\right), \] so the sample mean increment has precision \(T/\sigma^2\). Thus the posterior precision is the sum of the prior precision and the data precision: \[ \frac1{s_T^2} = \frac1{s_0^2}+\frac{T}{\sigma^2}. \] Moreover, \[ m_T = \frac{s_0^{-2}}{s_0^{-2}+T\sigma^{-2}}\,m_0 + \frac{T\sigma^{-2}}{s_0^{-2}+T\sigma^{-2}}\, \overline{\Delta x}_T. \] The two coefficients are nonnegative and sum to one. More precise prior information pulls \(m_T\) toward \(m_0\); more increments, or less noisy increments, pull it toward \(\overline{\Delta x}_T\).

Under \(\P_\delta\), the known-drift forecast is \[ X_{T+h}\mid X_{0:T}=x_{0:T} \sim N(x_T+h\delta,h\sigma^2). \] Definition 4.1 therefore gives the posterior predictive distribution under \(\P^\pi\): \[ X_{T+h}\mid X_{0:T}=x_{0:T} \sim N\!\left( x_T+hm_T, \ h\sigma^2+h^2s_T^2 \right). \] The term \(h\sigma^2\) comes from future innovations. The term \(h^2s_T^2\) comes from the uncertain drift. ◊

The Plug-In Comparison

Replacing \(\delta\) by \(m_T\) gives the plug-in law \[ N(x_T+hm_T,h\sigma^2). \] Its mean agrees with the posterior predictive mean in Example 4.3 because the known-drift forecast is linear in \(\delta\). Its variance omits \(h^2s_T^2\). The omission grows quadratically with the forecast horizon.

Workbook 3 showed that moments alone give only a conservative interval. Example 4.3 established the entire Gaussian posterior predictive law, not only its mean and variance. Therefore, for a fixed horizon \(h\), the equal-tail interval \[ x_T+hm_T \ \pm\ z_{1-\alpha/2}\sqrt{h\sigma^2+h^2s_T^2} \] has posterior predictive probability \(1-\alpha\). Intervals drawn over several horizons form a pointwise predictive fan; they do not form a simultaneous region for the full future path.

R code
T_obs <- 12
H <- 20
delta_true <- 0.35
sigma <- 1.2
m0 <- 0
s0 <- 0.5

x_obs <- c(0, cumsum(delta_true + rnorm(T_obs, 0, sigma)))
mean_increment <- mean(diff(x_obs))
sT2 <- 1 / (1 / s0^2 + T_obs / sigma^2)
mT <- sT2 * (m0 / s0^2 + T_obs * mean_increment / sigma^2)

h <- seq_len(H)
pred_mean <- tail(x_obs, 1) + h * mT
pred_sd <- sqrt(h * sigma^2 + h^2 * sT2)
plugin_sd <- sqrt(h * sigma^2)
pred_lo <- pred_mean - qnorm(0.975) * pred_sd
pred_hi <- pred_mean + qnorm(0.975) * pred_sd
plugin_lo <- pred_mean - qnorm(0.975) * plugin_sd
plugin_hi <- pred_mean + qnorm(0.975) * plugin_sd

B <- 40
paths <- replicate(B, {
  delta_draw <- rnorm(1, mT, sqrt(sT2))
  tail(x_obs, 1) + cumsum(delta_draw + rnorm(H, 0, sigma))
})
paths <- rbind(rep(tail(x_obs, 1), B), paths)

par(mfrow = c(1, 2), mar = c(4, 4, 2.5, 1), bg = "gray92")
future_t <- T_obs + h
plot(0:T_obs, x_obs, type = "o", pch = 19,
     xlim = c(0, T_obs + H),
     ylim = range(x_obs, pred_lo, pred_hi),
     xlab = "Time", ylab = "X", main = "Pointwise prediction")
polygon(c(future_t, rev(future_t)), c(pred_lo, rev(pred_hi)),
        border = NA, col = adjustcolor(4, alpha.f = 0.16))
lines(future_t, pred_mean, col = 4, lwd = 2)
lines(future_t, plugin_lo, col = 2, lty = 2)
lines(future_t, plugin_hi, col = 2, lty = 2)
abline(v = T_obs, lty = 3)
legend("topleft",
       legend = c("Observed", "Predictive mean", "Posterior predictive", "Plug-in bounds"),
       col = c(1, 4, adjustcolor(4, alpha.f = 0.35), 2),
       lty = c(1, 1, 1, 2), lwd = c(1, 2, 6, 1), pch = c(19, NA, NA, NA),
       bty = "n")

matplot(0:H, paths, type = "l", lty = 1,
        col = adjustcolor("gray35", alpha.f = 0.35),
        xlab = "Forecast horizon", ylab = "X",
        main = "Joint predictive paths")
lines(0:H, c(tail(x_obs, 1), pred_mean), col = 4, lwd = 2)
Figure 1: Posterior prediction for a simulated random walk with unknown drift. Left: the 95% posterior predictive fan and the narrower plug-in bounds. Right: joint posterior predictive paths, each generated with one posterior drift draw. The intervals are pointwise. Simulated data.

In this simulation, the mean observed increment is 0.44. The posterior for the drift is \(N(0.30, 0.081)\), where the second reported number is the posterior variance.

The left panel describes each forecast horizon separately. The right panel shows draws of the entire future path. Exercise 4.3 determines how such paths must be generated when the same unknown drift governs every future increment.

What Does 95% Mean?

Calling a prediction interval 95% does not yet say how the 95% is calculated. A Bayesian statement conditions on the observed window and uses the posterior predictive law. A repeated-sample statement holds the parameter fixed and asks how often the interval captures the future across repeated observed and future paths. Throughout this section, every interval predicts the same future target \(F\). What changes is the probability experiment used to calibrate or evaluate the interval.

Posterior Predictive Probability

After observing \(W=w\), a 95% posterior predictive set \(C_{\mathrm B}(w)\) satisfies \[ \P^\pi\{F\in C_{\mathrm B}(w)\mid W=w\}=0.95. \] The observed window is now fixed. The probability averages over posterior uncertainty about the parameter and over future process variation.

Coverage When the Parameter Is Fixed

Let \(\pi(\theta)\) denote the prior density used to construct the posterior predictive rule. For a fixed value \(\theta\), define that rule’s repeated-sample coverage by \[ c_{\mathrm B}(\theta) := \P_\theta\{F\in C_{\mathrm B}(W)\}. \] This probability keeps \(\theta\) fixed and repeatedly generates both the observed window \(W\) and the future target \(F\) under \(\P_\theta\). Although \(C_{\mathrm B}\) was constructed using the prior, \(c_{\mathrm B}(\theta)\) evaluates its performance when this one value of \(\theta\) generates every repetition.

Now suppose \(C_{\mathrm B}(W)\) has posterior predictive probability 0.95 for every observed window for which the posterior is defined. Under the complete hierarchical model, \(\theta\) is first drawn from the density \(\pi\), after which \(W\) and \(F\) are drawn under the fixed law \(\P_\theta\).

For an event \(A\) involving \(W\) and \(F\), the definition of \(\P^\pi\) gives \[ \P^\pi(A) = \int_\Theta \P_\theta(A)\pi(\theta)\,d\theta. \] Applying this identity to the event \(A=\{F\in C_{\mathrm B}(W)\}\) gives the first equality below: \[\begin{align*} \int_\Theta c_{\mathrm B}(\theta)\pi(\theta)\,d\theta &= \P^\pi\{F\in C_{\mathrm B}(W)\}\\ &= \E^\pi\!\left[ \P^\pi\{F\in C_{\mathrm B}(W)\mid W\} \right]\\ &=0.95. \end{align*}\] The first equality averages the fixed-parameter coverage over the possible values of \(\theta\), using the prior as their distribution. The second is the tower property under \(\P^\pi\), now conditioning on \(W\). The final equality uses the assumed 95% posterior predictive probability at each observed window.

Thus Bayesian calibration gives 95% coverage on average over the prior. An average can equal 0.95 while individual values \(c_{\mathrm B}(\theta)\) differ substantially from 0.95. It does not follow that the prediction rule has 95% coverage at each fixed \(\theta\).

A finite-sample prediction rule \(C_{\mathrm H}(W)\) is honest over \(\Theta\) at level 95% if \[ \inf_{\theta\in\Theta} \P_\theta\{F\in C_{\mathrm H}(W)\} \geq 0.95. \] In words, its repeated-sample coverage is at least 95% at every parameter value in the declared model class. This is a worst-case requirement, not a prior average. It still averages over repeated versions of \(W\) and \(F\) with \(\theta\) fixed; it is not a conditional statement after one particular \(W=w\) has been observed.

A posterior predictive set can also be honest. Posterior predictive probability alone does not prove that it is: the fixed-parameter coverage function must be checked separately.

A Simple Predictive Comparison

Example 4.4 (Prior-average coverage can hide poor fixed-parameter coverage). Think of \(W\) as one observed increment and \(F\) as the next increment in Example 4.3. Take \[ \delta\sim N(0,1), \qquad W=\delta+Z_1, \qquad F=\delta+Z_2, \] where \(Z_1,Z_2\overset{\mathrm{iid}}{\sim}N(0,1)\) and are independent of \(\delta\). Under the resulting hierarchical law \(\P^\pi\), normal conjugacy gives \[ \delta\mid W=w \sim N\!\left(\frac w2,\frac12\right), \qquad F\mid W=w \sim N\!\left(\frac w2,\frac32\right). \] Let \(z_{0.975}\) denote the 97.5th percentile of the standard normal distribution. After averaging over the posterior of \(\delta\), the 95% posterior predictive interval for the next increment is \[ C_{\mathrm B}(w) = \frac w2\pm z_{0.975}\sqrt{\frac32}. \]

Now hold the drift fixed rather than drawing it from the prior. Under \(\P_\delta\), \[ F-\frac W2 = \frac\delta2+Z_2-\frac{Z_1}{2}, \] and therefore \[ F-\frac W2 \sim N\!\left(\frac\delta2,\frac54\right). \] Writing \(a:=z_{0.975}\sqrt{3/2}\) and letting \(\Phi\) denote the standard normal distribution function, the fixed-drift coverage is \[ c_{\mathrm B}(\delta) = \Phi\!\left(\frac{a-\delta/2}{\sqrt{5/4}}\right) - \Phi\!\left(\frac{-a-\delta/2}{\sqrt{5/4}}\right). \] For example, \[ c_{\mathrm B}(0)\approx0.968, \qquad c_{\mathrm B}(2)\approx0.894, \qquad c_{\mathrm B}(4)\approx0.640. \] The interval has exactly 95% posterior predictive probability for every observed \(w\), and exactly 95% coverage after averaging \(\delta\) over its \(N(0,1)\) prior. It is nevertheless not an honest 95% prediction interval over \(\R\). Its fixed-drift coverage is above 95% near the prior mean and falls below 95% sufficiently far from it.

For comparison, consider \[ C_{\mathrm H}(W) = W\pm z_{0.975}\sqrt2. \] Under every fixed law \(\P_\delta\), \[ F-W=Z_2-Z_1\sim N(0,2). \] Therefore \[ \P_\delta\{F\in C_{\mathrm H}(W)\}=0.95 \qquad\text{for every }\delta\in\R. \] Both \(C_{\mathrm B}\) and \(C_{\mathrm H}\) predict the same future increment \(F\). The second interval is honest over the independent, unit-variance Gaussian fixed-parameter family just specified. Its half-width is about \(2.77\), compared with about \(2.40\) for the Bayesian predictive interval. The Bayesian interval uses the prior to shrink its center and shorten its width; the honest interval makes a uniform fixed-drift guarantee.

Neither guarantee is model-free. The Bayesian statement relies on both the fixed-parameter family and the prior. The honest statement ranges over every parameter value, but only inside the stated Gaussian family with known variance. Neither one establishes that this family generated the data. Under additional regularity, posterior predictive intervals may attain approximately 95% coverage at each fixed parameter as the sample grows. Uniform honesty over \(\Theta\) is stronger and requires a uniform result; neither conclusion follows from the posterior formula alone.

The Bayesian Trade-off

What it buys. Bayesian forecasting does not replace the unknown parameter by one fitted value and then pretend that value is known. It averages the fixed-parameter forecasts over the posterior: \[ p^\pi(f\mid w) = \int_\Theta p_\theta(f\mid w)\pi(\theta\mid w)\,d\theta. \] The resulting predictive law carries parameter uncertainty into forecasts of future observations and paths. Point forecasts, prediction intervals, predictive densities, and path-event probabilities can all be derived from this law.

What it requires. The analyst must specify:

  • a family of fixed-parameter laws for the observations and future, whose observed-data densities supply the likelihood; and
  • a prior for the unknown parameter.

This is a stronger commitment than specifying only a few moments. Predictions may depend materially on the model and prior, especially when the observed path is weakly informative. Bayesian updating handles parameter uncertainty within the chosen model; it does not establish that the model or prior is appropriate. Model criticism and sensitivity analysis remain necessary.

What it may cost. Example 4.3 has a closed-form solution because the normal prior is conjugate to the Gaussian likelihood. Conjugacy is convenient but special. In other models, the posterior predictive integral must be approximated using numerical integration or simulation, such as Markov chain Monte Carlo. The quality of that approximation must then be checked.

The trade-off: Bayesian forecasting provides a unified treatment of parameter uncertainty in prediction, at the price of a fuller probability model, a prior, and sometimes substantial computation.

Looking Ahead

The posterior and predictive laws are defined for every finite \(T\) whenever the required densities exist. Their existence does not show that the posterior concentrates or that prior sensitivity disappears as \(T\) grows. The workbook on one-path learnability asks which population features time averages can recover; the workbook on identifiability asks whether those features determine a unique parameter.

Exercises

Exercise 4.1 (A four-increment forecast). Let \(\sigma^2=1\) and \(\delta\sim N(0,1)\). Starting from \(x_0=0\), suppose the four observed increments are \(1.2,-0.4,0.7,0.1\).

Method and report. Derive the quantities by hand; a calculator or R may be used for the final arithmetic and the normal quantile. Report \(m_4\), \(s_4^2\), the predictive mean and variance, the endpoints of both intervals to two decimal places, and the omitted variance term.

  1. Compute \(m_4\) and \(s_4^2\).

  2. Compute the posterior predictive mean and variance of \(X_9\).

  3. Compute a pointwise 95% posterior predictive interval for \(X_9\).

  4. Compute the corresponding plug-in interval using \(m_4\) as if it were the known drift. Identify the omitted variance term.

Exercise 4.2 (An uncertain AR coefficient). For each \(\phi\in\R\), let \(\P_\phi\) be the law under which \[ X_{t+1}=\phi X_t+Z_{t+1}, \qquad Z_t\overset{\mathrm{iid}}{\sim}N(0,\sigma^2), \] where \(X_0=x_0\) is fixed and \(\sigma^2\) is known. Suppose a prior \(\pi\) has produced the posterior density \(\pi(\phi\mid X_{0:T})\); \(\P^\pi\) denotes the resulting hierarchical law.

Method and report. Work by hand. Give the one- and two-step predictive means and variances as symbolic expressions in \(x_T\), \(\sigma^2\), and the required posterior moments of \(\phi\). End with one or two sentences identifying which posterior moments a plug-in forecast omits. No R output is required.

  1. Express the one-step posterior predictive mean in terms of the first posterior moment of \(\phi\).

  2. Use Proposition 4.2 to show that the one-step posterior predictive variance is \[ \sigma^2+X_T^2\Var^\pi(\phi\mid X_{0:T}). \]

  3. Derive the two-step posterior predictive mean and variance using posterior moments of \(\phi\). Do not assume that the posterior mean of \(\phi^2\) equals the square of the posterior mean of \(\phi\).

  4. Explain why a posterior mean plug-in forecast need not even preserve the posterior predictive mean at horizons greater than one.

Exercise 4.3 (One drift draw per path). Under Example 4.3, define \(D_j:=X_{T+j}-X_{T+j-1}\).

Method and report. Work by hand. Report the predictive mean, variance, and cross-covariance; explain the resulting dependence in complete sentences; and give an ordered simulation algorithm that distinguishes draws made once per path from draws made once per horizon. Finish by explaining how the simulated paths answer a path-level probability question. No R output is required.

  1. Find the posterior predictive mean and variance of \(D_j\).

  2. Show that, for \(j\ne k\), \[ \Cov^\pi(D_j,D_k\mid X_{0:T})=s_T^2. \]

  3. Explain why the future increments are independent under \(\P_\delta\) but dependent after the fixed-law forecasts are averaged over the posterior for \(\delta\).

  4. Write an algorithm for simulating the joint vector \((X_{T+1},\ldots,X_{T+H})\). State where a new random draw is made and where the same draw is retained.

  5. Explain how repeated joint path draws could approximate \[ \P^\pi\!\left(\max_{1\le h\le H}X_{T+h}>c\mid X_{0:T}\right), \] and why a collection of pointwise predictive intervals does not answer this question.

Exercise 4.4 (Two predictive calibration standards). Let \(C_{\mathrm B}(W)\) be an exact \(1-\alpha\) posterior predictive set for a future target \(F\).

Method and report. Work by hand and keep every probability law or conditioning event explicit. Report the posterior predictive probability, the iterated-expectation proof, the fixed-parameter coverage function, and a short explanation of the honesty requirement. No R output is required.

  1. Write the posterior predictive probability statement that defines \(C_{\mathrm B}(W)\). After conditioning on \(W\), which sources of uncertainty remain in this probability?

  2. Use iterated expectation to prove that \(C_{\mathrm B}(W)\) has \(1-\alpha\) coverage under \(\P^\pi\). State which quantities this probability averages over.

  3. Define the fixed-parameter coverage function \[ c_{\mathrm B}(\theta) := \P_\theta\{F\in C_{\mathrm B}(W)\} \] and express the result of part (b) as a prior-weighted average of \(c_{\mathrm B}(\theta)\).

  4. State the condition that would make \(C_{\mathrm B}(W)\) honest over \(\Theta\). Explain why the identity in part (c) does not establish that condition.

  5. Explain why even an honest prediction interval over \(\Theta\) is not a model-free guarantee.

Exercise 4.5 (What would make the prior fade?). In Example 4.3, keep \(m_0,s_0^2\), and \(\sigma^2\) fixed while \(T\) grows.

Method and report. Work by hand. Show the two limits, state the law of large numbers with the random variables to which it applies, and end with two or three sentences explaining what the posterior algebra alone cannot establish. No R output is required.

  1. Show that \(s_T^2\to0\).

  2. Show that \(m_T-\overline{\Delta X}_T\to0\) in probability under \(\P^\pi\).

  3. Under \(\P_\delta\), state the law of large numbers that makes \(\overline{\Delta X}_T\to\delta\).

  4. Explain why the algebraic existence of the posterior alone would not have proved the conclusion in part (c).

Exercise 4.6 (How fast is the land warming?). The series astsa::gtemp_land contains annual land-surface temperature anomalies in degrees centigrade, 1850–2023. Restrict to 1973 onward, so that \(x_0\) is the 1973 value and the increments are the \(T=50\) annual changes through 2023, and apply Example 4.3 with \(\sigma\) treated as known at the sample standard deviation of the increments.

Method and report. Use R for every numerical calculation and for the plot. Derive the three increment formulas in (d) by hand. Submit one time plot of the restricted level series and a numerical summary, organized by parts (a)–(e), containing every quantity explicitly requested below. Answer the modeling questions in (c) and (e) in complete sentences; code and numerical output alone do not answer them.

  1. Plot the restricted series and report \(\overline{\Delta x}_T\) and the sample standard deviation of the increments. Also report \(T\), \(x_0\), \(x_T\), the total rise \(x_T-x_0\), and the ratio of the increment standard deviation to \(|\overline{\Delta x}_T|\). Numerically verify that the total rise equals \(T\overline{\Delta x}_T\).

  2. Take the prior \(\delta\sim N(0,0.5^2)\). Report the two precision terms in \(1/s_T^2\), their ratio, \(m_T\), \(s_T^2\), and \(\P^\pi(\delta<0\mid X_{0:T}=x_{0:T})\). State whether the data precision dominates the prior precision.

  3. Compute the 95% posterior predictive interval for the 2033 anomaly (\(h=10\)) and the corresponding plug-in interval. Report the common forecast center, both pairs of interval endpoints, and the fraction of the posterior predictive variance contributed by the drift term. In two or three sentences, explain whether either interval gives an informative forecast of the 2033 anomaly.

  4. Consider the alternative description \(X_t=a+bt+\varepsilon_t\) with \((\varepsilon_t)\) i.i.d. \(N(0,\sigma_\varepsilon^2)\). Derive the mean, variance, and lag-one correlation of its increments \(\Delta X_t=b+\varepsilon_t-\varepsilon_{t-1}\) by hand. Use R to compute the sample correlation of consecutive observed increments, report both correlations, and state which increment description this diagnostic favors.

  5. Interpret in complete sentences. Both descriptions estimate the drift by the same number \((x_T-x_0)/T=\overline{\Delta x}_T\). Under Example 4.3 this average has standard deviation \(\sigma/\sqrt T\); under the description in (d) it telescopes to \(b+(\varepsilon_T-\varepsilon_0)/T\). Evaluate both standard deviations, using (d) to express \(\sigma_\varepsilon\) through the increment standard deviation, and report their ratio. In one paragraph, explain why the two descriptions disagree about the precision of the same estimate and state every assumption on which the posterior probability in

  6. remains conditional. Hint: under Example 4.3 every annual change is permanent, so noise accumulates; under (d) the noise in \(x_T-x_0\) never accumulates beyond its two endpoints.


Sources for this workbook: West and Harrison (1997), Bayesian Forecasting and Dynamic Models, 2nd ed., Springer, for posterior prediction; Gelman et al. (2013), Bayesian Data Analysis, 3rd ed., CRC Press, for posterior predictive distributions and variance decomposition. Data: R package astsa (D. Stoffer).