
---
title: "Workbook 4 — Bayesian Forecasting Under Parameter Uncertainty"
subtitle: "STA 542 · Introduction to Time Series Analysis"
format:
  html:
    embed-resources: true
    toc: true
    toc-depth: 2
    number-sections: false
    code-copy: true
    fig-width: 9
    fig-height: 3.4
    fig-dpi: 150
execute:
  warning: false
  message: false
editor:
  markdown:
    wrap: 72
---

::: hidden
$$\def\E{\mathbb{E}}
\def\P{\mathbb{P}}
\def\R{\mathbb{R}}
\def\Z{\mathbb{Z}}
\def\N{\mathbb{N}}
\def\cH{\mathcal{H}}
\def\cP{\mathcal{P}}
\def\Var{\operatorname{Var}}
\def\Cov{\operatorname{Cov}}
\def\Corr{\operatorname{Corr}}$$
:::

```{=html}
<style>
.exercise { border-left: 3px solid #00539B; padding: 0.7em 1.1em; background: #f6f8fa; margin: 1.4em 0; border-radius: 0 4px 4px 0; }
.exercise p:first-child { margin-top: 0; }
figure figcaption { color: #555; font-size: 0.9em; }
</style>
```

```{r}
#| label: setup
#| include: false
set.seed(542)
```

Workbook 3 assumed that one process law and all of its parameters were known.
We now have a family of possible laws. A prior assigns weights to those laws,
and forecasting carries the resulting uncertainty about the law into the
future.

## A Prior Completes the Probability Model

Let $\Theta$ be the parameter space. We write the statistical model as
$$
\{\P_\theta:\theta\in\Theta\}.
$$
For each fixed $\theta$, $\P_\theta$ is the complete joint law of the observed
window $W$ and a future target $F$. We write $\E_\theta$ and $\Var_\theta$ for
expectation and variance under that law. The subscript selects one candidate
law; it is not a conditioning statement. In Example 4.3, for instance,
$\P_\delta$ will denote the random-walk law with the drift held fixed at
$\delta$.

Uppercase $W$ and $F$ denote random quantities. Lowercase $w$ denotes the
observed value of $W$, and $f$ denotes a possible value of $F$.

A prior density $\pi$ assigns weights to the candidate parameter values. The
model family and prior together induce a *hierarchical law*, denoted by
$\P^\pi$, for the parameter, the observed window, and the future target. Under
this law, one value $\theta$ is drawn from $\pi$, and the pair $(W,F)$ is then
generated according to $\P_\theta$. The same value of $\theta$ governs the
observed window and the future. For any event $A$ involving only $W$ and $F$,
$$
\P^\pi(A)
=
\int_\Theta \P_\theta(A)\pi(\theta)\,d\theta.
$$
We write $\E^\pi$, $\Var^\pi$, and $\Cov^\pi$ for moments under this law. The
superscript records the prior used to form the hierarchical law. Bare $\P$
retains its generic meaning from the earlier workbooks; it is not shorthand
for $\P^\pi$.

Suppose that $(W,F)$ has joint density $p_\theta(w,f)$ under $\P_\theta$.
Write $p_\theta(w)$ for the marginal density of the observed window and
$p_\theta(f\mid w)$ for the fixed-law predictive density. Once $w$ is
observed, $p_\theta(w)$ viewed as a function of $\theta$ is the *likelihood*.
It compares the support that the observed window gives to the candidate laws;
it is not itself a probability density on $\Theta$.

**Definition 4.1 (Posterior predictive law).** Assume
$$
0
<
\int_\Theta p_{\theta'}(w)\pi(\theta')\,d\theta'
<
\infty.
$$
After observing $W=w$, Bayes' rule gives the posterior density
$$
\pi(\theta\mid w)
=
\frac{p_\theta(w)\pi(\theta)}
{\int_\Theta p_{\theta'}(w)\pi(\theta')\,d\theta'},
$$
and the conditional density of $F$ under $\P^\pi$ is the *posterior
predictive density*
$$
p^\pi(f\mid w)
=
\int_\Theta p_\theta(f\mid w)\pi(\theta\mid w)\,d\theta.
$$
For a discrete parameter space, replace the prior density and the integrals
over $\Theta$ by prior masses and sums.

For each fixed $\theta$, $p_\theta(f\mid w)$ is the known-law forecast from
Workbook 3. Posterior prediction averages these forecasts using the parameter
values still plausible after observing $w$. The observed path changes the
weights assigned to the candidate laws, but the forecast does not replace the
unknown parameter by one fitted value.

The notation keeps three operations separate:

- a subscript $\theta$ holds the parameter fixed and selects $\P_\theta$;
- a superscript $\pi$ marks the hierarchical law induced by the prior; and
- $\mid W=w$ conditions on the observed random window.

Thus we write $\P_\theta(A)$ for a fixed-law probability; we do not express
the choice of a law with a conditioning bar. When prior averaging and
conditioning on the observed window are both intended, we write
$\P^\pi(A\mid W=w)$.

## Two Sources of Forecast Uncertainty

**Proposition 4.2 (Posterior predictive mean and variance).** Suppose $F$ is
square-integrable under $\P^\pi$. For each fixed $\theta\in\Theta$, write
$$
m_\theta(W):=\E_\theta[F\mid W],
\qquad
v_\theta(W):=\Var_\theta(F\mid W),
$$
and define
$$
\overline m(W)
:=
\int_\Theta m_\theta(W)\pi(\theta\mid W)\,d\theta.
$$
Then
$$
\E^\pi[F\mid W]=\overline m(W)
$$
and
$$
\Var^\pi(F\mid W)
=
\int_\Theta v_\theta(W)\pi(\theta\mid W)\,d\theta
+
\int_\Theta
\{m_\theta(W)-\overline m(W)\}^2
\pi(\theta\mid W)\,d\theta.
$$

*Proof.* The mean identity follows by integrating the fixed-parameter
conditional mean against the posterior density. For the variance, write
$$
F-\overline m(W)
=
\{F-m_\theta(W)\}
+
\{m_\theta(W)-\overline m(W)\}.
$$
The remaining conditional-expectation calculation is worked through below.

::: {.callout-note collapse="true"}
## More details

Fix an observed value $W=w$. Because
$\overline m(w)=\E^\pi[F\mid W=w]$, the posterior predictive variance is
$$
\Var^\pi(F\mid W=w)
=
\E^\pi\!\left[
\{F-\overline m(w)\}^2
\mid W=w
\right].
$$

First hold $\theta$ fixed. Squaring the decomposition gives
$$
\begin{aligned}
\E_\theta\!\left[
\{F-\overline m(w)\}^2
\mid W=w
\right]
={}&
\E_\theta\!\left[
\{F-m_\theta(w)\}^2
\mid W=w
\right]\\
&+2\{m_\theta(w)-\overline m(w)\}
\E_\theta\!\left[
F-m_\theta(w)
\mid W=w
\right]\\
&+\{m_\theta(w)-\overline m(w)\}^2.
\end{aligned}
$$
The first expectation is $v_\theta(w)$. The expectation in the cross-product
is
$$
\E_\theta[F-m_\theta(w)\mid W=w]
=m_\theta(w)-m_\theta(w)
=0.
$$
The final squared term is constant once $w$ and $\theta$ are fixed. Therefore
$$
\E_\theta\!\left[
\{F-\overline m(w)\}^2
\mid W=w
\right]
=
v_\theta(w)
+
\{m_\theta(w)-\overline m(w)\}^2.
$$

Now average this fixed-$\theta$ identity using the posterior density
$\pi(\theta\mid w)$. The posterior predictive mixture gives
$$
\begin{aligned}
\Var^\pi(F\mid W=w)
={}&
\int_\Theta v_\theta(w)\pi(\theta\mid w)\,d\theta\\
&+
\int_\Theta
\{m_\theta(w)-\overline m(w)\}^2
\pi(\theta\mid w)\,d\theta.
\end{aligned}
$$
Thus integrating over $F$ at fixed $\theta$ produces the two terms above,
and integrating over the posterior combines them into the posterior predictive
variance. Since the calculation holds for every observed $w$, it proves the
stated identity as a function of $W$.
:::

$\square$

The first variance term averages future variation at fixed parameter values.
The second records variation in the parameter-dependent conditional mean. A
plug-in forecast generally omits the second term.

## A Random Walk with Unknown Drift

**Example 4.3 (Unknown drift).** For each $\delta\in\R$, let $\P_\delta$ be
the law under which
$$
X_t=X_{t-1}+\delta+Z_t,
\qquad
Z_t\overset{\mathrm{iid}}{\sim}N(0,\sigma^2),
$$
where $X_0$ is observed and $\sigma^2$ is known. Place the prior
$$
\delta\sim N(m_0,s_0^2)
$$
on the drift and let $\P^\pi$ denote the resulting hierarchical law. The drift
draw is independent of the innovations used to construct the path. Under
$\P_\delta$, the process increments satisfy
$$
\Delta X_t:=X_t-X_{t-1}
\overset{\mathrm{iid}}{\sim}N(\delta,\sigma^2).
$$

For observations $x_0,\ldots,x_T$, write
$$
X_{0:T}:=(X_0,\ldots,X_T),
\qquad
x_{0:T}:=(x_0,\ldots,x_T),
$$
and define the observed increments and their average by
$$
\Delta x_t:=x_t-x_{t-1},
\qquad
\overline{\Delta x}_T
:=
\frac1T\sum_{t=1}^T\Delta x_t.
$$

The increment likelihood and prior have kernels
$$
p_\delta(\Delta x_1,\ldots,\Delta x_T)
\propto
\exp\!\left\{
-\frac1{2\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2
\right\}
$$
and
$$
\pi(\delta)
\propto
\exp\!\left\{-\frac{(\delta-m_0)^2}{2s_0^2}\right\}.
$$
Their product is therefore proportional to
$$
\exp\!\left[
-\frac12\left\{
\frac{(\delta-m_0)^2}{s_0^2}
+
\frac1{\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2
\right\}
\right].
$$

Collecting the terms that involve $\delta$ gives
$$
\frac{(\delta-m_0)^2}{s_0^2}
+
\frac1{\sigma^2}\sum_{t=1}^T(\Delta x_t-\delta)^2
=
A\delta^2-2B\delta+C,
$$
where $C$ does not depend on $\delta$ and
$$
A
=
\frac1{s_0^2}+\frac{T}{\sigma^2},
\qquad
B
=
\frac{m_0}{s_0^2}
+
\frac{T\overline{\Delta x}_T}{\sigma^2}.
$$
Now
$$
A\delta^2-2B\delta+C
=
A\left(\delta-\frac BA\right)^2
+
\left(C-\frac{B^2}{A}\right).
$$
The final parenthesis does not depend on $\delta$, so it is absorbed into the
normalizing constant. Hence, under $\P^\pi$, the posterior distribution is
$$
\delta\mid X_0=x_0,\ldots,X_T=x_T
\sim N(m_T,s_T^2),
$$
where
$$
s_T^2
=
\left(\frac1{s_0^2}+\frac{T}{\sigma^2}\right)^{-1},
\qquad
m_T
=
s_T^2\left(
\frac{m_0}{s_0^2}
+\frac{T\overline{\Delta x}_T}{\sigma^2}
\right).
$$

The *precision* of a normal distribution is the reciprocal of its variance.
Under $\P_\delta$,
$$
\overline{\Delta X}_T
:=
\frac1T\sum_{t=1}^T\Delta X_t
\sim
N\!\left(\delta,\frac{\sigma^2}{T}\right),
$$
so the sample mean increment has precision $T/\sigma^2$. Thus the posterior
precision is the sum of the prior precision and the data precision:
$$
\frac1{s_T^2}
=
\frac1{s_0^2}+\frac{T}{\sigma^2}.
$$
Moreover,
$$
m_T
=
\frac{s_0^{-2}}{s_0^{-2}+T\sigma^{-2}}\,m_0
+
\frac{T\sigma^{-2}}{s_0^{-2}+T\sigma^{-2}}\,
\overline{\Delta x}_T.
$$
The two coefficients are nonnegative and sum to one. More precise prior
information pulls $m_T$ toward $m_0$; more increments, or less noisy
increments, pull it toward $\overline{\Delta x}_T$.

Under $\P_\delta$, the known-drift forecast is
$$
X_{T+h}\mid X_{0:T}=x_{0:T}
\sim N(x_T+h\delta,h\sigma^2).
$$
Definition 4.1 therefore gives the posterior predictive distribution under
$\P^\pi$:
$$
X_{T+h}\mid X_{0:T}=x_{0:T}
\sim
N\!\left(
x_T+hm_T,
\ h\sigma^2+h^2s_T^2
\right).
$$
The term $h\sigma^2$ comes from future innovations. The term $h^2s_T^2$
comes from the uncertain drift. ◊

## The Plug-In Comparison

Replacing $\delta$ by $m_T$ gives the plug-in law
$$
N(x_T+hm_T,h\sigma^2).
$$
Its mean agrees with the posterior predictive mean in Example 4.3 because the
known-drift forecast is linear in $\delta$. Its variance omits $h^2s_T^2$.
The omission grows quadratically with the forecast horizon.

Workbook 3 showed that moments alone give only a conservative interval.
Example 4.3 established the entire Gaussian posterior predictive law, not only
its mean and variance. Therefore, for a fixed horizon $h$, the equal-tail
interval
$$
x_T+hm_T
\ \pm\
z_{1-\alpha/2}\sqrt{h\sigma^2+h^2s_T^2}
$$
has posterior predictive probability $1-\alpha$. Intervals drawn over several
horizons form a pointwise predictive fan; they do not form a simultaneous
region for the full future path.

```{r}
#| label: fig-bayesian-drift
#| fig-height: 4.2
#| code-fold: true
#| code-summary: "R code"
#| fig-cap: "Posterior prediction for a simulated random walk with unknown drift. Left: the 95% posterior predictive fan and the narrower plug-in bounds. Right: joint posterior predictive paths, each generated with one posterior drift draw. The intervals are pointwise. Simulated data."
T_obs <- 12
H <- 20
delta_true <- 0.35
sigma <- 1.2
m0 <- 0
s0 <- 0.5

x_obs <- c(0, cumsum(delta_true + rnorm(T_obs, 0, sigma)))
mean_increment <- mean(diff(x_obs))
sT2 <- 1 / (1 / s0^2 + T_obs / sigma^2)
mT <- sT2 * (m0 / s0^2 + T_obs * mean_increment / sigma^2)

h <- seq_len(H)
pred_mean <- tail(x_obs, 1) + h * mT
pred_sd <- sqrt(h * sigma^2 + h^2 * sT2)
plugin_sd <- sqrt(h * sigma^2)
pred_lo <- pred_mean - qnorm(0.975) * pred_sd
pred_hi <- pred_mean + qnorm(0.975) * pred_sd
plugin_lo <- pred_mean - qnorm(0.975) * plugin_sd
plugin_hi <- pred_mean + qnorm(0.975) * plugin_sd

B <- 40
paths <- replicate(B, {
  delta_draw <- rnorm(1, mT, sqrt(sT2))
  tail(x_obs, 1) + cumsum(delta_draw + rnorm(H, 0, sigma))
})
paths <- rbind(rep(tail(x_obs, 1), B), paths)

par(mfrow = c(1, 2), mar = c(4, 4, 2.5, 1), bg = "gray92")
future_t <- T_obs + h
plot(0:T_obs, x_obs, type = "o", pch = 19,
     xlim = c(0, T_obs + H),
     ylim = range(x_obs, pred_lo, pred_hi),
     xlab = "Time", ylab = "X", main = "Pointwise prediction")
polygon(c(future_t, rev(future_t)), c(pred_lo, rev(pred_hi)),
        border = NA, col = adjustcolor(4, alpha.f = 0.16))
lines(future_t, pred_mean, col = 4, lwd = 2)
lines(future_t, plugin_lo, col = 2, lty = 2)
lines(future_t, plugin_hi, col = 2, lty = 2)
abline(v = T_obs, lty = 3)
legend("topleft",
       legend = c("Observed", "Predictive mean", "Posterior predictive", "Plug-in bounds"),
       col = c(1, 4, adjustcolor(4, alpha.f = 0.35), 2),
       lty = c(1, 1, 1, 2), lwd = c(1, 2, 6, 1), pch = c(19, NA, NA, NA),
       bty = "n")

matplot(0:H, paths, type = "l", lty = 1,
        col = adjustcolor("gray35", alpha.f = 0.35),
        xlab = "Forecast horizon", ylab = "X",
        main = "Joint predictive paths")
lines(0:H, c(tail(x_obs, 1), pred_mean), col = 4, lwd = 2)
```

In this simulation, the mean observed increment is
`r sprintf("%.2f", mean_increment)`. The posterior for the drift is
$N(`r sprintf("%.2f", mT)`, `r sprintf("%.3f", sT2)`)$, where the second
reported number is the posterior variance.

The left panel describes each forecast horizon separately. The right panel
shows draws of the entire future path. Exercise 4.3 determines how such paths
must be generated when the same unknown drift governs every future increment.

## What Does 95% Mean?

Calling a prediction interval *95%* does not yet say how the 95% is
calculated. A Bayesian statement conditions on the observed window and uses
the posterior predictive law. A repeated-sample statement holds the parameter
fixed and asks how often the interval captures the future across repeated
observed and future paths. Throughout this section, every interval predicts
the same future target $F$. What changes is the probability experiment used
to calibrate or evaluate the interval.

### Posterior Predictive Probability

After observing $W=w$, a 95% posterior predictive set
$C_{\mathrm B}(w)$ satisfies
$$
\P^\pi\{F\in C_{\mathrm B}(w)\mid W=w\}=0.95.
$$
The observed window is now fixed. The probability averages over posterior
uncertainty about the parameter and over future process variation.

### Coverage When the Parameter Is Fixed

Let $\pi(\theta)$ denote the prior density used to construct the posterior
predictive rule. For a fixed value $\theta$, define that rule's
repeated-sample coverage by
$$
c_{\mathrm B}(\theta)
:=
\P_\theta\{F\in C_{\mathrm B}(W)\}.
$$
This probability keeps $\theta$ fixed and repeatedly generates both the
observed window $W$ and the future target $F$ under $\P_\theta$. Although
$C_{\mathrm B}$ was constructed using the prior, $c_{\mathrm B}(\theta)$
evaluates its performance when this one value of $\theta$ generates every
repetition.

Now suppose $C_{\mathrm B}(W)$ has posterior predictive probability 0.95 for
every observed window for which the posterior is defined. Under the complete
hierarchical model, $\theta$ is first drawn from the density $\pi$, after
which $W$ and $F$ are drawn under the fixed law $\P_\theta$.

For an event $A$ involving $W$ and $F$, the definition of $\P^\pi$ gives
$$
\P^\pi(A)
=
\int_\Theta \P_\theta(A)\pi(\theta)\,d\theta.
$$
Applying this identity to the event
$A=\{F\in C_{\mathrm B}(W)\}$ gives the first equality below:
\begin{align*}
\int_\Theta c_{\mathrm B}(\theta)\pi(\theta)\,d\theta
&=
\P^\pi\{F\in C_{\mathrm B}(W)\}\\
&=
\E^\pi\!\left[
\P^\pi\{F\in C_{\mathrm B}(W)\mid W\}
\right]\\
&=0.95.
\end{align*}
The first equality averages the fixed-parameter coverage over the possible
values of $\theta$, using the prior as their distribution. The second is the
tower property under $\P^\pi$, now conditioning on $W$. The final equality
uses the assumed 95% posterior predictive probability at each observed
window.

Thus Bayesian calibration gives 95% coverage *on average over the prior*. An
average can equal 0.95 while individual values $c_{\mathrm B}(\theta)$ differ
substantially from 0.95. It does not follow that the prediction rule has 95%
coverage at each fixed $\theta$.

A finite-sample prediction rule $C_{\mathrm H}(W)$ is *honest over*
$\Theta$ at level 95% if
$$
\inf_{\theta\in\Theta}
\P_\theta\{F\in C_{\mathrm H}(W)\}
\geq 0.95.
$$
In words, its repeated-sample coverage is at least 95% at every parameter
value in the declared model class. This is a worst-case requirement, not a
prior average. It still averages over repeated versions of $W$ and $F$ with
$\theta$ fixed; it is not a conditional statement after one particular
$W=w$ has been observed.

A posterior predictive set can also be honest. Posterior predictive
probability alone does not prove that it is: the fixed-parameter coverage
function must be checked separately.

### A Simple Predictive Comparison

**Example 4.4 (Prior-average coverage can hide poor fixed-parameter
coverage).** Think of $W$ as one observed increment and $F$ as the next
increment in Example 4.3. Take
$$
\delta\sim N(0,1),
\qquad
W=\delta+Z_1,
\qquad
F=\delta+Z_2,
$$
where $Z_1,Z_2\overset{\mathrm{iid}}{\sim}N(0,1)$ and are independent of
$\delta$. Under the resulting hierarchical law $\P^\pi$, normal conjugacy
gives
$$
\delta\mid W=w
\sim N\!\left(\frac w2,\frac12\right),
\qquad
F\mid W=w
\sim N\!\left(\frac w2,\frac32\right).
$$
Let $z_{0.975}$ denote the 97.5th percentile of the standard normal
distribution. After averaging over the posterior of $\delta$, the 95%
posterior predictive interval for the next increment is
$$
C_{\mathrm B}(w)
=
\frac w2\pm z_{0.975}\sqrt{\frac32}.
$$

Now hold the drift fixed rather than drawing it from the prior. Under
$\P_\delta$,
$$
F-\frac W2
=
\frac\delta2+Z_2-\frac{Z_1}{2},
$$
and therefore
$$
F-\frac W2
\sim
N\!\left(\frac\delta2,\frac54\right).
$$
Writing $a:=z_{0.975}\sqrt{3/2}$ and letting $\Phi$ denote the standard
normal distribution function, the fixed-drift coverage is
$$
c_{\mathrm B}(\delta)
=
\Phi\!\left(\frac{a-\delta/2}{\sqrt{5/4}}\right)
-
\Phi\!\left(\frac{-a-\delta/2}{\sqrt{5/4}}\right).
$$
For example,
$$
c_{\mathrm B}(0)\approx0.968,
\qquad
c_{\mathrm B}(2)\approx0.894,
\qquad
c_{\mathrm B}(4)\approx0.640.
$$
The interval has exactly 95% posterior predictive probability for every
observed $w$, and exactly 95% coverage after averaging $\delta$ over its
$N(0,1)$ prior. It is nevertheless not an honest 95% prediction interval over
$\R$. Its fixed-drift coverage is above 95% near the prior mean and falls
below 95% sufficiently far from it.

For comparison, consider
$$
C_{\mathrm H}(W)
=
W\pm z_{0.975}\sqrt2.
$$
Under every fixed law $\P_\delta$,
$$
F-W=Z_2-Z_1\sim N(0,2).
$$
Therefore
$$
\P_\delta\{F\in C_{\mathrm H}(W)\}=0.95
\qquad\text{for every }\delta\in\R.
$$
Both $C_{\mathrm B}$ and $C_{\mathrm H}$ predict the same future increment
$F$. The second interval is honest over the independent, unit-variance
Gaussian fixed-parameter family just specified. Its half-width is about
$2.77$,
compared with about $2.40$ for the Bayesian predictive interval. The Bayesian
interval uses the prior to shrink its center and shorten its width; the honest
interval makes a uniform fixed-drift guarantee.

Neither guarantee is model-free. The Bayesian statement relies on both the
fixed-parameter family and the prior. The honest statement ranges over every
parameter value, but only inside the stated Gaussian family with known
variance. Neither one establishes that this family generated the data.
Under additional regularity, posterior predictive intervals may attain
approximately 95% coverage at each fixed parameter as the sample grows.
Uniform honesty over $\Theta$ is stronger and requires a uniform result;
neither conclusion follows from the posterior formula alone.

## The Bayesian Trade-off

**What it buys.** Bayesian forecasting does not replace the unknown parameter
by one fitted value and then pretend that value is known. It averages the
fixed-parameter forecasts over the posterior:
$$
p^\pi(f\mid w)
=
\int_\Theta p_\theta(f\mid w)\pi(\theta\mid w)\,d\theta.
$$
The resulting predictive law carries parameter uncertainty into forecasts of
future observations and paths. Point forecasts, prediction intervals,
predictive densities, and path-event probabilities can all be derived from
this law.

**What it requires.** The analyst must specify:

- a family of fixed-parameter laws for the observations and future, whose
  observed-data densities supply the likelihood; and
- a prior for the unknown parameter.

This is a stronger commitment than specifying only a few moments. Predictions
may depend materially on the model and prior, especially when the observed
path is weakly informative. Bayesian updating handles parameter uncertainty
within the chosen model; it does not establish that the model or prior is
appropriate. Model criticism and sensitivity analysis remain necessary.

**What it may cost.** Example 4.3 has a closed-form solution because the
normal prior is conjugate to the Gaussian likelihood. Conjugacy is convenient
but special. In other models, the posterior predictive integral must be
approximated using numerical integration or simulation, such as Markov chain
Monte Carlo. The quality of that approximation must then be checked.

**The trade-off:** Bayesian forecasting provides a unified treatment of
parameter uncertainty in prediction, at the price of a fuller probability
model, a prior, and sometimes substantial computation.

## Looking Ahead

The posterior and predictive laws are defined for every finite $T$ whenever
the required densities exist. Their existence does not show that the posterior
concentrates or that prior sensitivity disappears as $T$ grows. The workbook
on one-path learnability asks which population features time averages can
recover; the workbook on identifiability asks whether those features determine
a unique parameter.

## Exercises


::: {.exercise}
**Exercise 4.1 (A four-increment forecast).** Let $\sigma^2=1$ and
$\delta\sim N(0,1)$. Starting from $x_0=0$, suppose the four observed
increments are $1.2,-0.4,0.7,0.1$.

*Method and report.* Derive the quantities by hand; a calculator or R may be
used for the final arithmetic and the normal quantile. Report $m_4$, $s_4^2$,
the predictive mean and variance, the endpoints of both intervals to two
decimal places, and the omitted variance term.

(a) Compute $m_4$ and $s_4^2$.

(b) Compute the posterior predictive mean and variance of $X_9$.

(c) Compute a pointwise 95% posterior predictive interval for $X_9$.

(d) Compute the corresponding plug-in interval using $m_4$ as if it were the
known drift. Identify the omitted variance term.
:::



::: {.exercise}
**Exercise 4.2 (An uncertain AR coefficient).** For each $\phi\in\R$, let
$\P_\phi$ be the law under which
$$
X_{t+1}=\phi X_t+Z_{t+1},
\qquad
Z_t\overset{\mathrm{iid}}{\sim}N(0,\sigma^2),
$$
where $X_0=x_0$ is fixed and $\sigma^2$ is known. Suppose a prior $\pi$ has
produced the posterior density $\pi(\phi\mid X_{0:T})$; $\P^\pi$ denotes the
resulting hierarchical law.

*Method and report.* Work by hand. Give the one- and two-step predictive means
and variances as symbolic expressions in $x_T$, $\sigma^2$, and the required
posterior moments of $\phi$. End with one or two sentences identifying which
posterior moments a plug-in forecast omits. No R output is required.

(a) Express the one-step posterior predictive mean in terms of the first
posterior moment of $\phi$.

(b) Use Proposition 4.2 to show that the one-step posterior predictive variance
is
$$
\sigma^2+X_T^2\Var^\pi(\phi\mid X_{0:T}).
$$

(c) Derive the two-step posterior predictive mean and variance using posterior
moments of $\phi$. Do not assume that the posterior mean of $\phi^2$ equals
the square of the posterior mean of $\phi$.

(d) Explain why a posterior mean plug-in forecast need not even preserve the
posterior predictive mean at horizons greater than one.
:::



::: {.exercise}
**Exercise 4.3 (One drift draw per path).** Under Example 4.3, define
$D_j:=X_{T+j}-X_{T+j-1}$.

*Method and report.* Work by hand. Report the predictive mean, variance, and
cross-covariance; explain the resulting dependence in complete sentences; and
give an ordered simulation algorithm that distinguishes draws made once per
path from draws made once per horizon. Finish by explaining how the simulated
paths answer a path-level probability question. No R output is required.

(a) Find the posterior predictive mean and variance of $D_j$.

(b) Show that, for $j\ne k$,
$$
\Cov^\pi(D_j,D_k\mid X_{0:T})=s_T^2.
$$

(c) Explain why the future increments are independent under $\P_\delta$ but
dependent after the fixed-law forecasts are averaged over the posterior for
$\delta$.

(d) Write an algorithm for simulating the joint vector
$(X_{T+1},\ldots,X_{T+H})$. State where a new random draw is made and where the
same draw is retained.

(e) Explain how repeated joint path draws could approximate
$$
\P^\pi\!\left(\max_{1\le h\le H}X_{T+h}>c\mid X_{0:T}\right),
$$
and why a collection of pointwise predictive intervals does not answer this
question.
:::



::: {.exercise}
**Exercise 4.4 (Two predictive calibration standards).** Let
$C_{\mathrm B}(W)$ be an exact $1-\alpha$ posterior predictive set for a
future target $F$.

*Method and report.* Work by hand and keep every probability law or
conditioning event explicit. Report the posterior predictive probability,
the iterated-expectation proof, the fixed-parameter coverage function, and a
short explanation of the honesty requirement. No R output is required.

(a) Write the posterior predictive probability statement that defines
$C_{\mathrm B}(W)$. After conditioning on $W$, which sources of uncertainty
remain in this probability?

(b) Use iterated expectation to prove that $C_{\mathrm B}(W)$ has
$1-\alpha$ coverage under $\P^\pi$. State which
quantities this probability averages over.

(c) Define the fixed-parameter coverage function
$$
c_{\mathrm B}(\theta)
:=
\P_\theta\{F\in C_{\mathrm B}(W)\}
$$
and express the result of part (b) as a prior-weighted average of
$c_{\mathrm B}(\theta)$.

(d) State the condition that would make $C_{\mathrm B}(W)$ honest over
$\Theta$. Explain why the identity in part (c) does not establish that
condition.

(e) Explain why even an honest prediction interval over $\Theta$ is not a
model-free guarantee.
:::



::: {.exercise}
**Exercise 4.5 (What would make the prior fade?).** In Example 4.3, keep
$m_0,s_0^2$, and $\sigma^2$ fixed while $T$ grows.

*Method and report.* Work by hand. Show the two limits, state the law of large
numbers with the random variables to which it applies, and end with two or
three sentences explaining what the posterior algebra alone cannot establish.
No R output is required.

(a) Show that $s_T^2\to0$.

(b) Show that $m_T-\overline{\Delta X}_T\to0$ in probability under
$\P^\pi$.

(c) Under $\P_\delta$, state the law of large numbers that makes
$\overline{\Delta X}_T\to\delta$.

(d) Explain why the algebraic existence of the posterior alone would not have
proved the conclusion in part (c).
:::



::: {.exercise}
**Exercise 4.6 (How fast is the land warming?).** The series
`astsa::gtemp_land` contains annual land-surface temperature anomalies
in degrees centigrade, 1850--2023. Restrict to 1973 onward, so that
$x_0$ is the 1973 value and the increments are the $T=50$ annual
changes through 2023, and apply Example 4.3 with $\sigma$ treated as
known at the sample standard deviation of the increments.

*Method and report.* Use R for every numerical calculation and for the plot.
Derive the three increment formulas in (d) by hand. Submit one time plot of the
restricted level series and a numerical summary, organized by parts (a)--(e),
containing every quantity explicitly requested below. Answer the modeling
questions in (c) and (e) in complete sentences; code and numerical output alone
do not answer them.

(a) Plot the restricted series and report $\overline{\Delta x}_T$ and
the sample standard deviation of the increments. Also report $T$, $x_0$,
$x_T$, the total rise $x_T-x_0$, and the ratio of the increment standard
deviation to $|\overline{\Delta x}_T|$. Numerically verify that the total
rise equals $T\overline{\Delta x}_T$.

(b) Take the prior $\delta\sim N(0,0.5^2)$. Report the two precision
terms in $1/s_T^2$, their ratio, $m_T$, $s_T^2$, and
$\P^\pi(\delta<0\mid X_{0:T}=x_{0:T})$. State whether the data precision
dominates the prior precision.

(c) Compute the 95% posterior predictive interval for the 2033 anomaly
($h=10$) and the corresponding plug-in interval. Report the common
forecast center, both pairs of interval endpoints, and the fraction of the
posterior predictive variance contributed by the drift term. In two or three
sentences, explain whether either interval gives an informative forecast of
the 2033 anomaly.

(d) Consider the alternative description
$X_t=a+bt+\varepsilon_t$ with $(\varepsilon_t)$ i.i.d.
$N(0,\sigma_\varepsilon^2)$. Derive the mean, variance, and lag-one
correlation of its increments
$\Delta X_t=b+\varepsilon_t-\varepsilon_{t-1}$ by hand. Use R to compute
the sample correlation of consecutive observed increments, report both
correlations, and state which increment description this diagnostic favors.

(e) Interpret in complete sentences. Both descriptions estimate the
drift by the same number $(x_T-x_0)/T=\overline{\Delta x}_T$. Under
Example 4.3 this average has standard deviation $\sigma/\sqrt T$;
under the description in (d) it telescopes to
$b+(\varepsilon_T-\varepsilon_0)/T$. Evaluate both standard
deviations, using (d) to express $\sigma_\varepsilon$ through the
increment standard deviation, and report their ratio. In one paragraph,
explain why the two descriptions disagree about the precision of the same
estimate and state every assumption on which the posterior probability in
(b) remains conditional. *Hint: under Example 4.3
every annual change is permanent, so noise accumulates; under (d) the
noise in $x_T-x_0$ never accumulates beyond its two endpoints.*
:::


------------------------------------------------------------------------

*Sources for this workbook: West and Harrison (1997), Bayesian Forecasting and
Dynamic Models, 2nd ed., Springer, for posterior prediction; Gelman et al.
(2013), Bayesian Data Analysis, 3rd ed., CRC Press, for posterior predictive
distributions and variance decomposition. Data: R package `astsa`
(D. Stoffer).*
