# Compare two AR(1) solutions driven by the same noise sequence.
set.seed(7)
tt <- -20:40
z <- rnorm(281) # i.i.d. N(0,1) draws for t = -240..40
xs <- as.numeric(stats::filter(z, 0.9, method = "recursive"))[221:281] # X* on -20..40 (long burn-in)
par(bg = "gray92")
plot(tt, xs + 0.9^tt * 3, type = "l", col = 2, xlab = "t", ylab = expression(x[t]),
main = "One equation, two solutions")
lines(tt, xs, col = 4)
legend("topright", legend = c(expression(X^"*"), expression(X^"*" + phi^t %.% 3)),
col = c(4, 2), lty = 1, bty = "n")Workbook 2 — Solutions of Generating Equations
STA 542 · Introduction to Time Series Analysis
\[\def\E{\mathbb{E}} \def\P{\mathbb{P}} \def\R{\mathbb{R}} \def\Z{\mathbb{Z}} \def\N{\mathbb{N}} \def\cH{\mathcal{H}} \def\cP{\mathcal{P}} \def\Var{\operatorname{Var}} \def\Cov{\operatorname{Cov}} \def\Corr{\operatorname{Corr}}\]
Workbook 1 wrote generating equations without establishing that any process satisfies them. A time series model contains process laws, each specified by the joint law of every finite window. We need finite ways to describe these infinite objects. We use two: specify the window laws directly, or specify an equation and determine its solutions.
Mathematical Preliminary: Mean-Square Convergence
Suppose that \(Y,Y_1,Y_2,\ldots\) are random variables with finite second moments. We say that \(Y_n\) converges in mean square to \(Y\) if \[ \E[(Y_n-Y)^2]\longrightarrow 0. \]
The sequence \((Y_n)\) is Cauchy in mean square if \[ \E[(Y_m-Y_n)^2]\longrightarrow 0 \qquad\text{as }m,n\longrightarrow\infty. \] The first definition applies when a candidate limit \(Y\) is already available. The Cauchy condition lets us establish that a limit exists without knowing it in advance, just like the Cauchy condition for convergence of real number sequences.
Lemma 2.1 (Mean-square Cauchy criterion). Let \(Y_1,Y_2,\ldots\) be random variables with finite second moments. The sequence converges in mean square to some random variable if and only if it is Cauchy in mean square. Proof left as an exercise.
Finally, a random series \[ \sum_{j=0}^{\infty}U_j \] converges in mean square if its partial sums \[ S_m=\sum_{j=0}^{m}U_j \] converge in mean square. By Lemma 2.1, it is enough to show that \[ \E[(S_m-S_n)^2]\longrightarrow0 \qquad\text{as }m,n\longrightarrow\infty. \]
Specifying a Model Directly
For an i.i.d. model, one marginal CDF and independence determine every window law. A general time series requires one CDF for every finite set of times, and these CDFs must be consistent. Choosing them separately is therefore impractical.
Gaussian laws provide enough structure to make direct specification manageable.
Example 2.2 (Direct specification, Gaussian). Choose a mean function \(\mu : \Z \to \R\) and a symmetric function \(\Sigma : \Z \times \Z \to \R\). For every \(t_1 < \cdots < t_k\), specify \[ (X_{t_1}, \dots, X_{t_k}) \;\sim\; N\!\Big( \bigl(\mu(t_1), \dots, \mu(t_k)\bigr),\ \bigl[ \Sigma(t_i, t_j) \bigr]_{i,j=1}^{k} \Big) \qquad \text{for all } t_1 < \cdots < t_k . \]
The display defines a distribution only if every matrix \(\Sigma_k := [\Sigma(t_i,t_j)]_{i,j=1}^k\) is nonnegative definite: \(a^\mathsf{T}\Sigma_k a \ge 0\) for every \(a \in \R^k\). Under this condition, the multivariate normal distributions are valid and consistent. Marginalizing a multivariate normal selects the corresponding entries of its mean vector and covariance matrix (check this!). Thus the pair \((\mu,\Sigma)\) determines one Gaussian process law.
In Gaussian-process notation, we write \[ X \sim \operatorname{GP}(\mu,\Sigma), \] and call \(\Sigma\) the covariance kernel. A Bayesian Gaussian-process prior assigns this joint law to an unknown trajectory \(t\mapsto X_t\). This is stronger than merely assigning a Gaussian prior to a finite-dimensional parameter.
Covariance functions of existing processes automatically satisfy the condition. For \(v>0\), examples include \(\Sigma(s,t)=v\mathbf{1}\{s=t\}\) for Gaussian white noise, \(\Sigma(s,t)=v\cos\{\lambda(s-t)\}\) for the trigonometric process of Exercise 1.5, and \(\Sigma(s,t)=v\rho^{|s-t|}\) for \(|\rho|<1\) for the autoregression solved below. A Gaussian model is any chosen collection of valid pairs \((\mu,\Sigma)\).
The condition cannot be omitted. Suppose \(\Sigma(s,t)\) equals \(1\) when \(s=t\), \(0.8\) when \(|s-t|=1\), and \(0\) otherwise. For three consecutive times, \[ \bigl[ \Sigma(t_i, t_j) \bigr] \;=\; \begin{pmatrix} 1 & 0.8 & 0 \\ 0.8 & 1 & 0.8 \\ 0 & 0.8 & 1 \end{pmatrix}, \qquad \det \;=\; 1 - 2 \cdot 0.8^2 \;=\; -0.28 \;<\; 0 , \] so this matrix is not nonnegative definite. No process, Gaussian or otherwise, has the proposed covariances. When correlations vanish beyond lag one, longer windows restrict the neighbor correlation to at most \(1/2\). General criteria for valid covariance functions require the same nonnegative-definiteness check in every dimension. The stationarity section below returns to the Gaussian covariance function. ◊
A second structure restricts memory. A process is Markov of order \(p\) if the conditional law of \(X_t\) given its past depends only on \((X_{t-1},\dots,X_{t-p})\). A time-homogeneous Markov specification gives one transition CDF \[ K(c \mid x_{t-1}, \dots, x_{t-p}) := \P\bigl(X_t \le c \mid X_{t-1}=x_{t-1},\dots,X_{t-p}=x_{t-p}\bigr), \] the same at every time point.
For \(p=1\), the linear Gaussian kernel \[ X_t \mid X_{t-1}=x \;\sim\; N(\phi x,\sigma^2) \] is a specification of this type: \(K(\,\cdot\mid x)\) is the \(N(\phi x,\sigma^2)\) distribution function.
Together with a law for \((X_1,\dots,X_p)\), the transition CDF determines every later window recursively. For \(t>p\), \[ F_{1, \dots, t}(c_1, \dots, c_t) \;=\; \E\Bigl[ \mathbf{1}\{X_1 \le c_1, \dots, X_{t-1} \le c_{t-1}\}\; K(c_t \mid X_{t-1}, \dots, X_{t-p}) \Bigr], \] where the expectation uses the window law already constructed and the final factor is \(K(c_t\mid X_{t-1},\dots,X_{t-p})\). Setting \(c_t=+\infty\) reduces the right-hand side to \(F_{1,\dots,t-1}(c_1,\dots,c_{t-1})\), so the window laws are consistent (check this!).
The other route specifies \(X_t\) as a function of earlier values and an auxiliary random sequence. We now define this object and the laws it determines.
Generating Equations, Defined
Example 2.3 (ARMA(1,1)). Combining the autoregressive process (AR) and the moving average (MA) from the previous workbook, we obtain the ARMA process. The equation \[ X_t \;=\; \phi X_{t-1} + Z_t + \theta Z_{t-1}, \qquad (Z_t) \sim \mathrm{WN}(0, \sigma^2), \] contains two processes. \((X_t)\) models the observed series, whereas \((Z_t)\) is auxiliary. The three terms use the previous value of the process, the current noise, and the previous noise. The white-noise condition specifies means, variances, and correlations; it does not make \(Z_t\) independent of the past. The feedback through \(X_{t-1}\) leaves open whether a process satisfies the equation. ◊
Definition 2.4 (Generating equation). Fix nonnegative integers \(p\) and \(q\). A generating equation, also called a structural equation, consists of
- a rule \(g:\R^{1+q}\times\R^p\to\R\); and
- a noise specification \(\mathcal Q\), a set of candidate process laws for a real-valued process \((Z_t)_{t\in\Z}\).
The equation requires a pair \((X,Z)\) to satisfy \[ X_t \;=\; g\bigl(Z_t, Z_{t-1}, \dots, Z_{t-q},\ X_{t-1}, \dots, X_{t-p}\bigr) \qquad \text{for every } t \in \Z . \] The process \((Z_t)\) is the driving sequence. Its law must belong to \(\mathcal Q\). The integers \(p\) and \(q\) count process lags and noise lags, respectively. The equation is a collection of constraints, and we must determine which pairs \((X,Z)\) satisfy them.
In Example 2.3, \(g(z,z',x)=\phi x+z+\theta z'\), \(p=q=1\), and \(\mathcal Q\) is the set of laws satisfying \(\mathrm{WN}(0,\sigma^2)\).
For \(\R^d\) valued processes, the same definition with the obvious change applies: a vector state \(Y_t\in\R^d\) and a vector-valued rule \(g\). One coordinate is designated as the observed process \(X_t\); the others are auxiliary state variables. For GARCH, \(Y_t=(X_t,V_t)\) and \[ X_t=\sqrt{V_t}\,Z_t,\qquad V_t=\omega+\alpha X_{t-1}^2+\beta V_{t-1}, \] where \((V_t)\) is an auxiliary state sequence.
The equation is a joint statement about \(X\) and \(Z\). A law for \(X\) alone therefore cannot be its solution.
Definition 2.5 (Solution; model defined by an equation). A solution is a joint process \(((X_t,Z_t))_{t\in\Z}\) such that
- the law of \((Z_t)\) belongs to \(\mathcal Q\); and
- with probability one, the displayed equation holds for every \(t\in\Z\).
The model defined by the equation is the set of \(X\)-process laws obtained from all such solutions. For a vector state \(Y_t\), the solution is the joint process \((Y,Z)\), and the model retains the law of the designated observed coordinate.
Existence means that at least one consistent family of finite-dimensional distributions for \((X,Z)\) satisfies both conditions. Uniqueness can refer either to the joint law of \((X,Z)\) or to the induced law of \(X\); we will state which is meant.
The noise specifications we have seen so far form a strict nesting \[ \bigl\{\, \text{i.i.d. } N(0, \sigma^2) \,\bigr\} \;\subsetneq\; \bigl\{\, \text{i.i.d., mean } 0,\ \text{variance } \sigma^2 \,\bigr\} \;\subsetneq\; \bigl\{\, \mathrm{WN}(0, \sigma^2) \,\bigr\} : \] the left set contains one process law, the middle contains one for each admissible marginal, and the right also contains dependent processes. A noise specification is complete if it contains one law and partial if it contains more than one. Refining a partial specification changes the model defined by the equation.
Moving Averages
The moving-average family uses a finite window of the driving sequence and no feedback.
Definition 2.6 (The moving-average family). An MA(\(q\)) equation is a generating equation whose rule uses \(q\) noise lags and no process lags. For coefficients \(\theta_1,\dots,\theta_q\), it requires \[ X_t \;=\; Z_t + \theta_1 Z_{t-1} + \cdots + \theta_q Z_{t-q}, \qquad t \in \Z . \] The noise specification is chosen separately, as in Definition 2.4.
The course uses the one-sided convention: present and past values of the driving sequence. For each candidate law in \(\mathcal Q\), the right-hand side defines \(X_t\) directly and hence determines one joint law for \((X,Z)\). Every MA(\(q\)) equation therefore has one solution per candidate noise law.
Autoregressions
The autoregressive family feeds earlier values of the process back into the current value.
Definition 2.7 (The autoregressive family). An AR(\(p\)) equation is a generating equation whose rule uses \(p\) process lags and no noise lags. For coefficients \(\phi_1,\dots,\phi_p\), it requires \[ X_t \;=\; \phi_1 X_{t-1} + \cdots + \phi_p X_{t-p} + Z_t, \qquad t \in \Z , \] with characteristic polynomial \(\phi(z)=1-\phi_1z-\cdots-\phi_pz^p\). The noise specification is chosen separately, as in Definition 2.4.
Because the right-hand side contains unknown process values, Definition 2.7 does not construct a solution. The order-one case shows how existence and uniqueness must be handled.
Example 2.8 (The AR(1), solved). Fix \(0<|\phi|<1\) and let \((Z_t)\sim\mathrm{WN}(0,\sigma^2)\). Iterating \(X_t=\phi X_{t-1}+Z_t\) backward suggests \[ X_t^{\star} \;=\; \sum_{j=0}^{\infty} \phi^{\,j} Z_{t-j} . \]
For \(m<m'\), the partial sums satisfy \[ \E\left[\left(\sum_{m<j\le m'}\phi^jZ_{t-j}\right)^2\right] =\sigma^2\sum_{m<j\le m'}\phi^{2j} \longrightarrow 0 . \] The equality uses the common variance and zero correlations of white noise. Thus the partial sums are Cauchy in mean square. By Lemma 2.1, the series has a mean-square limit for every \(t\).
Substitution gives \[ \phi X_{t-1}^{\star}+Z_t =\phi\sum_{j=0}^{\infty}\phi^jZ_{t-1-j}+Z_t =\sum_{j=0}^{\infty}\phi^jZ_{t-j} =X_t^{\star}. \] Hence \((X^\star,Z)\) is a solution. A solution is causal if each \(X_t\) is a function of \((Z_t,Z_{t-1},\dots)\); \(X^\star\) is causal.
The equation is unique when \(\phi=0\), because then \(X_t=Z_t\). For the nonzero \(\phi\) in Example 2.8, let \(W\) be any random variable defined jointly with the driving sequence. Then \[ X_t \;=\; X_t^{\star} + \phi^{\,t}\, W \] also solves the equation. Conversely, the difference \(H\) of any two solutions satisfies \(H_t=\phi H_{t-1}\), so \(H_t=\phi^tH_0\). These are all the solutions with the given driving process.
Among these solutions, \(X^\star\) is the only one satisfying \[ \sup_{t\in\Z}\E[X_t^2]<\infty. \] Its variance is \(\sigma^2/(1-\phi^2)\). For every nonzero \(W\), the term \(\phi^tW\) is unbounded in mean square as \(t\to-\infty\) (check this!). We use bounded second moments as a selection criterion, which leaves one solution for each candidate noise law. The stationarity section below identifies it as the unique weakly stationary solution. ◊
Remark 2.9 (♠ Beyond the boundary). For \(|\phi|>1\), the backward candidate diverges. The unique solution with bounded second moments is \[ X_t=-\sum_{j\ge1}\phi^{-j}Z_{t+j}, \] which depends on future values of the driving sequence (check that it solves the equation!). This solution is noncausal. Causality is an additional requirement we can impose on solutions. ◊
A causal solution therefore has the form \[ X_t=f(Z_t,Z_{t-1},Z_{t-2},\ldots) \] for a function \(f\) of the present and past driving variables \((Z_t,Z_{t-1},Z_{t-2},\ldots)\).
If \((Z_t)\) is i.i.d., the variables \(X_{t-1},X_{t-2},\dots\) are functions of \(Z_{t-1},Z_{t-2},\dots\), so \(Z_t\) is independent of \((X_{t-1},X_{t-2},\ldots)\). Consequently, when \(\E[Z_t]=0\), \[ \E[Z_t\mid X_{t-1},X_{t-2},\ldots]=0, \qquad \E[X_{t-h}Z_t]=0,\qquad h\ge 1, \] whenever the relevant moments exist. We call this the innovation property. It fails for the noncausal solution in Remark 2.9 because \(X_{t-1}\) contains \(Z_t\).
Under the complete specification \(Z_t\overset{\mathrm{iid}}{\sim}N(0,\sigma^2)\), write \(\psi=(\phi,\sigma^2)\) and let \(\varphi\) denote the standard normal PDF. The selected solution has density \[ p_\psi(x_1,\dots,x_T) =\frac{\sqrt{1-\phi^2}}{\sigma} \varphi\!\left(\frac{\sqrt{1-\phi^2}\,x_1}{\sigma}\right) \prod_{t=2}^{T}\frac{1}{\sigma} \varphi\!\left(\frac{x_t-\phi x_{t-1}}{\sigma}\right). \] The product factors follow from the innovation property.
Under the partial white-noise specification, the coefficients do not determine a density. They do determine the selected solution’s second moments, such as \[ \Var(X_t^\star)=\frac{\sigma^2}{1-\phi^2}, \qquad \Cov(X_t^\star,X_{t-1}^\star) =\frac{\phi\sigma^2}{1-\phi^2}. \] These identities hold for every candidate noise law for \((Z_t)\) (check this!). They support moment-based methods such as the least-squares fit to the hare–lynx series in Workbook 1.
For \(\phi\neq 0\), the AR(1) polynomial \(\phi(z)=1-\phi z\) of Definition 2.7 has the single root \(1/\phi\). That root lies outside the closed unit disk when \(|\phi|<1\), recovering the causal solution of Example 2.8. It lies inside when \(|\phi|>1\), recovering the noncausal solution of Remark 2.9.
For an AR(\(p\)) equation, the polynomial \(\phi(z)=1-\phi_1z-\cdots-\phi_pz^p\) plays the same role. A unique solution with bounded second moments exists when no root of \(\phi\) lies on the unit circle. That solution is causal exactly when every root lies outside the closed unit disk (see Brockwell and Davis, 1991, Proposition 3.1.1, for a proof, ♠).
The ARMA(1,1) equation of Example 2.3 uses the same polynomial \(\phi(z)=1-\phi z\); the coefficient \(\theta\) does not enter the root condition. For \(|\phi|<1\), the causal candidate is \[ X_t=\sum_{j=0}^{\infty}\phi^j \bigl(Z_{t-j}+\theta Z_{t-j-1}\bigr), \] and the proof from Example 2.8 applies (check this!). The ARMA workbook develops the general family; the stability workbook proves the root conditions.
The Random Walk
Remark 2.10 (The random walk is defined only up to a constant). At \(\phi=1\), let \((Z_t)\sim\mathrm{WN}(0,\sigma^2)\). The homogeneous equation is \(H_t=H_{t-1}\), so \[ X_t=X_t^{(0)}+W \] for any particular solution \(X^{(0)}\) and any random variable \(W\) defined jointly with the driving sequence. No member has second moments bounded over \(\Z\): the variance of \(X_t-X_s\) is \(|t-s|\sigma^2\) under white-noise increments. The selection criterion used for \(|\phi|<1\) therefore selects no solution.
We can choose an anchored solution by setting \(X_0=0\) and accumulating increments in both time directions. A different anchor adds a constant to the entire path. Consequently, statements about \[ X_t-X_s=\sum_{s<j\le t}Z_j,\qquad s<t, \] do not depend on the anchor. Workbook 1 used this construction on \(t=1,2,\dots\) with an anchor at the start of the record. ◊
# Compare random walks with common increments and different anchors.
set.seed(542)
tt <- -40:40
z <- rnorm(81)
s <- c(-rev(cumsum(rev(z[1:40]))), 0, cumsum(z[42:81])) # increments accumulated from t = 0
par(bg = "gray92")
plot(NULL, xlim = c(-40, 40), ylim = range(s) + c(-6, 6), xlab = "t", ylab = expression(x[t]),
main = "Random-walk solutions with different anchors")
for (w in c(-5, 0, 5)) lines(tt, s + w, col = c(4, 1, 2)[match(w, c(-5, 0, 5))])
legend("topleft", legend = c("X[0] = -5", "X[0] = 0", "X[0] = 5"),
col = c(4, 1, 2), lty = 1, bty = "n")The equation and the increment law do not determine the distribution of \(X_1\), so they do not determine a full likelihood for \((x_1,\dots,x_T)\). Suppose the increments are i.i.d. with density \(f_\psi\). Under the anchored construction that fixes \(X_1=x_1\) independently of future increments, the conditional likelihood is \[ p_\psi(x_2,\dots,x_T\mid x_1) =\prod_{t=2}^{T}f_\psi(x_t-x_{t-1}). \] Thus conditioning on the first observation and modeling the differences give the same likelihood. The workbook on nonstationarity develops both approaches.
An Equation With No Solution
Example 2.11 (An equation with no solution). Consider \[ x_t \;=\; 1 + a\,\lvert x_{t-1} \rvert, \qquad t \in \Z, \quad a \ge 1 . \] Any solution must satisfy \(x_t\ge1\). Hence \(|x_{t-1}|=x_{t-1}\) and backward iteration gives \[ x_t\ge 1+a+\cdots+a^m \] for every \(m\). When \(a\ge1\), the lower bound diverges with \(m\), so no real-valued sequence satisfies the equation. For \(0<a<1\), the solutions \((1-a)^{-1}+a^tw\) with \(w\ge0\) show that the boundary changes the conclusion. ◊
Now let \(X_t=1+A_t|X_{t-1}|\), where \((A_t)\) is i.i.d., \(A_t>0\), and \(\E[|\log A_t|]<\infty\). Any solution is positive, and backward iteration gives \[ X_t \;\ge\; 1 + A_t + A_t A_{t-1} + A_t A_{t-1} A_{t-2} + \cdots , \] where the logarithm of the \(k\)th product is \(\log A_t+\cdots+\log A_{t-k+1}\). The law of large numbers makes its average converge to \(\E[\log A_1]\). If this expectation is negative, the series converges and its sum solves the equation (see Brandt, 1986, Theorem 1, for a proof, ♠). If the expectation is positive, the lower bound diverges and no solution exists. Thus the geometric mean of the random coefficient determines existence.
Definition 2.12 (GARCH(1,1)). Let \(V_t:=\sigma_t^2\) denote a positive scale state. For \(\omega>0\) and \(\alpha,\beta\ge0\), the GARCH(1,1) system is \[ X_t=\sqrt{V_t}\,Z_t,\qquad V_t=\omega+\alpha X_{t-1}^2+\beta V_{t-1}, \qquad t\in\Z. \] Its noise specification is the set of i.i.d. process laws whose marginal has mean zero and variance one. The unit variance fixes the scale: without it, rescaling \(Z_t\) and \(V_t\) would leave the law of \(X_t\) unchanged.
Example 2.13 (Existence of a GARCH solution). Substituting \(X_{t-1}=\sqrt{V_{t-1}}Z_{t-1}\) gives \[ V_t=\omega+\left(\alpha Z_{t-1}^2+\beta\right)V_{t-1}. \] With \(A_t:=\alpha Z_t^2+\beta\), the candidate is \[ V_t=\omega\left\{1+A_{t-1}+A_{t-1}A_{t-2}+\cdots\right\}. \] For a fixed candidate noise law, this series converges when \(\E[\log(\alpha Z_1^2+\beta)]<0\) and no positive finite solution exists when the expectation is positive (Nelson’s condition; see Nelson, 1990, Theorems 1 and 2, for a proof, ♠).
The expectation depends on the marginal law of \(Z_1\). At \((\alpha,\beta)=(0.3,0.75)\), standard Gaussian noise gives \(\E[\log(0.3Z_1^2+0.75)]\approx-0.0074\), whereas the two-point distribution \(\P(Z_1=-1)=\P(Z_1=1)=1/2\) gives \(\log(1.05)\approx0.0488\). Both marginals have mean zero and variance one, but only the Gaussian choice satisfies the existence condition. Definition 2.12 therefore admits candidate noise laws that lead to solutions and others that do not. ◊
# Compare convergent and divergent GARCH backward sums.
set.seed(542)
bs <- function(a, b, K = 60, omega = 0.05) {
A <- a * rnorm(K - 1)^2 + b
omega * cumsum(c(1, cumprod(A)))
}
left <- replicate(5, bs(0.15, 0.8))
right <- replicate(5, bs(1.2, 0.9))
par(mfrow = c(1, 2), mar = c(4, 4, 2.5, 1), bg = "gray92")
matplot(1:60, left, type = "l", lty = 1, col = 4,
xlab = "K", ylab = "Partial sum",
main = expression(alpha == 0.15 ~ "," ~ beta == 0.8))
matplot(1:60, right, type = "l", lty = 1, col = 2, log = "y",
xlab = "K", ylab = "Partial sum (log scale)",
main = expression(alpha == 1.2 ~ "," ~ beta == 0.9))With i.i.d. \(N(0,1)\) noise, \(V_t\) is a function of past noise and \(Z_t\) is independent of the past. Hence \[ X_t\mid X_{t-1},X_{t-2},\dots\sim N(0,V_t). \] For a finite record, the recursion requires a numerical initialization for \(V_1\) because its exact value depends on the unobserved infinite past. The product of the resulting conditional densities is the working Gaussian likelihood. If the noise marginal remains unspecified, the same product may be used as a quasi-likelihood, but it is not the likelihood of every law in the model.
The examples give four outcomes: a unique solution after a stated selection criterion, a family left undistinguished by that criterion, a bounded noncausal solution, and no solution. The stability workbook develops root and contraction conditions that distinguish these cases.
Can the Process Law Be Learned?
Assume that one candidate process law in the model governs the observed series. We still observe only one stretch \(x_1,\dots,x_T\) from one realization. Model membership alone does not make that one path informative.
In an i.i.d. sample, observations are repetitions because they have a common distribution. A time series may be dependent, but its shifted windows can still have the same distribution. Stationarity formalizes this comparability across time.
Definition 2.14 (Stationarity). A process \((X_t)_{t\in\Z}\) is strictly stationary if, for every \(k\), every \(t_1<\cdots<t_k\), and every \(h\in\Z\), \[ (X_{t_1},\dots,X_{t_k}) \overset{d}{=} (X_{t_1+h},\dots,X_{t_k+h}). \] It is weakly stationary if it has finite second moments and there are a constant \(\mu\) and a function \(\gamma:\Z\to\R\) such that \[ \E[X_t]=\mu,\qquad \Cov(X_t,X_{t+h})=\gamma(h) \qquad\text{for every }t,h\in\Z. \] The function \(\gamma\) is the autocovariance function (ACVF).
Strict stationarity makes the law of every finite window invariant to a common time shift. Weak stationarity asks only for invariant means and covariances. Strict stationarity with finite second moments implies weak stationarity. For a Gaussian process, the converse also holds because its window laws are determined by their means and covariance matrices.
Example 2.15 (Stationarity of the earlier processes).
- Direct Gaussian specification. The process in Example 2.2 is stationary when \(\mu(t)=\mu\) and \(\Sigma(s,t)=\gamma(t-s)\). For a Gaussian process, these moment conditions determine all shifted window laws.
- Moving average. If the driving sequence is i.i.d., an MA(\(q\)) process is strictly stationary: shifting time applies the same window rule to a shifted i.i.d. sequence. Under a white-noise specification alone, the MA(\(q\)) is weakly stationary.
- Stable autoregression. Under i.i.d. noise, the causal solution \(X_t^\star=\sum_{j\ge0}\phi^jZ_{t-j}\) is strictly stationary. Under white noise it is weakly stationary. Among solutions with finite second moments, stationarity selects \(X^\star\), since every weakly stationary solution has second moments bounded over \(\Z\).
- Random walk. For the solution anchored at \(X_0=0\), \(\Var(X_t)=|t|\sigma^2\), so the level process is not weakly stationary. Its increments \(X_t-X_{t-1}=Z_t\) inherit the stationarity of the driving sequence.
The equation in Example 2.11 has no solution, so it supplies no process whose stationarity can be checked. ◊
Example 2.16 (Stationary without repetition). Let \(Y\) have mean \(\mu\) and variance \(\eta^2>0\), and set \(X_t:=Y\) for every \(t\). Every shifted window is the same random vector, so the process is strictly stationary. Nevertheless, \[ \bar X_T:=\frac{1}{T}\sum_{t=1}^T X_t=Y \qquad\text{for every }T. \] One path repeats one draw from the distribution of \(Y\); in particular, \(\bar X_T\) does not converge to \(\mu\). Its ACVF is \(\gamma(h)=\eta^2\) at every lag. ◊
Stationarity makes observations comparable across time. We also need distant observations to supply new information.
Definition 2.17 (Mean ergodicity). A weakly stationary process is mean ergodic if \[ \E\bigl[(\bar X_T-\mu)^2\bigr]\longrightarrow0. \] Mean ergodicity says that the sample mean converges to \(\mu\) in mean square, and hence in probability.
Proposition 2.18 (Covariance decay makes the mean learnable). If a weakly stationary process satisfies \(\gamma(h)\to0\) as \(|h|\to\infty\), then it is mean ergodic.
Proof. Weak stationarity gives \[ \begin{aligned} \Var(\bar X_T) &=\frac{1}{T^2}\sum_{s=1}^T\sum_{t=1}^T\gamma(t-s)\\ &=\frac{1}{T}\left[ \gamma(0)+2\sum_{h=1}^{T-1} \left(1-\frac{h}{T}\right)\gamma(h)\right]. \end{aligned} \] The first term tends to zero. For the second, \[ \left|\frac{2}{T}\sum_{h=1}^{T-1} \left(1-\frac{h}{T}\right)\gamma(h)\right| \le \frac{2}{T}\sum_{h=1}^{T-1}|\gamma(h)| \longrightarrow0, \] because the averages of the convergent sequence \(|\gamma(h)|\to0\) also converge to zero. Thus \(\E[(\bar X_T-\mu)^2]=\Var(\bar X_T)\to0\). \(\square\)
Example 2.19 (Which means can be learned?). For an MA(\(q\)), the ACVF vanishes beyond lag \(q\), so Proposition 2.18 applies. For the stable AR(1), \[ \gamma(h)=\frac{\sigma^2\phi^{|h|}}{1-\phi^2}, \] which tends to zero geometrically. The repeated-level process of Example 2.16 fails the condition and satisfies \(\bar X_T=Y\). The random walk is outside the proposition because it is not stationary. ◊
Mean ergodicity answers one learning question about one feature of the law. General ergodicity concerns time averages of broader functions of the process. Whether distinct parameter values determine distinct process laws is the separate problem of identifiability. Both require more than the stationarity definition itself.
Looking Ahead
The workbook on forecasting under a known process law takes the laws in this workbook as given and asks what they imply about unobserved values. Later work returns to the statistical problem: whether one observed path can reveal the law and its parameters.
Exercises
Exercise 2.1 (Verifying AR(1) solutions). Let \((Z_t)\) be i.i.d. with mean zero and variance \(\sigma^2>0\), fix \(0<|\phi|<1\), and consider \[ X_t=\phi X_{t-1}+Z_t,\qquad t\in\Z. \] For \(m\ge0\), define \[ X_t^{(m)}=\sum_{j=0}^{m}\phi^jZ_{t-j}, \qquad X_t^\star=\sum_{j=0}^{\infty}\phi^jZ_{t-j}, \] where Example 2.8 shows that the second series converges in mean square.
Compute the residual \(X_t^{(m)}-\phi X_{t-1}^{(m)}-Z_t\) and show that it converges to zero in mean square.
Verify that \(X_t^\star+\phi^tW\) solves the equation for any random variable \(W\) defined jointly with \((Z_t)\). Show that the difference \(H_t:=X_t^\star-\tilde{X}_t^\star\) of any two solutions \(X^\star\) and \(\tilde{X}^\star\) with the same driving sequence satisfies \(H_t=\phi H_{t-1}\). Show that only \(W=0\) can yield second moments bounded over \(\Z\).
Let \(Y_t=\sum_{j\ge0}\phi^jZ_{t+j}\). Compute \(Y_t-\phi Y_{t-1}\) and show that \(Y\) does not solve the AR(1) equation.
Explain why an infinite-series candidate requires two separate arguments: convergence of the series and verification of the equation.
Exercise 2.2 (One AR(1) equation, two noise specifications). For each noise specification below, select the causal solution of \[ X_t=0.8X_{t-1}+Z_t,\qquad t\in\Z. \] Compare (i) \(Z_t\overset{\mathrm{iid}}{\sim}N(0,1)\) and (ii) i.i.d. \(Z_t\) with mean zero and variance one.
Under (i), state the conditional distribution of \(X_t\) given \((X_{t-1},X_{t-2},\ldots)\). Under (ii), state the corresponding conditional mean and variance and explain why the conditional distribution is not fixed.
Use standard normal noise and two-point noise with \(\P(Z_t=-1)=\P(Z_t=1)=1/2\). Simulate the causal solution under each (\(T=400\), seed
542, burn-in 100). Plot both paths and both histograms.
# Simulate AR(1) paths under two driving distributions.
set.seed(542)
# z_gaussian <- rnorm(500)
# z_twopoint <- sample(c(-1, 1), 500, replace = TRUE)
# x_gaussian_all <- as.numeric(
# stats::filter(z_gaussian, 0.8, method = "recursive")
# )
# x_twopoint_all <- as.numeric(
# stats::filter(z_twopoint, 0.8, method = "recursive")
# )
# x_gaussian <- x_gaussian_all[-(1:100)]
# x_twopoint <- x_twopoint_all[-(1:100)]Report the sample mean, variance, and lag-one covariance for each path. Explain which similarities follow from the common first two moments of the noise.
Explain the set relation between the two sets of induced \(X\)-process laws. State one claim valid under both specifications and one valid only under the Gaussian specification.
Exercise 2.3 (The walk and its anchors). Consider \(X_t=X_{t-1}+Z_t\) on \(\Z\), where \((Z_t)\) is i.i.d. with mean zero and variance \(\sigma^2>0\).
Show that every solution has the form \(X_t=X_t^{(0)}+W\), where \(X^{(0)}\) is anchored at zero. Show that no solution satisfies \(\sup_{t\in\Z}\E[X_t^2]<\infty\). Hint: use \(\E[(X_t-X_s)^2]=|t-s|\sigma^2\) and \((a-b)^2\le2a^2+2b^2\).
For \(t>0\), verify the following statements under the zero anchor. Then determine which hold for every allowed anchor \(W\), including random anchors, and which can change or cease to be defined: \[ \E[X_t]=0,\qquad \Var(X_t)=t\sigma^2,\qquad X_t-X_{t-1}\sim Z_1,\qquad \Corr(X_t,X_{2t})=1/\sqrt2. \]
Simulate three solutions with common i.i.d. \(N(0,1)\) increments, \(t=-40,\dots,40\), and anchors \(-5\), \(0\), and \(5\). Compare their pairwise differences with those of the AR(1) paths following Example 2.8.
# Construct random walks with common increments and different anchors.
set.seed(542)- Explain why bounded second moments select a solution for the AR(1) equation when \(|\phi|<1\) and select none when \(\phi=1\). Show that first differencing, \(\Delta X_t:=X_t-X_{t-1}\), turns the random-walk equation into \(\Delta X_t=Z_t\).
Exercise 2.4 (♠ Backward iteration and GARCH existence). Part (a) is a deterministic warm-up. In parts (b)–(e), assume \(Z_t\overset{\mathrm{iid}}{\sim}N(0,1)\) and use the existence criterion stated in Example 2.13; no proof of Nelson’s condition is required.
For \(x_t=1+a|x_{t-1}|\) on \(\Z\), derive \(x_t\ge1+a+\cdots+a^m\) and conclude that no solution exists when \(a\ge1\). For \(0<a<1\), verify directly that \[ x_t=\frac{1}{1-a}+a^tw,\qquad w\ge0, \] is a solution.
Using \(10^6\) standard normal draws and seed 542, estimate \(\E[\log(\alpha Z_1^2+\beta)]\) for \((\alpha,\beta)=(0.15,0.8)\), \((0.3,0.75)\), and \((1.2,0.9)\). Report the three estimates and their signs in a table. For purposes of applying the existence criterion, take the three population signs to be negative, negative, and positive, respectively. Identify the parameter pairs for which the criterion asserts existence of a positive finite solution.
Set \(A_s:=\alpha Z_s^2+\beta\) and define the \(K\)-term truncation of the backward series for \(V_t\) by \[ S_{K,t}=\omega\left\{1+\sum_{k=1}^{K-1} \prod_{j=1}^k A_{t-j}\right\}. \] For each parameter pair, use \(\omega=0.05\) to simulate 20 independent paths of \((S_{K,t}:K=1,\ldots,80)\). Plot them, using a logarithmic vertical scale for the last two parameter pairs, and compare their finite-\(K\) behavior with part (b). Explain why the plots illustrate, but do not prove, convergence or divergence.
# Compute GARCH backward sums from simulated multipliers.
set.seed(542)
# A <- alpha * rnorm(79)^2 + beta
# S <- 0.05 * cumsum(c(1, cumprod(A)))Iterate \[ V_t=\omega+(\alpha Z_{t-1}^2+\beta)V_{t-1} \] backward \(K\) times. Show that \(V_t\ge S_{K,t}\). Deduce that \(S_{K,t}\to\infty\) rules out every positive finite solution. Explain why the same monotone-lower-bound argument does not apply to a random walk.
Continue with \(\omega=0.05\). For \((\alpha,\beta)=(0.3,0.75)\), set \(V_1=\omega\), simulate \(20{,}000\) observations after a burn-in of \(5{,}000\), and plot the running mean of \(X_t^2\). Then suppose, for contradiction, that the stationary solution satisfies \(\E[V_t]<\infty\). Take expectations in the variance recursion and show why \(\alpha+\beta=1.05\) rules this out. Conclude that existence need not imply a finite second moment.
Exercise 2.5 (Stationarity and one-path learning). Let \((Z_t)\) be i.i.d. with mean zero and variance \(\sigma^2>0\), let \(Y\) have mean \(\mu\) and variance \(\eta^2>0\), and fix \(|\phi|<1\). Consider \[ \begin{aligned} M_t&=Z_t+\theta Z_{t-1},& U_t&=\sum_{j=0}^{\infty}\phi^jZ_{t-j},\\ R_t&=Y,& L_t&=L_{t-1}+Z_t,\qquad L_0=0. \end{aligned} \] Here \(U\) is the causal AR(1) process of Example 2.8.
Determine which processes are strictly stationary and which are weakly stationary. Justify each answer from Definition 2.14.
Compute the ACVF of \(R\). For \(M\), show that \[ \gamma_M(0)=\sigma^2(1+\theta^2),\qquad \gamma_M(1)=\gamma_M(-1)=\theta\sigma^2, \] and \(\gamma_M(h)=0\) for \(|h|>1\).
Use Proposition 2.18 and the AR(1) ACVF in Example 2.19 to determine which of \(M\), \(U\), and \(R\) are mean ergodic. Explain why the proposition does not apply to \(L\).
Interpret the comparison in complete sentences: what does stationarity provide, what additional conclusion does mean ergodicity provide, and why does the process \(R_t=Y\) separate the two ideas?
Sources for this workbook: van der Vaart (2010), Time Series, lecture notes, VU Amsterdam, §1.1 and Lemma 1.28; Brockwell and Davis (1991), Time Series: Theory and Methods, 2nd ed., Springer, Ch. 3; Brandt (1986), “The stochastic equation \(Y_{n+1}=A_nY_n+B_n\) with stationary coefficients,” Advances in Applied Probability 18, for the random-coefficient recursion; Nelson (1990), “Stationarity and persistence in the GARCH(1,1) model,” Econometric Theory 6, for the ♠ existence condition; Diaconis and Freedman (1999), “Iterated random functions,” SIAM Review 41, for backward iteration in general.