3. Processes with Stationary Increments#

Download PDF here

Authors: Lars Peter Hansen (University of Chicago) and Thomas J. Sargent (NYU)

\(\newcommand{\eqdef}{\stackrel{\text{def}}{=}}\)

../_images/ch3_outset.jpg

3.1. Introduction#

Earlier chapters explain why we like statistical models that generate stationary stochastic processes: stationarity brings a Law of Large Numbers that helps us identify statistical models and make inferences about model parameters. However, logarithms of many economic time series appear not to be stationary. Instead they grow systematically. This situation motivates us to study models that generate stochastic processes with stationary increments. Multivariate versions of such models possess stochastic process versions of balanced growth paths. Applied econometricians sometimes study permanent shocks that contribute to stochastic growth. We shall describe how to pose a central limit theory and to characterize its implications in terms of processes with stationary increments.

The mathematical formulation in this chapter opens the door to studying these topics using a unified set of tools. In this chapter we return to the mathematical formulation used in Chapter 1, while in the next chapter we will assume a Markov structure.

3.2. Basic setup#

We adopt assumptions from Section Inventing an Infinite Past of Chapter 1 that allow an infinite past. Again let \( {\mathfrak A}\) be a subsigma algebra of \({\mathfrak F}\) and

\[{\mathfrak A}_t = \left\{ \Lambda_t \in {\mathfrak F} : \Lambda_t = \{ \omega \in \Omega : {\mathbb S}^t(\omega) \in \Lambda \} \textrm{ for some } \Lambda \in {\mathfrak A} \right\} .\]

The event collection \({\mathfrak A}\) includes invariant events as well as past information. Thus we assume that

\[{\mathfrak A}_{t} \subset {\mathfrak A}_{t+1},\]

so information accumulates over time. Formally, \(\{ {\mathfrak A}_t \}\) is called a filtration, which captures the dynamic information flow.

Let \(X\) be a scalar measurement function that is \({\mathfrak A} = {\mathfrak A}_0\) measurable. Assume that \(Y_0\) is \({\mathfrak A}_0\) measurable, and consider a scalar process \(\{Y_t : t=0,1,... \}\) with stationary increments \(\{X_t\}\):

(3.1)#\[Y_{t+1} - Y_t = X_{t+1}\]

for \(t=0,1, \ldots\). Notice that both \(X_{t+1}\) and \(Y_{t+1}\) are in the implied date \(t+1\) collection of information. Let

\[\eta = E\left(X_{t+1} \vert {\mathfrak I} \right),\]

and

\[U_{t+1} = X_{t+1} - E\left(X_{t+1} \vert {\mathfrak A}_t \right).\]

We interpret the above equations as providing two contributions to the \(\{Y_{t}: t \ge 0\}\) process. Component \(U_{t+1}\) is unpredictable and represents new information about \(Y_{t+1}\) that arrives at date \(t+1\). Component \(\eta\) is the trend rate of growth or decay in \(\{Y_{t} : t \ge 0\}\) conditioned on the invariant events. In the following sections, we present a full decomposition of a stationary increment process that will be useful both in connecting to sources of permanent versus transitory shocks and to central limit theorems.

3.3. A martingale decomposition#

A special class of stationary increment processes consisting of what we call additive martingales interests us.

Definition 3.1

The process \(\{Y_t^m : t=0,1,... \}\) is said to be an additive martingale relative to \(\{ {\mathfrak A}_{t} : t=0,1,... \}\) if for \(t=0,1,... \)

  • \(Y_t^m\) is \({\mathfrak A}_{t}\) measurable, and

  • \(E\left(Y_{t+1}^m \vert {\mathfrak A}_t \right) = Y_t^m\) .

Notice that by the Law of Iterated Expectations, for a martingale \(\{Y_{t}^m : t \ge 0\}\), best forecasts satisfy:

\[E \left (Y_{t+j}^m \mid {\mathfrak A}_t \right) = Y_t^m\]

for \(j \ge 1\). Under suitable additional restrictions on the increment process \(\{X_t : t \ge 0 \}\), we can deploy a construction of [Gordin, 1969] to build a martingale component of the \(\{Y_t^m : t=0,1, ... \}\) process.[1] Let \({\mathcal H}\) denote the set of all scalar random variables \(X\) such that \(E(X^2) < \infty\) and such that[2]

\[H_t = \sum_{j=0}^\infty E\left( X_{t+j} - \eta \vert {\mathfrak A}_t \right)\]

is well defined as a mean-square convergent series. Convergence of the infinite sum on the right side limits temporal dependence of the process \(\{ X_t \}\). For example, it can exclude so-called long memory processes.[3]

Construct the one-period ahead forecast of \(H_{t+1}\) conditioned on date \(t\) information:

\[H_t^+ = E\left( H_{t+1} \mid {\mathfrak A}_{t} \right).\]

Notice that

\[X_t - \eta = H_t - H_t^+ = G_t + \left( H_{t-1}^+ - H_t^+ \right)\]

where

(3.2)#\[G_{t} \eqdef H_{t} - H_{t-1}^+ = H_t - E\left( H_{t} \mid {\mathfrak A}_{t-1} \right). \]

Since \(G_t\) is a forecast error,

\[E \left( G_{t+1} \vert {\mathfrak A}_{t} \right) = 0.\]

Assembling these parts, we have

(3.3)#\[Y_{t+1} - Y_t = X_{t+1} = \eta + G_{t+1} + H_t^+ - H_{t+1}^+ .\]

Let

\[Y^m_t \eqdef \sum_{j=1}^t G_j .\]

Since \(Y_t^m\) is \({\mathfrak A}_{t}\) measurable, the equality

\[E \left( \sum_{j=1}^{t+1} G_j \mid {\mathfrak A}_t \right) = \sum_{j=1}^{t} G_j \]

implies that the process \(\{Y_t^m : t \ge 0 \}\) is an additive martingale.

For a given stationary increment process, \(\{Y_t : t \ge 0\}\), express the martingale increment as

(3.4)#\[\begin{split} G_{t} &= \sum_{j=0}^\infty \left[ E\left( X_{t+j} \mid {\mathfrak A}_{t} \right) - E\left( X_{t+j} \mid {\mathfrak A}_{t-1} \right) \right] \cr &= \lim_{j \rightarrow \infty} \left[ E\left(Y_{t+j} \vert {\mathfrak A}_{t} \right) - E\left(Y_{t+j} \vert {\mathfrak A}_{t-1} \right) \right] . \end{split}\]

So the increment to the martingale component of \(\{Y_t : t \ge 0 \}\) provides new information about the limiting optimal forecast of \(Y_{t+j}\) as \(j \rightarrow + \infty\).

By accumulating equation (3.3) forward, we arrive at:

Proposition 3.1

If \(X\) is in \({\mathcal H}\), the stationary increments process \(\{Y_t : t=0,1,...\}\) satisfies the additive decomposition

\[\begin{matrix} Y_{t} & = & \underbrace{t\eta} & + & Y_t^m & - &\underbrace{ H_t^+} & + & \underbrace{Y_0 + H_0^+}.\cr &&\textbf{trend} &&\textbf{martingale}&& \textbf{stationary} && \textbf{invariant} \end{matrix}\]

where we impose the normalization, \(Y_0^m = 0.\) The two stationary increment processes, \(\{ t\eta : t \ge 0\}\) and \(\{Y_{t}^m : t\ge 0 \},\) are deterministic and martingale components of the original process. The contribution, \(\{H_{t}^+\},\) is a stationary component. The other components are constant over time.

Proposition 3.1 decomposes a stationary-increment process into a linear time trend, a martingale, and a transitory component. A permanent shock is the increment to the martingale. The martingale and transitory contributions are typically correlated. Some decomposition methods go one step further by adjusting the decomposition to remove the correlation between these two components as we will illustrate in an example that follows.

With this mathematical structure in place, we construct an operator \(\mathbb{Md}\) that maps an admissible increment process in \(\mathcal{H}\) into the innovation in a martingale component. Let \(\mathcal{G}\) be the set of all random variables \(G\) with finite second moments that satisfy the conditions that i) \(G\) is \(\mathfrak{A}\) measurable and that ii) \(E(G_1 \vert \mathfrak{A}) = 0\) where \(G_t = G \circ \mathbb{S}^t\). Define \(\mathbb{Md}: \mathcal{H} \rightarrow \mathcal{G}\)

\[ \mathbb{Md} (X) = G \]

for \(G = G_0\) given by (3.4) for \(t = 0\). Both \(\mathcal{G}\) and \(\mathcal{H}\) are linear spaces of random variables and \(\mathbb{Md}\) is a linear transformation. The operator \(\mathbb{Md}\) plays a prominent role in some of the analysis that follows. Chapter 4 supplies a constructive algorithm for computing \(\mathbb{Md}\) when the increments are generated by a stationary Markov process.

3.4. Permanent shocks#

In this construction, we impose a moving-average structure on the underlying time series.
Specifically, consider again the Example 1.8 moving-average process:

(3.5)#\[X_{t} = \sum_{j=0}^\infty \alpha_j \cdot W_{t-j} .\]

Use this \(\{X_t\}\) process as the increment for \(\{ Y_t : t \ge 0 \}\) in formula (3.1). New information about the unpredictable component of \(X_{t+j}\) for \(j \ge 0\) that arrives at date \(t\) is

\[E \left( X_{t+j} \mid {\mathfrak A}_{t} \right)- E \left( X_{t+j} \mid {\mathfrak A}_{t-1} \right)= \alpha_{j} \cdot W_{t}\]

Summing these terms over \(j\) gives

\[G_{t} = \alpha(1) \cdot W_{t} \]

where

\[\alpha(1) = \sum_{j=0}^\infty \alpha_j\]

provided that the coefficient sequence \(\{ \alpha_j : j\ge 0\}\) is summable, a condition that restricts temporal dependence of the increment process \(\{X_t\}\). Indeed, it is possible for \(\alpha(1) = \infty\) or for it not to be well defined while

\[ \sum_{j=0}^\infty |\alpha_j|^2 < \infty\]

ensuring that \(X_t\) is well defined. This possibility opened the door to the literature on long-memory processes that allow for \(\alpha(1)\) to be infinite as discussed in [Granger and Joyeux, 1980] and elsewhere.

In what follows, we presume that \(\alpha(1)\) is finite. This sum of the coefficients \(\{\alpha_j: j\ge 0 \}\) in moving-average representation (3.5) for the first difference \(Y_{t+1} - Y_t = X_{t+1}\) of \(\{ Y_t : t=0,1,.... \}\) reveals the permanent effect of \(W_{t+1}\) on future values of the level of \(Y\), i.e., the effect on \(\lim_{j\rightarrow + \infty} Y_{t+j}\). Models of [Blanchard and Quah, 1989] and [Shapiro and Watson, 1988] build on this property.

The variance of the random variable \(\alpha(1) \cdot W_{t+1}\) conditioned on the invariant events in \({\mathfrak I}\) is \(|\alpha(1)|^2\). The overall variance of \(X_{t}\) is

\[\sum_{j=0}^\infty|\alpha_j|^2 ,\]

where \(|\cdot |\) is the Euclidean norm. The two magnitudes typically differ:

\[\sum_{j=0}^\infty|\alpha_j|^2 \ne |\alpha(1)|^2 .\]

In what follows, we assume that \(\alpha(1) \ne 0\). To form a permanent-transitory shock decomposition, construct the scalar permanent shock as:

\[W_{t+1}^p \eqdef \left( \frac {1}{|\alpha(1)|} \right) \alpha(1) \cdot W_{t+1}\]

where we introduce an additional scaling so the permanent shock has variance one. Form

\[\begin{split} W_{t+1}^{tr} & = W_{t+1} - \left[\frac 1 {\mid \alpha(1) \mid^2} \alpha(1) \right] \alpha(1) \cdot W_{t+1} \cr & = W_{t+1} - \left[\frac 1 {\mid \alpha(1) \mid} \alpha(1) \right] W_{t+1}^p, \end{split}\]

which by construction will be uncorrelated with \(W_{t+1}^p\). The terms in square brackets are the population least squares regression coefficients where the regressors are the variables that follow. The covariance matrix for \(W_{t+1}^{tr}\) is given by the idempotent matrix:

\[{\mathbb I} - \left[\frac 1 {\mid \alpha(1) \mid^2}\right] \alpha(1) \alpha(1)' \]

and has reduced rank. As a consequence, the entries of \(W_{t+1}^{tr}\) can be expressed as linear combinations of a vector of transitory shocks with unit variances and one fewer dimension.

3.5. Central limit approximation#

In this section, we produce a central limit approximation for temporally dependent processes originally due to [Gordin, 1969]. We view Gordin’s result as an application of Proposition 3.1.

To form a central limit approximation, construct the following scaled partial sum that nets out trend growth

\[{\frac 1 {\sqrt{t}}}(Y_t - \eta t) = {\frac 1 {\sqrt{t}}} Y_t^m - {\frac 1 {\sqrt t}} H_{t}^+ + {\frac 1 {\sqrt{t}}} (H_0^+ + Y_0) \]

where

\[Y_t^m= \sum_{j=1}^t G_j\]

From [Billingsley, 1961]’s central limit theorem for martingales

\[{\frac 1 {\sqrt{t}}} Y_t^m \Rightarrow \mathcal{N} \left(0, E\left[ \mathbb{Md}(X)^2 \vert \mathfrak{I} \right] \right)\]

where \(\Rightarrow\) denotes weak convergence, meaning convergence in distribution. Clearly, \(\{(1/ {\sqrt t}) H_{t}^+\}\) and \(\{(1/{\sqrt{t}}) (H_0^+ + Y_0) \}\) both converge in mean square to zero.

Proposition 3.2

For all stationary increment processes \(\{Y_t : t=0,1,2, ...\}\) represented by \(X\) in \(\mathcal{H}\)

\[{\frac 1 {\sqrt{t}}}(Y_t - \eta t) \Rightarrow {\mathcal{N}} \left( 0, E\left[ \mathbb{Md}(X)^2 \vert {\mathfrak{I}} \right] \right) .\]

Furthermore,

\[E\left[ \mathbb{Md}(X)^2 \mid {\mathfrak{I}} \right] = \lim_{t \rightarrow \infty} E \left[ \left({\frac 1 {\sqrt{t}}} \left(Y_t - t \eta \right) \right)^2 \mid {\mathfrak{I}} \right].\]

The variance in Proposition 3.2 is conditioned on the invariant events, so it is a random variable rather than a number. When \({\mathbb S}\) is ergodic, Proposition 1.1 of Chapter 1 makes it the constant \(E\left[{\mathbb Md}(X)^2\right]\) and the limit is an ordinary normal. Otherwise the limit is a mixture of normals, one for each statistical model, mixed by the probabilities that \({\text{Pr}}\) assigns to the invariant events. A long time series discloses the variance of whichever statistical model generated it, not the mixture.

This finding has a straightforward extension to a multivariate counterpart of \(X\) through the study of all linear combinations.

Observe that the variance in the central limit approximation is the variance of the martingale difference:

\[E\left[ \mathbb{Md}(X)^2 \mid {\mathfrak{I}} \right] \ne E\left[ X^2 \mid \mathfrak{I} \right].\]

Consider the moving-average example in Section Permanent shocks. Then

(3.6)#\[\begin{align} E\left[ \mathbb{Md}(X)^2 \mid {\mathfrak{I}} \right] & = \left| \sum_{j=0}^\infty \alpha_j \right|^2 \cr E\left[ X^2 \mid {\mathfrak{I}} \right] & = \sum_{j=0}^\infty \left| \alpha_j \right|^2, \end{align}\]

which are typically distinct. The first of these computations is the variance pertinent for the central limit approximation.

3.6. Cointegration#

Linear combinations of stationary increment processes \(Y_t^1\) and \(Y_t^2\) have stationary increments. For real-valued scalars \(r_1\) and \(r_2\), form

\[Y_{t} = r_1 Y_{t}^1 + r_2 Y_{t}^2\]

where

\[\begin{split}\begin{align*} Y_{t+1}^1 - Y_t^1 & = X_{t+1}^1 \\ Y_{t+1}^2 - Y_t^2 & = X_{t+1}^2. \end{align*}\end{split}\]

The increment in \(\{Y_t : t=0, 1, \ldots \}\) is

\[X_{t+1} = r_1 X_{t+1}^1 + r_2 X_{t+1}^2\]

and

\[Y_0 = r_1 Y_0^1 + r_2 Y_0^2.\]

The Proposition 3.1 martingale component of \(\{ Y_t : t \ge 0 \}\) is the corresponding linear combination of the martingale components of \(\{ Y_t^1 : t =0,1,...\}\) and \(\{ Y_t^2 : t =0,1,...\}\). The Proposition 3.1 trend component of \(\{ Y_t : t =0,1, \ldots \}\) is the corresponding linear combination of the trend components of \(\{ Y_t^1 : t =0,1, \ldots \}\) and \(\{ Y_t^2 : t =0,1, \ldots \}\).

Proposition 3.1 sheds light on the cointegration concept of [Engle and Granger, 1987] that is associated with linear combinations of stationary increment processes whose trend and martingale components are both zero. [Engle and Granger, 1987] call two processes cointegrated if there exists a linear combination of them that is stationary.[4] That situation prevails when there exist real-valued scalars \(r_1\) and \(r_2\) such that

\[\begin{split}\begin{eqnarray*} r_1 \eta_1 + r_2 \eta_2 & = & 0 \\ r_1 \mathbb{Md}(X^1) + r_2 \mathbb{Md}(X^2) & = & 0, \end{eqnarray*}\end{split}\]

where the \(\eta\)’s correspond to the trend components in Proposition 3.1. These two zero restrictions imply that the time trend and the martingale component for the linear combination \(Y_t\) are both zero.[5] When \(r_1 = 1\) and \(r_2 = - 1\), the stationary increment processes \(Y_{t}^1\) and \(Y_{t}^2\) share a common growth component.

This notion of cointegration provides one way to formalize balanced growth paths in stochastic environments through determining a linear combination of growing time series for which stochastic growth is absent. Section Expected discounted cash flows of Chapter 8 shows how this same restriction shapes the long-term valuation of stochastic cash flows.

3.7. Summary#

Proposition 3.1 decomposes a process with stationary increments into a linear trend, a martingale, a stationary component, and a component fixed at date zero. The martingale increment is the permanent shock: by (3.4) it is the revision in the forecast of \(Y_{t+j}\) as \(j\) grows without bound. Which variance governs long-horizon behavior follows from this. Proposition 3.2 scales partial sums by \(1/\sqrt{t}\) and finds the variance of the martingale increment, not the variance of the increment process itself.

The construction requires summable moving-average coefficients, so that \(\alpha(1)\) is finite; without that, the distinction between permanent and transitory dissolves. Chapter 4 obtains the same decomposition constructively when the increments are Markov, and Section Expected discounted cash flows of Chapter 8 shows that the cointegration restriction developed at the end of this chapter governs the long-term valuation of stochastic cash flows.

3.8. Exercises#

Exercise 3.1 (The Gordin decomposition of a first-order moving average)

Let \(\{W_t : -\infty < t < \infty\}\) be an i.i.d. sequence of standard normal random variables and let the increment process be the first-order moving average

\[X_{t} = W_{t} + \theta W_{t-1} ,\]

with \(Y_{t+1} - Y_t = X_{t+1}\) as in (3.1). Carry out the construction of Section Permanent shocks by hand.

(a) Compute the trend coefficient \(\eta\).

(b) Compute \(H_t\) and \(H_t^+\), and then the martingale increment \(G_t\) from (3.2).

(c) Verify directly that the four terms of Proposition 3.1 sum to \(Y_t\), by comparing with \(Y_0 + \sum_{j=1}^t X_j\). Identify which of the four terms is the trend, which the martingale, which the stationary component, and which is invariant.

(d) Confirm that your \(G_t\) agrees with the moving-average formula \(G_t = \alpha(1)\cdot W_t\) of Section Permanent shocks.

Exercise 3.2 (Which variance appears in the central limit theorem)

Continue with the first-order moving average of Exercise 3.1.

(a) Compute \({\mathbb E}\left[{\mathbb Md}(X)^2 \mid {\mathfrak I}\right]\) and \({\mathbb E}\left[X^2 \mid {\mathfrak I}\right]\) and confirm that they differ, as the chapter asserts after Proposition 3.2.

(b) For which values of \(\theta\) is the long-run variance larger than the one-period variance, and for which is it smaller? Where are they equal?

(c) Take \(\theta = -1\). Show that \(\alpha(1) = 0\), that the martingale component vanishes, and that \(\{Y_t\}\) is in fact a stationary process. Exhibit \(Y_t\) explicitly.

(d) A researcher estimates the one-period variance of \(X\) accurately and uses it to scale a confidence interval for \(Y_t/\sqrt{t}\). Using (b), describe the two ways this can go wrong and say which is the more dangerous in practice.

Exercise 3.3 (Constructing a cointegrated pair)

Let \(\{W_t\}\) be an i.i.d. sequence of standard normal random vectors and consider two processes with stationary increments,

\[Y_{t+1}^i - Y_t^i = \eta_i + X_{t+1}^i, \qquad X_t^i = \sum_{j=0}^\infty \alpha_j^i \cdot W_{t-j}, \qquad i = 1,2,\]

with both coefficient sequences summable.

(a) State, in terms of \(\eta_i\) and \(\alpha^i(1)\), the conditions under which real scalars \(r_1, r_2\) make \(r_1 Y^1 + r_2 Y^2\) stationary.

(b) Take \(\alpha_j^1 = \rho^j {\sf c}\) and \(\alpha_j^2 = \lambda^j {\sf d}\) for vectors \({\sf c}, {\sf d}\) and scalars \(|\rho|, |\lambda| < 1\). Find the restriction on \(({\sf c}, {\sf d}, \rho, \lambda)\) under which \((1,-1)\) is a cointegrating vector, given \(\eta_1 = \eta_2\).

(c) Under your restriction in (b), are the two increment processes identical? Are their impulse responses identical? Explain what cointegration does and does not require of the short-run dynamics.

(d) Section Expected discounted cash flows of Chapter 8 studies cash flows of the form \(G_t f(X_t)\) and observes that, stated in logarithms, they are all cointegrated with vector \((1,-1)\). Using (c), explain why a whole family of cash flows can share one long-run growth component while differing arbitrarily at short horizons.

Exercise 3.4 (When the martingale decomposition is unavailable)

The chapter notes that \(\alpha(1)\) may fail to be finite even when \(\sum_j |\alpha_j|^2 < \infty\), so that \(X_t\) is well defined by (1.9) while the construction of Section Permanent shocks breaks down. This exercise makes that concrete.

Take scalar coefficients \(\alpha_j = (j+1)^{-{\sf p}}\) for a constant \({\sf p}\).

(a) For which \({\sf p}\) is \(\sum_j |\alpha_j|^2\) finite? For which is \(\sum_j \alpha_j\) finite? Exhibit a range of \({\sf p}\) for which the first holds and the second fails.

(b) For such a \({\sf p}\), is \(X_t\) a well-defined stationary process? Is \(X\) in the set \({\mathcal H}\) defined in Section A martingale decomposition? Which requirement fails, and at which step of the construction does the failure first bite?

(c) Since Proposition 3.1 is unavailable, Proposition 3.2 does not apply. What does that say about the scaling \(1/\sqrt{t}\) used to form the partial sums?

(d) The chapter observes that this possibility “opened the door to the literature on long-memory processes.” Explain in a sentence or two why a modeler might nonetheless prefer to impose summability, and what is being assumed away by doing so.

3.9. Answers#