3. Processes with Stationary Increments#
Authors: Lars Peter Hansen (University of Chicago) and Thomas J. Sargent (NYU)
\(\newcommand{\eqdef}{\stackrel{\text{def}}{=}}\)
3.1. Introduction#
Earlier chapters explain why we like statistical models that generate stationary stochastic processes: stationarity brings a Law of Large Numbers that helps us identify statistical models and make inferences about model parameters. However, logarithms of many economic time series appear not to be stationary. Instead they grow systematically. This situation motivates us to study models that generate stochastic processes with stationary increments. Multivariate versions of such models possess stochastic process versions of balanced growth paths. Applied econometricians sometimes study permanent shocks that contribute to stochastic growth. We shall describe how to pose a central limit theory and to characterize its implications in terms of processes with stationary increments.
The mathematical formulation in this chapter opens the door to studying these topics using a unified set of tools. In this chapter we return to the mathematical formulation used in Chapter 1, while in the next chapter we will assume a Markov structure.
3.2. Basic setup#
We adopt assumptions from Section Inventing an Infinite Past of Chapter 1 that allow an infinite past. Again let \( {\mathfrak A}\) be a subsigma algebra of \({\mathfrak F}\) and
The event collection \({\mathfrak A}\) includes invariant events as well as past information. Thus we assume that
so information accumulates over time. Formally, \(\{ {\mathfrak A}_t \}\) is called a filtration, which captures the dynamic information flow.
Let \(X\) be a scalar measurement function that is \({\mathfrak A} = {\mathfrak A}_0\) measurable. Assume that \(Y_0\) is \({\mathfrak A}_0\) measurable, and consider a scalar process \(\{Y_t : t=0,1,... \}\) with stationary increments \(\{X_t\}\):
for \(t=0,1, \ldots\). Notice that both \(X_{t+1}\) and \(Y_{t+1}\) are in the implied date \(t+1\) collection of information. Let
and
We interpret the above equations as providing two contributions to the \(\{Y_{t}: t \ge 0\}\) process. Component \(U_{t+1}\) is unpredictable and represents new information about \(Y_{t+1}\) that arrives at date \(t+1\). Component \(\eta\) is the trend rate of growth or decay in \(\{Y_{t} : t \ge 0\}\) conditioned on the invariant events. In the following sections, we present a full decomposition of a stationary increment process that will be useful both in connecting to sources of permanent versus transitory shocks and to central limit theorems.
3.3. A martingale decomposition#
A special class of stationary increment processes consisting of what we call additive martingales interests us.
Definition 3.1
The process \(\{Y_t^m : t=0,1,... \}\) is said to be an additive martingale relative to \(\{ {\mathfrak A}_{t} : t=0,1,... \}\) if for \(t=0,1,... \)
\(Y_t^m\) is \({\mathfrak A}_{t}\) measurable, and
\(E\left(Y_{t+1}^m \vert {\mathfrak A}_t \right) = Y_t^m\) .
Notice that by the Law of Iterated Expectations, for a martingale \(\{Y_{t}^m : t \ge 0\}\), best forecasts satisfy:
for \(j \ge 1\). Under suitable additional restrictions on the increment process \(\{X_t : t \ge 0 \}\), we can deploy a construction of [Gordin, 1969] to build a martingale component of the \(\{Y_t^m : t=0,1, ... \}\) process.[1] Let \({\mathcal H}\) denote the set of all scalar random variables \(X\) such that \(E(X^2) < \infty\) and such that[2]
is well defined as a mean-square convergent series. Convergence of the infinite sum on the right side limits temporal dependence of the process \(\{ X_t \}\). For example, it can exclude so-called long memory processes.[3]
Construct the one-period ahead forecast of \(H_{t+1}\) conditioned on date \(t\) information:
Notice that
where
Since \(G_t\) is a forecast error,
Assembling these parts, we have
Let
Since \(Y_t^m\) is \({\mathfrak A}_{t}\) measurable, the equality
implies that the process \(\{Y_t^m : t \ge 0 \}\) is an additive martingale.
For a given stationary increment process, \(\{Y_t : t \ge 0\}\), express the martingale increment as
So the increment to the martingale component of \(\{Y_t : t \ge 0 \}\) provides new information about the limiting optimal forecast of \(Y_{t+j}\) as \(j \rightarrow + \infty\).
By accumulating equation (3.3) forward, we arrive at:
Proposition 3.1
If \(X\) is in \({\mathcal H}\), the stationary increments process \(\{Y_t : t=0,1,...\}\) satisfies the additive decomposition
where we impose the normalization, \(Y_0^m = 0.\) The two stationary increment processes, \(\{ t\eta : t \ge 0\}\) and \(\{Y_{t}^m : t\ge 0 \},\) are deterministic and martingale components of the original process. The contribution, \(\{H_{t}^+\},\) is a stationary component. The other components are constant over time.
Proposition 3.1 decomposes a stationary-increment process into a linear time trend, a martingale, and a transitory component. A permanent shock is the increment to the martingale. The martingale and transitory contributions are typically correlated. Some decomposition methods go one step further by adjusting the decomposition to remove the correlation between these two components as we will illustrate in an example that follows.
With this mathematical structure in place, we construct an operator \(\mathbb{Md}\) that maps an admissible increment process in \(\mathcal{H}\) into the innovation in a martingale component. Let \(\mathcal{G}\) be the set of all random variables \(G\) with finite second moments that satisfy the conditions that i) \(G\) is \(\mathfrak{A}\) measurable and that ii) \(E(G_1 \vert \mathfrak{A}) = 0\) where \(G_t = G \circ \mathbb{S}^t\). Define \(\mathbb{Md}: \mathcal{H} \rightarrow \mathcal{G}\)
for \(G = G_0\) given by (3.4) for \(t = 0\). Both \(\mathcal{G}\) and \(\mathcal{H}\) are linear spaces of random variables and \(\mathbb{Md}\) is a linear transformation. The operator \(\mathbb{Md}\) plays a prominent role in some of the analysis that follows. Chapter 4 supplies a constructive algorithm for computing \(\mathbb{Md}\) when the increments are generated by a stationary Markov process.
3.4. Permanent shocks#
In this construction, we impose a moving-average structure on the underlying time series.
Specifically, consider again the Example 1.8 moving-average process:
Use this \(\{X_t\}\) process as the increment for \(\{ Y_t : t \ge 0 \}\) in formula (3.1). New information about the unpredictable component of \(X_{t+j}\) for \(j \ge 0\) that arrives at date \(t\) is
Summing these terms over \(j\) gives
where
provided that the coefficient sequence \(\{ \alpha_j : j\ge 0\}\) is summable, a condition that restricts temporal dependence of the increment process \(\{X_t\}\). Indeed, it is possible for \(\alpha(1) = \infty\) or for it not to be well defined while
ensuring that \(X_t\) is well defined. This possibility opened the door to the literature on long-memory processes that allow for \(\alpha(1)\) to be infinite as discussed in [Granger and Joyeux, 1980] and elsewhere.
In what follows, we presume that \(\alpha(1)\) is finite. This sum of the coefficients \(\{\alpha_j: j\ge 0 \}\) in moving-average representation (3.5) for the first difference \(Y_{t+1} - Y_t = X_{t+1}\) of \(\{ Y_t : t=0,1,.... \}\) reveals the permanent effect of \(W_{t+1}\) on future values of the level of \(Y\), i.e., the effect on \(\lim_{j\rightarrow + \infty} Y_{t+j}\). Models of [Blanchard and Quah, 1989] and [Shapiro and Watson, 1988] build on this property.
The variance of the random variable \(\alpha(1) \cdot W_{t+1}\) conditioned on the invariant events in \({\mathfrak I}\) is \(|\alpha(1)|^2\). The overall variance of \(X_{t}\) is
where \(|\cdot |\) is the Euclidean norm. The two magnitudes typically differ:
In what follows, we assume that \(\alpha(1) \ne 0\). To form a permanent-transitory shock decomposition, construct the scalar permanent shock as:
where we introduce an additional scaling so the permanent shock has variance one. Form
which by construction will be uncorrelated with \(W_{t+1}^p\). The terms in square brackets are the population least squares regression coefficients where the regressors are the variables that follow. The covariance matrix for \(W_{t+1}^{tr}\) is given by the idempotent matrix:
and has reduced rank. As a consequence, the entries of \(W_{t+1}^{tr}\) can be expressed as linear combinations of a vector of transitory shocks with unit variances and one fewer dimension.
3.5. Central limit approximation#
In this section, we produce a central limit approximation for temporally dependent processes originally due to [Gordin, 1969]. We view Gordin’s result as an application of Proposition 3.1.
To form a central limit approximation, construct the following scaled partial sum that nets out trend growth
where
From [Billingsley, 1961]’s central limit theorem for martingales
where \(\Rightarrow\) denotes weak convergence, meaning convergence in distribution. Clearly, \(\{(1/ {\sqrt t}) H_{t}^+\}\) and \(\{(1/{\sqrt{t}}) (H_0^+ + Y_0) \}\) both converge in mean square to zero.
Proposition 3.2
For all stationary increment processes \(\{Y_t : t=0,1,2, ...\}\) represented by \(X\) in \(\mathcal{H}\)
Furthermore,
The variance in Proposition 3.2 is conditioned on the invariant events, so it is a random variable rather than a number. When \({\mathbb S}\) is ergodic, Proposition 1.1 of Chapter 1 makes it the constant \(E\left[{\mathbb Md}(X)^2\right]\) and the limit is an ordinary normal. Otherwise the limit is a mixture of normals, one for each statistical model, mixed by the probabilities that \({\text{Pr}}\) assigns to the invariant events. A long time series discloses the variance of whichever statistical model generated it, not the mixture.
This finding has a straightforward extension to a multivariate counterpart of \(X\) through the study of all linear combinations.
Observe that the variance in the central limit approximation is the variance of the martingale difference:
Consider the moving-average example in Section Permanent shocks. Then
which are typically distinct. The first of these computations is the variance pertinent for the central limit approximation.
3.6. Cointegration#
Linear combinations of stationary increment processes \(Y_t^1\) and \(Y_t^2\) have stationary increments. For real-valued scalars \(r_1\) and \(r_2\), form
where
The increment in \(\{Y_t : t=0, 1, \ldots \}\) is
and
The Proposition 3.1 martingale component of \(\{ Y_t : t \ge 0 \}\) is the corresponding linear combination of the martingale components of \(\{ Y_t^1 : t =0,1,...\}\) and \(\{ Y_t^2 : t =0,1,...\}\). The Proposition 3.1 trend component of \(\{ Y_t : t =0,1, \ldots \}\) is the corresponding linear combination of the trend components of \(\{ Y_t^1 : t =0,1, \ldots \}\) and \(\{ Y_t^2 : t =0,1, \ldots \}\).
Proposition 3.1 sheds light on the cointegration concept of [Engle and Granger, 1987] that is associated with linear combinations of stationary increment processes whose trend and martingale components are both zero. [Engle and Granger, 1987] call two processes cointegrated if there exists a linear combination of them that is stationary.[4] That situation prevails when there exist real-valued scalars \(r_1\) and \(r_2\) such that
where the \(\eta\)’s correspond to the trend components in Proposition 3.1. These two zero restrictions imply that the time trend and the martingale component for the linear combination \(Y_t\) are both zero.[5] When \(r_1 = 1\) and \(r_2 = - 1\), the stationary increment processes \(Y_{t}^1\) and \(Y_{t}^2\) share a common growth component.
This notion of cointegration provides one way to formalize balanced growth paths in stochastic environments through determining a linear combination of growing time series for which stochastic growth is absent. Section Expected discounted cash flows of Chapter 8 shows how this same restriction shapes the long-term valuation of stochastic cash flows.
3.7. Summary#
Proposition 3.1 decomposes a process with stationary increments into a linear trend, a martingale, a stationary component, and a component fixed at date zero. The martingale increment is the permanent shock: by (3.4) it is the revision in the forecast of \(Y_{t+j}\) as \(j\) grows without bound. Which variance governs long-horizon behavior follows from this. Proposition 3.2 scales partial sums by \(1/\sqrt{t}\) and finds the variance of the martingale increment, not the variance of the increment process itself.
The construction requires summable moving-average coefficients, so that \(\alpha(1)\) is finite; without that, the distinction between permanent and transitory dissolves. Chapter 4 obtains the same decomposition constructively when the increments are Markov, and Section Expected discounted cash flows of Chapter 8 shows that the cointegration restriction developed at the end of this chapter governs the long-term valuation of stochastic cash flows.
3.8. Exercises#
Exercise 3.1 (The Gordin decomposition of a first-order moving average)
Let \(\{W_t : -\infty < t < \infty\}\) be an i.i.d. sequence of standard normal random variables and let the increment process be the first-order moving average
with \(Y_{t+1} - Y_t = X_{t+1}\) as in (3.1). Carry out the construction of Section Permanent shocks by hand.
(a) Compute the trend coefficient \(\eta\).
(b) Compute \(H_t\) and \(H_t^+\), and then the martingale increment \(G_t\) from (3.2).
(c) Verify directly that the four terms of Proposition 3.1 sum to \(Y_t\), by comparing with \(Y_0 + \sum_{j=1}^t X_j\). Identify which of the four terms is the trend, which the martingale, which the stationary component, and which is invariant.
(d) Confirm that your \(G_t\) agrees with the moving-average formula \(G_t = \alpha(1)\cdot W_t\) of Section Permanent shocks.
Exercise 3.2 (Which variance appears in the central limit theorem)
Continue with the first-order moving average of Exercise 3.1.
(a) Compute \({\mathbb E}\left[{\mathbb Md}(X)^2 \mid {\mathfrak I}\right]\) and \({\mathbb E}\left[X^2 \mid {\mathfrak I}\right]\) and confirm that they differ, as the chapter asserts after Proposition 3.2.
(b) For which values of \(\theta\) is the long-run variance larger than the one-period variance, and for which is it smaller? Where are they equal?
(c) Take \(\theta = -1\). Show that \(\alpha(1) = 0\), that the martingale component vanishes, and that \(\{Y_t\}\) is in fact a stationary process. Exhibit \(Y_t\) explicitly.
(d) A researcher estimates the one-period variance of \(X\) accurately and uses it to scale a confidence interval for \(Y_t/\sqrt{t}\). Using (b), describe the two ways this can go wrong and say which is the more dangerous in practice.
Exercise 3.3 (Constructing a cointegrated pair)
Let \(\{W_t\}\) be an i.i.d. sequence of standard normal random vectors and consider two processes with stationary increments,
with both coefficient sequences summable.
(a) State, in terms of \(\eta_i\) and \(\alpha^i(1)\), the conditions under which real scalars \(r_1, r_2\) make \(r_1 Y^1 + r_2 Y^2\) stationary.
(b) Take \(\alpha_j^1 = \rho^j {\sf c}\) and \(\alpha_j^2 = \lambda^j {\sf d}\) for vectors \({\sf c}, {\sf d}\) and scalars \(|\rho|, |\lambda| < 1\). Find the restriction on \(({\sf c}, {\sf d}, \rho, \lambda)\) under which \((1,-1)\) is a cointegrating vector, given \(\eta_1 = \eta_2\).
(c) Under your restriction in (b), are the two increment processes identical? Are their impulse responses identical? Explain what cointegration does and does not require of the short-run dynamics.
(d) Section Expected discounted cash flows of Chapter 8 studies cash flows of the form \(G_t f(X_t)\) and observes that, stated in logarithms, they are all cointegrated with vector \((1,-1)\). Using (c), explain why a whole family of cash flows can share one long-run growth component while differing arbitrarily at short horizons.
Exercise 3.4 (When the martingale decomposition is unavailable)
The chapter notes that \(\alpha(1)\) may fail to be finite even when \(\sum_j |\alpha_j|^2 < \infty\), so that \(X_t\) is well defined by (1.9) while the construction of Section Permanent shocks breaks down. This exercise makes that concrete.
Take scalar coefficients \(\alpha_j = (j+1)^{-{\sf p}}\) for a constant \({\sf p}\).
(a) For which \({\sf p}\) is \(\sum_j |\alpha_j|^2\) finite? For which is \(\sum_j \alpha_j\) finite? Exhibit a range of \({\sf p}\) for which the first holds and the second fails.
(b) For such a \({\sf p}\), is \(X_t\) a well-defined stationary process? Is \(X\) in the set \({\mathcal H}\) defined in Section A martingale decomposition? Which requirement fails, and at which step of the construction does the failure first bite?
(c) Since Proposition 3.1 is unavailable, Proposition 3.2 does not apply. What does that say about the scaling \(1/\sqrt{t}\) used to form the partial sums?
(d) The chapter observes that this possibility “opened the door to the literature on long-memory processes.” Explain in a sentence or two why a modeler might nonetheless prefer to impose summability, and what is being assumed away by doing so.
3.9. Answers#
Solution to Exercise 3.1 (The Gordin decomposition of a first-order moving average)
(a) \(\eta = {\mathbb E}\left(X_{t+1}\mid{\mathfrak I}\right) = 0\), since each \(W\) has mean zero. There is no trend.
(b) Conditioning on \({\mathfrak A}_t\), which contains \(W_t\) and its past:
Summing, \(H_t = (1+\theta)W_t + \theta W_{t-1}\). Then
and from (3.2),
As required, \({\mathbb E}\left(G_{t+1}\mid{\mathfrak A}_t\right) = 0\).
(c) Proposition 3.1 with \(\eta = 0\) and \(Y_0^m = 0\) reads
Directly,
where the last step adds and subtracts \(\theta W_t\). The two expressions agree term by term. The martingale accumulates without bound, the \(-\theta W_t\) term is stationary, and \(Y_0 + \theta W_0\) is fixed at date zero.
(d) Here \(\alpha_0 = 1\), \(\alpha_1 = \theta\), and \(\alpha_j = 0\) for \(j \ge 2\), so \(\alpha(1) = 1 + \theta\) and the formula \(G_t = \alpha(1)\cdot W_t\) gives \((1+\theta)W_t\), matching (b).
Solution to Exercise 3.2 (Which variance appears in the central limit theorem)
(a) From Exercise 3.1, \({\mathbb Md}(X) = (1+\theta)W\), so
These differ, exactly as the chapter’s display asserts.
(b) \((1+\theta)^2 - \left(1+\theta^2\right) = 2\theta\). So the long-run variance exceeds the one-period variance when \(\theta > 0\), falls short when \(\theta < 0\), and the two coincide only at \(\theta = 0\), where \(X\) is i.i.d. and there is nothing to distinguish. Positive \(\theta\) means a shock is reinforced next period, so its effects on the level cumulate; negative \(\theta\) means it is partly undone.
(c) At \(\theta = -1\), \(\alpha(1) = 0\), so \(G_t = 0\) and the martingale component is identically zero. Then
which telescopes. The process is stationary once we set \(Y_0 = W_0\), and in any case has bounded variance: a process built as a cumulated sum of increments need not itself be growing. Notice that \(Y_t/\sqrt{t}\) converges to zero rather than to a nondegenerate normal, which is consistent with Proposition 3.2 since the limiting variance is \(|\alpha(1)|^2 = 0\).
(d) The scaling is wrong in one of two directions. If \(\theta > 0\) the researcher understates the long-run variance and the confidence interval is too narrow, so the true coverage falls short of the nominal level and the null is rejected too often. If \(\theta < 0\) the interval is too wide and the procedure is conservative.
The first is the more dangerous. A conservative interval costs power; an anti-conservative one produces confident false findings, and the error grows with \(\theta\). Since positive serial correlation in first differences is the common case for macroeconomic and financial series, this is the direction one usually faces. Proposition 3.2 says one must estimate \(|\alpha(1)|^2\), the variance of the martingale increment. No amount of accuracy about \({\mathbb E}\left[X^2\right]\) substitutes for it.
Solution to Exercise 3.3 (Constructing a cointegrated pair)
(a) Linear combinations of stationary increment processes have stationary increments, so \(Y_t = r_1Y_t^1 + r_2Y_t^2\) has increment \(r_1X_{t+1}^1 + r_2X_{t+1}^2\) and trend \(r_1\eta_1 + r_2\eta_2\). Applying \({\mathbb Md}\), which is linear, its martingale increment is \(r_1\alpha^1(1)\cdot W_t + r_2\alpha^2(1)\cdot W_t\). The combination is stationary exactly when both vanish:
(b) Summing the geometric sequences, \(\alpha^1(1) = \frac{{\sf c}}{1-\rho}\) and \(\alpha^2(1) = \frac{{\sf d}}{1-\lambda}\). With \(\eta_1 = \eta_2\) the first condition holds at \((1,-1)\), and the second requires
The two long-run responses must coincide as vectors.
(c) No and no. The increment processes are \(X_t^1 = {\sf c}\cdot\sum_j \rho^j W_{t-j}\) and \(X_t^2 = {\sf d}\cdot\sum_j\lambda^jW_{t-j}\), which are different processes whenever \(\rho \ne \lambda\). Their impulse responses, the sequences \(\left\{\rho^j{\sf c}\right\}\) and \(\left\{\lambda^j{\sf d}\right\}\), differ at every horizon; only their sums agree. One process may respond sharply and decay fast while the other responds mildly and decays slowly.
So cointegration is a restriction on long-run behavior only. It says the two series are driven by a common permanent component, and it is silent about the transient dynamics, which may be arbitrarily different.
(d) Chapter 8 builds cash flows as \(G_t f(X_t)\) with \(f\) ranging over a set of positive functions of a stationary state. In logarithms this is \(\log G_t + \log f(X_t)\): a common component \(\log G_t\) that carries all the growth, plus a stationary term that depends on which \(f\) was chosen. Differencing two members of the family cancels \(\log G_t\) and leaves a stationary difference, so \((1,-1)\) cointegrates any two of them.
Your answer to (c) explains why the family can be large. Since cointegration constrains only \(\alpha(1)\) and the trend, the choice of \(f\) is free to reshape the short-horizon response as much as it likes. Chapter 8 exploits this: it shows that all members of such a family share one limiting risk-return tradeoff and one limiting holding-period return, precisely because they differ only in a stationary factor.
Solution to Exercise 3.4 (When the martingale decomposition is unavailable)
(a) \(\sum_j (j+1)^{-2{\sf p}}\) converges when \(2{\sf p} > 1\), that is \({\sf p} > {\frac 1 2}\). \(\sum_j (j+1)^{-{\sf p}}\) converges when \({\sf p} > 1\). So for
the coefficients are square summable but not summable. For instance \({\sf p} = {\frac 3 4}\): \(\sum_j (j+1)^{-3/2}\) converges while \(\sum_j (j+1)^{-3/4}\) diverges.
(b) Yes, \(X_t\) is well defined: restriction (1.9) is exactly square summability, and it guarantees that the infinite sum converges in mean square, so \(X_t\) is a stationary process with finite variance \(\sum_j |\alpha_j|^2\).
But \(X\) is not in \({\mathcal H}\). Membership requires that
converge in mean square, and here the \(j^{th}\) term is \(\sum_{i \ge 0}\alpha_{i+j}W_{t-i}\), whose contributions accumulate like the divergent tail sums of \(\alpha\). The failure bites at the very first step of the construction, before any martingale has been extracted: \(H_t\) does not exist, so \(H_t^+\), \(G_t\), and the decomposition (3.3) are all unavailable. Equivalently, \(\alpha(1) = \infty\) and the candidate martingale increment \(\alpha(1)\cdot W_t\) has infinite variance.
(c) With Proposition 3.1 unavailable, there is no martingale to which [Billingsley, 1961]’s theorem can be applied, and Proposition 3.2 has no content. Concretely the scaling \(1/\sqrt{t}\) is wrong: \(Y_t - \eta t\) grows faster than \(\sqrt{t}\), and the correct normalization is \(t^{-{\sf H}}\) for an exponent \({\sf H} > {\frac 1 2}\) determined by \({\sf p}\). A researcher who uses \(\sqrt{t}\) will find the studentized statistic diverging and will reject essentially any null hypothesis in a long enough sample.
(d) Summability buys the whole apparatus of this chapter: a martingale decomposition, a well-defined permanent shock, a \(\sqrt{t}\) central limit theorem, and a cointegration concept resting on \(\alpha(1)\). Without it, none of these objects exists, and much of the machinery that later chapters build on the martingale component has nothing to attach to.
What is assumed away is the possibility that the influence of a shock decays so slowly that its cumulative effect is unbounded, so that the distinction between “permanent” and “transitory” itself dissolves. Whether economic time series are long memory in this sense is an empirical question that finite samples answer poorly, since summable coefficients that decay slowly and non-summable ones look much alike over any sample one has. Imposing summability is therefore a modeling choice made partly for tractability, and the honest position is to say so rather than to treat it as established.