Filtering Real Time Series

STA 542 · Homework 4 exploration workbook

This workbook supplies the observations for Homework 4, Problem 1. You are encouraged to explore all eleven series, but only need to submit answers to parts (a)–(e) for four. Try to choose four series for which different filtering techniques are appropriate.

  1. Choose a filter. Based on a plot of the time series, propose a filter that you think will make it weakly stationary, and explain which features of the plot motivate your choice.

  2. Does the filter need fitting? If so, explain what you fit and how. Apply the filter and plot the result.

  3. Is the filter invertible? What is lost? Explain your answer. If information is lost, give a scientific question that can no longer be answered from the filtered data alone.

  4. Examine the result. Plot the sample ACVF. Estimate the long-run variance using the Bartlett estimator and run KPSS at a few bandwidths. Explain your bandwidth choices and report the results.

  5. Assess stationarity. How convincing is a weakly stationary model for the filtered series? Support your answer using the plots and results.

KPSS assumes a short-memory stationary remainder with positive long-run variance; nonrejection does not establish all the requirements of weak stationarity.

Keep the data folder beside this Quarto file. Each code block loads and plots one series; the time coordinate and measurement units are described below.

1. Ocean Waves During a Falling Tide

Water depth above a pressure sensor near Los Angeles: 7,200 observations at four per second over 30 minutes during a falling tide on August 19, 2016. Values are in meters. They are pressure-derived depths without a correction for wave-pressure attenuation, so they understate surface-wave fluctuations.

d <- read.csv("data/ocean-waves.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Seconds from first observation",
     ylab = "Water depth (m)")
Figure 1: Real data: Ocean Waves During a Falling Tide.

Source: oceanwaves::wavedata, column swDepth.m. Dataset explanation and measurement processing.

2. Ancient Sediment Layers

Thicknesses of 634 consecutive annual sediment layers from Massachusetts, beginning nearly 12,000 years ago. Time counts years in the sequence; physical thickness units are not specified in the source.

d <- read.csv("data/varve.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Year in sequence",
     ylab = "Layer thickness (recorded units)")
Figure 2: Real data: Ancient Sediment Layers.

Source: astsa::varve. Dataset documentation.

3. The Cost of Computer Storage

Median listed retail cost per gigabyte of sampled hard drives, in dollars, for 29 years from 1980 through 2008. The historical prices are not adjusted for inflation.

d <- read.csv("data/storage-cost.csv")
x <- ts(d$value, start = 1980, frequency = 1)
plot(x, type = "l", col = "#00539B",
     xlab = "Year",
     ylab = "Dollars per GB")
Figure 3: Real data: The Cost of Computer Storage.

4. A Variable Star

A star’s magnitude at midnight on 600 consecutive days, on the source’s recorded scale. Shumway and Stoffer, Example 4.5, identify approximately 24-day and 29-day cycles in this record; these periods may be used here.

d <- read.csv("data/star.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Day in sequence",
     ylab = "Stellar magnitude (recorded scale)")
Figure 4: Real data: A Variable Star.

Source: astsa::star. Dataset documentation.

5. One Brain Signal During Repeated Stimulation

One cortical BOLD signal (cort1) from an fMRI experiment: 128 measurements two seconds apart. Stimulation alternated between 32 seconds on and 32 seconds off, giving a 32-observation cycle. Signal units are unspecified.

d <- read.csv("data/fmri.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Seconds from first observation",
     ylab = "BOLD signal (recorded units)")
Figure 5: Real data: One Brain Signal During Repeated Stimulation.

Source: astsa::fmri1, column cort1. Dataset documentation.

6. Annual Flow of the Nile

Annual water volume at Aswan, 1871–1970: 100 observations in units of 10⁸ cubic meters. The source notes an apparent level change near 1898.

d <- read.csv("data/nile.csv")
x <- ts(d$value, start = 1871, frequency = 1)
plot(x, type = "l", col = "#00539B",
     xlab = "Year",
     ylab = expression("Annual volume (" * 10^8 * " m"^3 * ")"))
Figure 6: Real data: Annual Flow of the Nile.

Source: datasets::Nile. Dataset documentation.

7. A Beaver’s Body Temperature

One female beaver’s body temperature in degrees Celsius, measured every ten minutes in Wisconsin on November 3–4, 1990. There are 100 observations; the record is shorter than one day.

d <- read.csv("data/beaver.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Minutes from first observation",
     ylab = "Body temperature (degrees C)")
Figure 7: Real data: A Beaver’s Body Temperature.

Source: datasets::beaver2, column temp. Dataset documentation.

8. An Earthquake Recording

A seismic trace with 2,048 observations. The source identifies the first 1,024 with the primary-wave arrival and the remainder with the shear-wave arrival. Sampling interval and physical amplitude units are unspecified.

d <- read.csv("data/earthquake.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Observation",
     ylab = "Seismic signal (recorded units)")
Figure 8: Real data: An Earthquake Recording.

Source: astsa::EQ5. Dataset documentation.

9. Monthly Temperatures in Nottingham

Monthly average air temperature at Nottingham Castle, England: 240 values from January 1920 through December 1939, in degrees Fahrenheit.

d <- read.csv("data/nottingham.csv")
x <- ts(d$value, start = c(1920, 1), frequency = 12)
plot(x, type = "l", col = "#00539B",
     xlab = "Year",
     ylab = "Temperature (degrees F)")
Figure 9: Real data: Monthly Temperatures in Nottingham.

Source: datasets::nottem. Dataset documentation.

10. The DAX Stock Index

Daily closing values of the German DAX stock index, in index points: 1,860 trading observations during 1991–1998. Weekends and holidays are omitted; one step is one trading observation.

d <- read.csv("data/dax.csv")
x <- d$value
plot(d$time, x, type = "l", col = "#00539B",
     xlab = "Trading observation",
     ylab = "DAX index points")
Figure 10: Real data: The DAX Stock Index.

Source: datasets::EuStockMarkets, column DAX. Dataset documentation.

11. Atmospheric Carbon Dioxide at Mauna Loa

Monthly atmospheric CO₂ concentration at Mauna Loa, in ppm: 468 values from January 1959 through December 1997. The R dataset fills February–April 1964 by linear interpolation; these supplied values are retained.

d <- read.csv("data/co2.csv")
x <- ts(d$value, start = c(1959, 1), frequency = 12)
plot(x, type = "l", col = "#00539B",
     xlab = "Year",
     ylab = "CO2 concentration (ppm)")
Figure 11: Real data: Atmospheric Carbon Dioxide at Mauna Loa.

Source: datasets::co2, from Scripps Institution of Oceanography records. Dataset documentation.