Identification and inference in high-dimensional econometric models

Sauvenier, Mathieu
(2023)

Files

These_V2-1.pdf
  • Open Access
  • Adobe PDF
  • 2.84 MB

Details

Authors
  • Sauvenier, MathieuUCLouvain
    author
Supervisors
Van Bellegem, Sébastien
Abstract
Big data sets are often called high-dimensional data sets in statistical literature. In a linear regression model, for instance, high dimensionality means that the number of covariates is greater than the number of observations. In a multivariate time series framework, high dimensionality refers to situations where many time series are studied jointly. The thesis contributes to the development of statistical analysis of both high-dimensional cross-sectional data and high-dimensional time series. The contributions relate both to the theoretical and to the empirical analysis of high-dimensional data. For high-dimensional cross-sections, we provide new results concerning the modeling, the selection, the estimation and the statistical testing of hypothesis in a high-dimensional linear sparse model. Chapter 2 establishes the connection between the Best Subset Selection and the Generalized Eigenvalue Problem in the context of high dimensional sparse linear regression. The chapter identifies the direction of the parameter of interest with the generalized eigenspace associated with the only nonnull generalized eigenvalue of a pair of measurable matrices. Based on this result, the chapter also introduces a new estimator for which the minimax convergence and a central limit theorem are established. Simulation results highlight the superiority of the suggested estimator over existing methods. Chapter 3 provides a goodness-of-fit test for variable selection in high-dimensional sparse linear models. This test is implemented using a Multiple-splitting Projection Test procedure and is designed to minimize a moment condition that is proven to reflect the sparsity constraint. The chapter demonstrates the test's power through Monte Carlo experiments and applies it to the problem of nowcasting the quarterly Belgian GDP via monthly time series. Chapter 6 briefly discusses high-dimensional data modelling from a more fundamental perspective and provides an early exploration of the connections between existing theoretical setups. For high-dimensional time series data we provide results for modeling and estimating the time varying second order structure of high-dimensional zero-mean processes through the definition, the identification and the estimation of a new class of multivariate multiscale models for locally stationary processes. Chapter 4 defines the class of cross-correlation wavelet functions and studies some of its important properties. The chapter then defines a new class of multivariate locally stationary wavelet processes and identifies the asymptotic representation of the time-varying second order structure of a vector time series in the space of the cross-correlation of wavelet functions. Chapter 5 studies the estimation of the representation of the second order structure of a vector time series in the space of cross-correlation of wavelet functions. An estimator is presented, and its convergence is established. The chapter also presents an economic application of the method that studies the time varying coordination between Saudi Arabia and the rest of OPEC in oil production by providing new measurements not reliant on the assumption of stationarity.
Affiliations

Citations

Sauvenier, M. (2023). Identification and inference in high-dimensional econometric models. https://hdl.handle.net/2078.5/106241