Analysis of high-dimensional numerical data : from principal component analysis to non-linear dimensionality reduction and blind source separation

(2003)

Files

Lee.pdf
  • Restricted Access
  • Adobe PDF
  • 27.63 MB

Details

Authors
Supervisors
Verleysen, Michel
Abstract
Numerical data often raise some difficulties during their analysis, inter alia because of their high dimensionality. When a large number of observed variables are available through sensors or other measurement processes, the analyst has to discover whether those variables depend on each other. If this is the case, then is it possible to summarise the data into a smaller set of so-called latent variables? Traditionally, such a transformation is computed by Principal Component Analysis (PCA) or metric multidimensional scaling (MDS). This thesis provides alternative answers to that question, in the form of recent methods. A first important part of the thesis explores methods that can achieve a non-linear dimensionality reduction of data. Because these methods can transform data in a non-linear way, they outperform PCA and metric MDS, which are linear only. Mainly geometric methods, based on either distance or topology preservation, are studied. Beside Sammon's non-linear mapping (NLM) and Kohonen's Self-Organising Maps (SOMs), the thesis deals with recent methods like Curvilinear Component Analysis, Isomap, LLE or Kernel PCA. Original contributions also include Curvilinear Distance Analysis, Isotop and modified update rules for SOMs. A second important part of the thesis goes toward true latent variable separation. Contrarily to non-linear dimensionality reduction which only seeks a compact representation of data, latent variable separation aims at isolating and retrieving the true hidden parameters that generate data. Blind Source Separation (BSS) and Independent Component Analysis (ICA) rely on reasonable assumptions like statistical independence in order to solve that problem. Several methods that implement BSS and ICA techniques are described and applied on electroencephalographic signals. Each part of the thesis also contains a discussion about the estimation of the data dimensionality. All methods described in the thesis are compared to each other according to various criterions and illustrated with toy examples as well as real applications, like image and sound processing. Finally, conclusions classify all used methods, propose a generic analysis schema and outline the current tendencies in the domain.
Affiliations

Citations

Lee, J. (2003). Analysis of high-dimensional numerical data : from principal component analysis to non-linear dimensionality reduction and blind source separation. https://hdl.handle.net/2078.5/98464