Homogeneity tests for functional data based on depth-depth plots with chemical applications

Published in Chemometrics and Intelligent Laboratory Systems, 2021

Recommended citation: Calle-Saldarriaga, A.; Laniado, H.; Zuluaga, F.; Leiva, V. (2021). "Homogeneity tests for functional data based on depth-depth plots with chemical applications." Chemometrics and Intelligent Laboratory Systems. 219(104420). https://acallesalda.github.io/files/ddplot.pdf

Abstract: One of the standard problems in statistics is determining if two samples come from the same population, that is, testing homogeneity for two samples. In this paper, we propose homogeneity tests in the context of functional data, adopting an idea from multivariate analysis corresponding to the depth-depth plot. This plot is a multivariate generalization of the quantile-quantile plot. We propose some statistics based on the depth-depth plot, and use bootstrapping to approximate their null distributions. We conduct simulations to state the empirical size and power of the proposed tests, obtaining better results than other homogeneity tests considered in the literature. We detect that our test has very high power in relation to other competing tests. We employ many different depths based on what is proposed in the literature to see which is more suitable for this kind of homogeneity testing. Finally, we illustrate the obtained results with chemical heterogeneous data to show potential applications, getting consistent results.

A bit of context: this came out of my MSc. work at EAFIT, where I worked on robust and nonparametric statistics for functional data. The idea is simple: data depth gives a center-outward ordering of a sample, so plotting the depths of one sample with respect to two different reference distributions (the DD-plot) turns a two-sample problem into a shape-comparison problem, and departures from the diagonal signal that the two samples differ. We build several statistics on top of that plot, calibrate them by bootstrap, and find they have higher power than competing homogeneity tests, illustrating everything on chemical spectrometry curves. I do not work on functional data anymore, but this is where I learned to think about distributions of infinite-dimensional objects, which turned out to be quite useful later.

Download paper here