This paper proposes a variant of the well-known boosting trees algorithm to estimate conditional distributions. Since regression trees partition observations into subgroups, the corresponding empirical distributions can be used to define the splitting criterion. Precisely, the parametric approach using Poisson deviance is replaced with a non-parametric one maximizing probabilistic distances between empirical distributions in child nodes. Proceeding inthis way, the actuary obtains an estimated conditional distribution for the response, from which a conditional mean can be derived as well as any other quantity of interest in risk management. The numerical performances of the proposed method are assessed with simulated data while a case study demonstrates its usefulness for insurance applications.
Denuit, M., Michaelides, M., Trufin, J., & Verelst, H. (2025). Wasserstein boosting trees algorithm for count data, with application to claim frequencies in motor insurance (LIDAM Discussion Paper ISBA 2025/24).