Files

ISBA_DP_2023-02.pdf
  • Open Access
  • Adobe PDF
  • 3.16 MB

Details

Authors
Abstract
The k-means algorithm and its variants are popular clustering techniques. Their purpose is to uncover group structures in a dataset. In actuarial applications, these partitioning methods detect clusters of policies with similar features and allow one to draw up a map of dominant risks. The main challenge lies in de􏰂ning a distance between two observations exclusively characterised by categorical variables. This research paper starts with a review of the k-means algorithm and develops an extension based on Burt's framework to manage categorical rating factors. We then focus on a mini-batch version that keeps computation time under control when analysing a large-scale dataset. We next broaden the scope of application of the fuzzy k-means to fully categorised datasets. Lastly, we conclude with a thorough introduction to spectral clustering and work around the dimensionality issue by reducing the size of the initial dataset with k-means.
Affiliations

Citations

Jamotton, C., Hainaut, D., & Hames, T. (2023). Insurance analytics with clustering techniques (LIDAM Discussion Paper ISBA 2023/02). https://hdl.handle.net/2078.5/128065