H141-0028
Using k-means cluster analysis to regionalize precipitation in the Tropical Andes

Monday, 14 December 2020
Poster
Mario Cordova, Johanna Orellana-Alvear and Rolando Célleri, Universidad de Cuenca, Departamento de Recursos Hídricos y Ciencias Ambientales, Cuenca, Ecuador
Abstract:
Regionalization of precipitation is fundamental for hydrologic and climate applications; however, it is not a trivial task. In the Tropical Andes (TA), this is especially challenging because of the geographical features and climate drivers that give rise to great zonal and meridional gradients: both the wettest (Chocó) and the driest (Atacama Desert) regions in the world are within the TA, also, a vast contrast exists between the dry western and the wet eastern slopes of the TA. In spite of the importance of defining homogenous regions of precipitation, this has not yet been done for the TA. In this context, classification machine learning algorithms emerge as fitting tools to regionalize precipitation. In special the k-means clustering algorithm, which has been successfully applied in other regions of the world for this purpose in the past, both because its capacity to handle unlabeled data and its rapid convergence even when large amounts of data are processed. Therefore, in this study, we used the k-means algorithm to regionalize precipitation in the TA. The input data were: i) the annual cycle of precipitation, calculated for the 1931 – 2016 period, using the 0.25° GPCC dataset for the pixels above 300 m a.s.l. in the TA, and ii) the latitude and longitude coordinates of each pixel to improve the spatial cohesion of the clusters. Additionally, we applied the elbow method to determine the optimum number of clusters, which was found to be 10. Our results show spatially compact precipitation regions that align well with large-scale geographical and circulation features. For instance, we identified the following regions in different clusters: the arid Atacama Desert; the wet Chocó; the contrasting dry and wet areas to the west and east of the Peruvian Andes; the rainfall hotspots located in the eastern Andes of Southern Peru and Northern Bolivia; and the equatorial Andes, which have a distinctive bimodal regime. Therefore, we conclude that the k-means cluster analysis is suitable to identify precipitation regions that are influenced by different climate drivers and circulation patterns in the TA. Currently, we are evaluating different machine learning algorithms to gain new insights on the teleconnections and circulation patterns that drive precipitation processes in each of the different clusters identified here.