H172-0002
A transformation-based method for automatic threshold selection in Peaks Over Threshold (POT) approach
A transformation-based method for automatic threshold selection in Peaks Over Threshold (POT) approach
Tuesday, 15 December 2020
Poster
Abstract:
A popular choice for frequency analysis of extremes (e.g., storms, floods), especially in data sparse situations, is Peaks Over Threshold (POT) approach. An unresolved issue in using the approach is identification of optimal threshold which ensures extraction of maximum information on extremes from a given time-series, without violating assumptions of the underlying extreme value theory. This paper contributes a novel transformation-based automatic threshold selection (TS) method. It involves mapping peaks over tentative thresholds (depicted by Generalized Pareto distributed random variable) from the original space to a non-dimensional space using a proposed transformation, where the mapped peaks theoretically follow standard exponential distribution. Optimal threshold is identified as that which minimizes Mahalanobis distance between sample L-statistics of the mapped peaks and those of the population (i.e., standard exponential distribution) in the non-dimensional space. To reduce computational effort in determining the population L-statistics, their asymptotic values are expressed as functions of sample size through Monte-Carlo simulations (MCS). Effectiveness of the proposed method is demonstrated over four existing automatic TS methods through MCS experiments and case studies over rainfall and streamflow datasets chosen from India, United Kingdom and Australia. The four methods include three based on goodness of fit (GoF) test statistics (of Anderson-Darling, and two non-parametric tests), and a recent one based on L-moment ratio diagram (LMRD) whose potential is unexplored in hydrology. This paper further provides insight into properties and effectiveness of the four TS methods, which is scanty in literature. The trade-offs between computational effort and accuracy gain/loss with each method are also assessed for various sample sizes. Results indicate that there is inconsistency in performance of GoF test-based methods across datasets exhibiting fat and thin tail behaviour, owing to their theoretical assumptions and uncertainty associated with sampling distribution of test statistics. Issues affecting performance of LMRD-based TS method are also identified. The proposed method overcomes those issues, while being computationally more efficient.