H195-0016
The Wisdom of the Crowd in Probabilistic Predictive Modelling: Large-Scale Application to Monthly Rainfall-Runoff Problems
The Wisdom of the Crowd in Probabilistic Predictive Modelling: Large-Scale Application to Monthly Rainfall-Runoff Problems
Wednesday, 16 December 2020
Poster
Abstract:
Probabilistic rainfall-runoff modelling is often performed with hydrological post-processing methodologies. A distinguishing feature of these methodologies is the allowance for exploitation of the information provided by process-based (else referred to as “physically-based”) hydrological models. This information takes the form of point predictions. We here present and extensively test a new two-stage methodology of this family. Adopting key concepts from a theoretically consistent blueprint for probabilistic hydrological modelling, this new methodology allows the utilization of flexible (e.g., machine learning) quantile regression algorithms under an ensemble learning strategy. Its main differences with respect to basic two-stage hydrological post-processing methodologies using the same type of regression algorithms are that (a) instead of a single point hydrological prediction it generates a large number of “sister predictions” (yet using a single process-based model), and that (b) it relies on the concept of combining probabilistic predictions via simple quantile averaging (simplest form of probabilistic ensemble learning). We conduct a large-scale benchmark experiment by using complete 50-year-long total monthly time series observed in 270 catchments in the contiguous United States. For each catchment, we combine numerous probabilistic predictions via simple quantile averaging. We assess the individual probabilistic predictions and their combinations using three scores, i.e., the reliability score, the average interval width and the average interval score. The equal-weight combiner is theoretically expected and empirically shown to offer larger robustness in performance than the individual predictors. It is also empirically proven to harness the “wisdom of the crowd” in terms of average interval score, i.e., the average of the individual predictions scores no worse –usually better− than the average of the scores of the same predictions.