SY011-0013
Bayesian quantification of errors in citizen science data: Application to Nepal rainfall observations
Bayesian quantification of errors in citizen science data: Application to Nepal rainfall observations
Tuesday, 8 December 2020
Poster
Abstract:
High quality citizen science can be instrumental in advancing scientific discoveries and developing a deeper understanding of under-observed phenomena. However, the error structure of citizen scientist data must be well-defined. Within a citizen science program, the error types in submitted observations vary, and their occurrence may depend on a variety of citizen scientist-specific variables, such as motivation. This study develops a graphical Bayesian inference model of error types in citizen science data. The model assumes that: (1) each citizen scientist observation is subject to a specific error type, each with its own bias and noise; and (2) an observation's error type depends on the error community of the citizen scientist, which in turn relates to characteristics of the citizen scientist submitting the observation. Given a set of citizen scientist observations and corresponding ground-truth values, the model can be calibrated for a specific application, yielding (i) the number of error types and communities, (ii) the bias and noise of each error type, (iii) the error distribution of each community, and (iv) the community to which each citizen scientist belongs. The model, applied to Nepal citizen scientist rainfall observations, identifies seven error types and sorts citizen scientists into four model-inferred communities. In the case study, 79% of citizen scientists committed errors in fewer than 6.3% of their observations. The remaining tended to commit unit, meniscus, and unknown errors. A citizen scientist’s assigned community, coupled with the model-inferred error probability, can identify observations that require verification. With such a system, the onus of validating citizen scientist data is partially transferred from human effort to machine-learned algorithms.