IN011-07
A surrogate modeling strategy to learn ecosystem control points

Tuesday, 8 December 2020: 19:18
Virtual
Karl Bernard Bernard Schaettle1, Elijah Hoffman2, Nicola Falco3, Baptiste Dafflon3, Craig Ulrich4, Karina Nugent4, Jay McEntire5, Haruko M Wainwright3 and James Bentley Brown3,6, (1)University of California Berkeley, Departments of Chemistry, Bioengineering, Chemical and Biomolecular Engineering, Berkeley, CA, United States, (2)Lawrence Berkeley National Laboratory, Berkeley, United States, (3)Lawrence Berkeley National Laboratory, Berkeley, CA, United States, (4)Lawrence Berkeley National Laboratory, Earth and Environmental Sciences, Berkeley, CA, United States, (5)ARVA Intelligence, Salt Lake City, UT, United States, (6)University of California Berkeley, Berkeley, CA, United States
Abstract:
Ecosystem control points are spatiotemporally localized processes that contribute substantially to particular functions. Learning control points directly from data is a central pursuit in molecular ecosystems biology. The science of “EcoImaging” – acquiring high-resolution, multi-modal, multi-scale, above- and below-ground measurements – now provides an unprecedented opportunity for discovery. We analogize EcoImaging to ‘omics technologies in biology – the goal is to measure as much of the system, in its totality, as possible – it is a “hypothesis free” approach. Here, we take advantage of an extensive dataset compiled at the AR1K.org field site from 2017 to 2019 for monoculture systems of soy along with their microbes and soil contexts. Monoculture lands constitute exceptional “reduced order” model ecosystems, and hence useful testbeds for new data science tools. We developed a three-stage machine learning algorithm for the discovery of control points for target ecosystem services – and here we focus on agricultural yield. In the first stage, an iterative Random Forest (iRF) is used to extract important interactions and processes that are predictive of yield. In the second stage, we use Multivariate Adaptive Regression Splines (MARS) to model interactions using functions that are differentiable almost everywhere. Finally, we fit a Reduced Order Surrogate Model (ROSM) by performing forward-backward regression under a L2 loss where each term in the model is itself a MARS response surface. The resulting hybrid learning machine achieves comparable performance to the iRF from which it is derived, and captures explicit relationships suitable for human exploration. We call our technique Surrogate Models through iRF (SMiRF), and here we describe it’s utility in obtaining a predictive understanding of yield in terms of ecosystem control points at the AR1K.org field lab. In future work, we will pursue the use of SMiRFs to construct mechanistic process models from data, and we describe some of these directions.