P064-0001
Algorithmic detection of elemental biosignatures

Tuesday, 15 December 2020
Poster
Jesse Murray1,2, Aivaras Vilutis3,4, Thomas Stucky5,6, Michael Furlong7,8, Jessica E. Koehne6, David Mauro6,9, Annmarie Schramm10 and Diana Gentry6, (1)University of Oxford, Statistics, Oxford, United Kingdom, (2)NASA Ames Research Center, NASA Internships, Fellowships & Scholarships, Moffett Field, CA, United States, (3)Vilnius University, Vilnius, Lithuania, (4)NASA Ames Research Center, NASA International Internships Program, Moffett Field, CA, United States, (5)SETI Institute Mountain View, Mountain View, CA, United States, (6)NASA Ames Research Center, Moffett Field, CA, United States, (7)NASA Ames Research Center, Intelligent Robotics Group, Moffett Field, CA, United States, (8)Stinger Ghaffarian Technologies, Greenbelt, MD, United States, (9)Millennium Engineering & Integration, Moffett Field, CA, United States, (10)NASA Ames Research Center, Wyle Laboratories, Moffett Field, CA, United States
Abstract:
Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features.

Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization.

The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.