H078-08
Quantitative analysis of produced water in Permian Basin-New Mexico using machine learning techniques
Abstract:
This work utilizes different machine learning algorithms to predict the produced water quantity from different types of oil and gas wells (vertical, horizontal, and directional). The data for the research is collected from the New Mexico Oil Conservation Division (OCD) and contains information for 81,444 distinctive wells in the years of 1900-2020. The prediction considers multiple factors including well latitude, longitude, true vertical depth, measured vertical depth, year of operations, county, formation, and oil or gas production amount. Both linear regression and non-linear regression approaches (Random Forest approach in particular) are deployed to conduct the analysis. The Mean Absolute Error (MAE) and R2 scores are reported as measurement metrics. Two types of analysis were conducted. The first analysis examines the effect of each individual factor on the produced water production. The results suggest that the well location (latitude and longitude), true well vertical depth, measured well vertical depth, and produced oil/gas amount play important roles in the prediction process. The second analysis utilizes all the factors to make the prediction. Prediction results from five-fold cross-validation show that the Random Forest model reported high prediction accuracy, with the highest R2 Score (0.91) for horizontal gas wells and lowest R2 score (0.70) for directional gas wells. The lower accuracy of directional gas wells is due to the smaller number of useful wells in the prediction. The machine learning techniques provide a valuable approach to quantitatively analyze, characterize, and project produced water quantity in the Permian Basin for sustainable management of produced water.