MR023-0004
Application of machine learning for predicting gas reservoir recovery factor

Wednesday, 16 December 2020
Poster
Alireza Roustazadeh, Kansas State University, Manhattan, KS, United States, Behzad Ghanbarian, Kansas State University, Geology, Manhattan, KS, United States, Mohammad B Shadmand, Kansas State University, electrical engineering, Manhattan, KS, United States, Vahid Taslimitehrani, Realtor.com, Santa Clara, United States and Larry W Lake, University of Texas at Austin, Austin, TX, United States
Abstract:
Natural gas is probably the cleanest source of fossil fuel energy in the world. Furthermore, gas recovery factor typically is relatively high depending on reservoir characteristics. Accordingly, gas reservoirs have been the active area of exploitation and exploration. However, it is necessary to determine in advance whether a reservoir is economically feasible to be drilled or not. Conventional methods for predicting gas recovery factor are time consuming and budget exhaustive. With the advancements in machine learning technology and big data analytics, one may use machine learning approaches to predict the gas recovery factor. The goal of this study was to predict the ultimate gas recovery factors of different gas reservoirs from other characteristics, such as permeability, porosity, water saturation, reservoir pressure and other lithological properties and production data. For this purpose, a large dataset consisted of more than 1200 reservoirs was used to extract tree-based relationships. The dataset was first cleaned up, then sorted ascendingly based on the target variable (i.e., recovery factor) and next prepared for imputation for the missing data in the dataset. Missing values were imputed using a new method based on which each parameter was divided into subsamples with at least 10 entries and less than 10% of missing data. Next, the missing values in each subsample were replaced with the mode of the corresponding parameter. The entire database was randomly divided into 2/3 and 1/3 fractions for training and testing purposes respectively with 1000 iterations. After that, the Tree-Based Gradient Boost (XGBoost) method was applied to train and test the model. For the training and test, we found the average RMSLE = 0.067 and 0.08 and R2 = 0.76 and 0.64, respectively. Comparing the calculated statistical parameter values including the correlation coefficient R2 with those reported in the literature for databases with similar diversity showed a substantial improvement in the prediction of gas recovery factor.