Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

A Comparative Study of Feature Selection Techniques for Customer Churn Prediction

Merit Chinonso Opara, Benjamin Chiemeka Opara, Uchechi Joyce Nneji, Chibueze, Favour Aririguzo, Confidence Chigozirim Olumba, Chidimma Grace Emmanuel, Prisca Chimezie Opara, Happy Nkanta Monday

Abstract

This paper investigates the effectiveness of feature selection techniques in optimizing supervised machine learning pipelines for customer churn prediction using the publicly available Customer Churn Dataset from Kaggle. Feature selection plays a crucial role in enhancing model interpretability and generalization by eliminating irrelevant or redundant variables. In this study, three optimized feature selection algorithms—Minimum Redundancy-Maximum Relevance , Multi Spatially Uniform Relief (MultiSURF), and Hilbert-Schmidt Independence Criterion —were implemented and compared alongside a baseline scenario without feature selection. Eight supervised machine learning classifiers, including Support Vector Machine , Logistic Regression, K-Nearest Neighbors , Decision Tree, AdaBoost, Bagging, Stacking, and Voting Classifier, were optimized using Randomized GridSearchCV. The dataset was divided into training and testing subsets with a 70:30 ratio, and model performance was evaluated using metrics such as accuracy, precision, recall, F1-score, ROC-AUC, and computational time. Experimental results reveal that all feature selection techniques achieved comparable performance across classifiers. This study concludes that optimized feature selection techniques contribute to more stable and efficient model training. P-ISSN 2695-186X

Keywords

Customer Churn Prediction; Feature Selection; Supervised Machine Learning; mRMR; MultiSURF; Hilbert–Schmidt Independence Criterion ; Ensemble Classifiers; Randomized GridSearchCV

References

Urbanowicz, R. J., Meeker, M., La Cava, W., Olson, R. S., & Moore, J. H. (2018). Relief-based feature selection: Introduction and review. Journal of Biomedical Informatics, 85, 189– 203. https://doi.org/10.1016/j.jbi.2018.07.014 Greene, C. S., Penrod, N. M., Kiralis, J., & Moore, J. H. (2009). Spatially uniform ReliefF for computationally efficient filtering of gene–gene interactions. BioData Mining, 2(1), Article 5. https://doi.org/10.1186/1756-0381-2-5 Peng, H., Long, F., & Ding, C. (2005). Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(8), 1226–1238. https://doi.org/10.1109/TPAMI.2005.159 Gretton, A., Bousquet, O., Smola, A., & Schölkopf, B. (2005). Measuring statistical dependence with Hilbert–Schmidt norms. In S. Jain, H. U. Simon, & E. Tomita (Eds.), Algorithmic learning theory (Lecture Notes in Computer Science, Vol. 3734, pp. 63–77). Springer. https://doi.org/10.1007/11564089_7 Kafrawy, P. E., Fathi, H., Qaraad, M., Kelany, A. K., & Chen, X. (2021). An efficient SVM- based feature selection model for cancer classification using high-dimensional microarray data. IEEE Access, 9, 155353–155369. https://doi.org/10.1109/ACCESS.2021.3123090 Peterson, L. (2009). K-nearest neighbor. Scholarpedia, 4(2), 1883. https://doi.org/10.4249/scholarpedia.1883 Song, Y., & Lu, Y. (2015). Decision tree methods: Applications for classification and prediction. Journal of Software Engineering, 27(2). Margineantu, D. D., & Dietterich, T. G. (1997). Pruning adaptive boosting. In Proceedings of the 14th International Conference on Machine Learning (pp. 211–218). Bauer, E., & Kohavi, R. (1999). An empirical comparison of voting classification algorithms: Bagging, boosting, and variants. Machine Learning, 36(1–2), 105–139. Džeroski, S., & Ženko, B. (2004). Is combining classifiers with stacking better than selecting the best one? Machine Learning, 54(3), 255–273. https://doi.org/10.1023/B:MACH.0000015881.36452.6e Breiman, L. (1996). Bagging predictors. Machine Learning, 24(2), 123–140. https://doi.org/10.1007/BF00058655 Ahmad, G. N., Fatima, H., Ullah, S., Saidi, A. S., & Imdadullah. (2022). Efficient medical diagnosis of human heart diseases using machine learning techniques with and without GridSearchCV. IEEE Access, 10, 80151–80173. https://doi.org/10.1109/ACCESS.2022.3165792

More Articles from IIARD INTERNATIONAL JOURNAL OF ECONOMICS AND BUSINESS MANAGEMENT