Machine Learning on Credit Risk for Lending Club Business
DOI:
https://doi.org/10.54691/bcpbm.v23i.1470Keywords:
Credit Risk, Feature Engineering, Machine Learning, Binary Classification.Abstract
Under the big data era, human beings attach much importance on machine learning. The purpose of this paper is to identify the way to use different methods of machine learning for data analysis and data mining when the companies in lending industry is faced with credit risk. To be specific, the methods are judged from the indicators (e.g., accuracy score and AUC score), and utilized to recognize which loans are bad loans, so as to facilitate the company to make better decisions. The methods including Random Forest, Gaussian Naive Bayes, and Artificial Neural Network are discussed. According to the analysis, all models have an accuracy around 60-70%, while each show different tendency in classifying results. Further optimization that can be applied in the future studies is suggested in the paper. The overall value and reputation of a company will be improved with good credit risk management. Therefore, a good method of credit risk management (e.g., accurately identifying good loans and bad loans) is classifying and analyzing the existing data of the company via different algorithms, and eventually compare them. These results shed light on guiding further exploration of evaluating the creditworthiness in the lending field.
Downloads
References
Basle Committee on Banking Supervision, Bank for International Settlements. Principles for the management of credit risk[M]. Bank for International Settlements, 2000.
Information on https://www.gdslink.com/what-are-the-different-types-of-credit-risk/
Kwabena A B M. Credit risk management in financial institutions: A case study of Ghana Commercial Bank Limited[J]. Research journal of finance and accounting, 2014, 5(23): 67-85.
Information on: https://corporatefinanceinstitute.com/resources/knowledge/finance/non-performing-loan-npl/
Information on: https://www.spglobal.com/marketintelligence/en/documents/machine_learning_and_credit_risk_modelling_november_2020.pdf
Breiman L. Random forests[J]. Machine learning, 2001, 45(1): 5-32.
Permission S. Generative and discriminative classifiers: naive bayes and logistic regression[J]. 2005.
Huang P J. Classification of imbalanced data using synthetic over-sampling techniques[M]. University of California, Los Angeles, 2015.
Himberg T. Loan Default Prediction with Machine Learning[J]. 2021.
Shi T, Horvath S. Unsupervised learning with random forest predictors[J]. Journal of Computational and Graphical Statistics, 2006, 15(1): 118-138.
Mohammed R, Rawashdeh J, Abdullah M. Machine learning with oversampling and undersampling techniques: overview study and experimental results[C]. 2020 11th international conference on information and communication systems (ICICS). IEEE, 2020: 243-248.
Dongare A D, Kharde R R, Kachare A D. Introduction to artificial neural network[J]. International Journal of Engineering and Innovative Technology (IJEIT), 2012, 2(1): 189-194.
Fawcett T. An introduction to ROC analysis[J]. Pattern recognition letters, 2006, 27(8): 861-874.
Rosebrock, A., 2022. Why is my validation loss lower than my training loss? - PyImageSearch. Information on: https://pyimagesearch.com/2019/10/14/why-is-my-validation-loss-lower-than-my-training-loss/> [Accessed 30 May 2022].
Kaggle.com. 2022. Lending Club Loan Defaulters Prediction. Information on: https://www.kaggle.com/code/faressayah/lending-club-loan-defaulters-prediction.






