Outcome Definition Before Model Choice in Credit Risk
Machine-Learning Models versus Ratio-Based Scorecards
DOI:
https://doi.org/10.58445/rars.4228Keywords:
Credit risk, Credit scoring, Machine learningAbstract
This paper examines how outcome definition affects the evaluation of traditional and machine-learning credit-risk models. Using U.S. Small Business Administration lending data, consumer-credit benchmarks, and controlled simulations, the study compares ratio-based credit scorecards with machine-learning models across discrimination, calibration, credit-allocation efficiency, and robustness to changes in the lending environment. The results show that machine-learning models can improve predictive performance over traditional scorecards, particularly in precision-recall performance and calibration. However, the analysis also demonstrates that model choice can matter less than how the credit outcome itself is defined. Restricting observations based on loan status or maturity can substantially alter measured default rates and produce apparently strong predictive performance while misrepresenting underlying credit risk. Additional analyses examine model stability across prediction horizons, calibration under distribution shift, and the relationship between predictive complexity and lending decisions. The findings suggest that financial institutions should define credit outcomes and observation windows before selecting predictive models, evaluate calibration alongside ranking metrics, and consider simpler models when their transparency and stability outweigh modest gains in predictive accuracy. Overall, the study shows that responsible adoption of AI in lending depends not only on choosing more powerful algorithms, but also on constructing the credit-risk problem correctly.
References
Akerlof, G. A. (1970). The market for "lemons": Quality uncertainty and the market mechanism. Quarterly Journal of Economics, 84(3), 488–500. https://doi.org/10.2307/1879431
Banasik, J., & Crook, J. (2007). Reject inference, augmentation, and sample selection. European Journal of Operational Research, 183(3), 1582–1594. https://doi.org/10.1016/j.ejor.2006.06.072
Barocas, S., & Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3), 671–732. https://doi.org/10.15779/Z38BG31
Bartlett, R., Morse, A., Stanton, R., & Wallace, N. (2022). Consumer-lending discrimination in the FinTech era. Journal of Financial Economics, 143(1), 30–56. https://doi.org/10.1016/j.jfineco.2021.05.047
Bellotti, T., & Crook, J. (2014). Retail credit stress testing using a discrete hazard model with macroeconomic factors. Journal of the Operational Research Society, 65(3), 340–350. https://doi.org/10.1057/jors.2013.91
Berg, T., Burg, V., Gombović, A., & Puri, M. (2020). On the rise of FinTechs: Credit scoring using digital footprints. Review of Financial Studies, 33(7), 2845–2897. https://doi.org/10.1093/rfs/hhz099
Björkegren, D., & Grissen, D. (2020). Behavior revealed in mobile phone usage predicts credit repayment. World Bank Economic Review, 34(3), 618–634. https://doi.org/10.1093/wber/lhz006
Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., & Liang, P. (2022). Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? In Advances in Neural Information Processing Systems 35. https://arxiv.org/abs/2211.13972
Botha, A., & Verster, T. (2026). Approaches for modelling the term-structure of default risk under IFRS 9: A tutorial using discrete-time survival analysis. International Journal of Data Science and Analytics, 22(1), Article 67. https://doi.org/10.1007/s41060-026-01032-w
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD (pp. 785–794). https://doi.org/10.1145/2939672.2939785
Coronavirus Aid, Relief, and Economic Security Act, Pub. L. No. 116-136, § 1112 (2020).
Dirick, L., Claeskens, G., & Baesens, B. (2017). Time to default in credit scoring using survival analysis: A benchmark study. Journal of the Operational Research Society, 68(6), 652–665. https://doi.org/10.1057/s41274-016-0128-9
Djeundje, V. B., & Crook, J. (2019). Dynamic survival models with varying coefficients for credit risks. European Journal of Operational Research, 275(1), 319–333. https://doi.org/10.1016/j.ejor.2018.11.029
Equal Credit Opportunity Act (Regulation B), 12 C.F.R. § 1002.9 (2024). https://www.ecfr.gov/current/title-12/part-1002/section-1002.9
Fuster, A., Goldsmith-Pinkham, P., Ramadorai, T., & Walther, A. (2022). Predictably unequal? The effects of machine learning on credit markets. Journal of Finance, 77(1), 5–47. https://doi.org/10.1111/jofi.13090
Grömping, U. (2019). South German credit data: Correcting a widely used data set (Report 04/2019). Beuth University of Applied Sciences Berlin.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th ICML (PMLR 70, pp. 1321–1330).
Hand, D. J., & Adams, N. M. (2014). Selection bias in credit scorecard evaluation. Journal of the Operational Research Society, 65(3), 408–415. https://doi.org/10.1057/jors.2013.55
Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems 29 (pp. 3315–3323). https://proceedings.neurips.cc/paper/2016/hash/9d2682367c3935defcb1f9e247a97c0d-Abstract.html
Hofmann, H. (1994). Statlog (German Credit Data) [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5NC77
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In NeurIPS 30 (pp. 3146–3154).
Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated errors in large language models. In Proceedings of the 42nd International Conference on Machine Learning (PMLR 267, pp. 30038–30066). https://proceedings.mlr.press/v267/kim25e.html
Kleinberg, J., & Raghavan, M. (2021). Algorithmic monoculture and social welfare. Proceedings of the National Academy of Sciences, 118(22), e2018340118. https://doi.org/10.1073/pnas.2018340118
Kozodoi, N., Lessmann, S., Alamgir, M., Moreira-Matias, L., & Papakonstantinou, K. (2025). Fighting sampling bias: A framework for training and evaluating credit scoring models. European Journal of Operational Research, 324(2), 616–628. https://doi.org/10.1016/j.ejor.2025.01.040
Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124–136. https://doi.org/10.1016/j.ejor.2015.05.030
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In NeurIPS 30 (pp. 4765–4774).
Medina-Olivares, V., Calabrese, R., Crook, J., & Lindgren, F. (2023). Joint models for longitudinal and discrete survival data in credit scoring. European Journal of Operational Research, 307(3), 1457–1473. https://doi.org/10.1016/j.ejor.2022.10.022
Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. In Proceedings of the AAAI Conference on Artificial Intelligence, 29(1), 2901–2907. https://doi.org/10.1609/aaai.v29i1.9602
Siddiqi, N. (2017). Intelligent credit scoring: Building and implementing better credit risk scorecards (2nd ed.). Wiley. https://doi.org/10.1002/9781119282396
Singer, J. D., & Willett, J. B. (1993). It’s about time: Using discrete-time survival analysis to study duration and the timing of events. Journal of Educational Statistics, 18(2), 155–195. https://doi.org/10.3102/10769986018002155
Stiglitz, J. E., & Weiss, A. (1981). Credit rationing in markets with imperfect information. American Economic Review, 71(3), 393–410.
U.S. Small Business Administration. (2021). 7(a) and 504 Section 1112 payment extension (Procedural Notice 5000-20079). https://www.sba.gov/document/procedural-notice-5000-20079-7a-504-section-1112-payment-extension
U.S. Small Business Administration. (2026). 7(a) & 504 FOIA [Data set]. https://data.sba.gov/dataset/7a-504-foia
Wang, H., Bellotti, T., Qu, R., & Bai, R. (2024). Discrete-time survival models with neural networks for age–period–cohort analysis of credit risk. Risks, 12(2), 31. https://doi.org/10.3390/risks12020031
Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the Eighth ACM SIGKDD (pp. 694–699). https://doi.org/10.1145/775047.775151
Downloads
Posted
Categories
License
Copyright (c) 2026 Research Archive of Rising Scholars

This work is licensed under a Creative Commons Attribution 4.0 International License.