Preprint / Version 1

Outcome Definition Before Model Choice in Credit Risk

Machine-Learning Models versus Ratio-Based Scorecards

##article.authors##

  • Saharsh Jangaon Polygence researcher
  • Darrell Robinson Research Mentor

DOI:

https://doi.org/10.58445/rars.4228

Keywords:

Credit risk, Credit scoring, Machine learning

Abstract

This paper examines how outcome definition affects the evaluation of traditional and machine-learning credit-risk models. Using U.S. Small Business Administration lending data, consumer-credit benchmarks, and controlled simulations, the study compares ratio-based credit scorecards with machine-learning models across discrimination, calibration, credit-allocation efficiency, and robustness to changes in the lending environment. The results show that machine-learning models can improve predictive performance over traditional scorecards, particularly in precision-recall performance and calibration. However, the analysis also demonstrates that model choice can matter less than how the credit outcome itself is defined. Restricting observations based on loan status or maturity can substantially alter measured default rates and produce apparently strong predictive performance while misrepresenting underlying credit risk. Additional analyses examine model stability across prediction horizons, calibration under distribution shift, and the relationship between predictive complexity and lending decisions. The findings suggest that financial institutions should define credit outcomes and observation windows before selecting predictive models, evaluate calibration alongside ranking metrics, and consider simpler models when their transparency and stability outweigh modest gains in predictive accuracy. Overall, the study shows that responsible adoption of AI in lending depends not only on choosing more powerful algorithms, but also on constructing the credit-risk problem correctly.

References

Akerlof, G. A. (1970). The market for "lemons": Quality uncertainty and the market mechanism. Quarterly Journal of Economics, 84(3), 488–500. https://doi.org/10.2307/1879431

Banasik, J., & Crook, J. (2007). Reject inference, augmentation, and sample selection. European Journal of Operational Research, 183(3), 1582–1594. https://doi.org/10.1016/j.ejor.2006.06.072

Barocas, S., & Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3), 671–732. https://doi.org/10.15779/Z38BG31

Bartlett, R., Morse, A., Stanton, R., & Wallace, N. (2022). Consumer-lending discrimination in the FinTech era. Journal of Financial Economics, 143(1), 30–56. https://doi.org/10.1016/j.jfineco.2021.05.047

Bellotti, T., & Crook, J. (2014). Retail credit stress testing using a discrete hazard model with macroeconomic factors. Journal of the Operational Research Society, 65(3), 340–350. https://doi.org/10.1057/jors.2013.91

Berg, T., Burg, V., Gombović, A., & Puri, M. (2020). On the rise of FinTechs: Credit scoring using digital footprints. Review of Financial Studies, 33(7), 2845–2897. https://doi.org/10.1093/rfs/hhz099

Björkegren, D., & Grissen, D. (2020). Behavior revealed in mobile phone usage predicts credit repayment. World Bank Economic Review, 34(3), 618–634. https://doi.org/10.1093/wber/lhz006

Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., & Liang, P. (2022). Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? In Advances in Neural Information Processing Systems 35. https://arxiv.org/abs/2211.13972

Botha, A., & Verster, T. (2026). Approaches for modelling the term-structure of default risk under IFRS 9: A tutorial using discrete-time survival analysis. International Journal of Data Science and Analytics, 22(1), Article 67. https://doi.org/10.1007/s41060-026-01032-w

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD (pp. 785–794). https://doi.org/10.1145/2939672.2939785

Coronavirus Aid, Relief, and Economic Security Act, Pub. L. No. 116-136, § 1112 (2020).

Dirick, L., Claeskens, G., & Baesens, B. (2017). Time to default in credit scoring using survival analysis: A benchmark study. Journal of the Operational Research Society, 68(6), 652–665. https://doi.org/10.1057/s41274-016-0128-9

Djeundje, V. B., & Crook, J. (2019). Dynamic survival models with varying coefficients for credit risks. European Journal of Operational Research, 275(1), 319–333. https://doi.org/10.1016/j.ejor.2018.11.029

Equal Credit Opportunity Act (Regulation B), 12 C.F.R. § 1002.9 (2024). https://www.ecfr.gov/current/title-12/part-1002/section-1002.9

Fuster, A., Goldsmith-Pinkham, P., Ramadorai, T., & Walther, A. (2022). Predictably unequal? The effects of machine learning on credit markets. Journal of Finance, 77(1), 5–47. https://doi.org/10.1111/jofi.13090

Grömping, U. (2019). South German credit data: Correcting a widely used data set (Report 04/2019). Beuth University of Applied Sciences Berlin.

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th ICML (PMLR 70, pp. 1321–1330).

Hand, D. J., & Adams, N. M. (2014). Selection bias in credit scorecard evaluation. Journal of the Operational Research Society, 65(3), 408–415. https://doi.org/10.1057/jors.2013.55

Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems 29 (pp. 3315–3323). https://proceedings.neurips.cc/paper/2016/hash/9d2682367c3935defcb1f9e247a97c0d-Abstract.html

Hofmann, H. (1994). Statlog (German Credit Data) [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5NC77

Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In NeurIPS 30 (pp. 3146–3154).

Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated errors in large language models. In Proceedings of the 42nd International Conference on Machine Learning (PMLR 267, pp. 30038–30066). https://proceedings.mlr.press/v267/kim25e.html

Kleinberg, J., & Raghavan, M. (2021). Algorithmic monoculture and social welfare. Proceedings of the National Academy of Sciences, 118(22), e2018340118. https://doi.org/10.1073/pnas.2018340118

Kozodoi, N., Lessmann, S., Alamgir, M., Moreira-Matias, L., & Papakonstantinou, K. (2025). Fighting sampling bias: A framework for training and evaluating credit scoring models. European Journal of Operational Research, 324(2), 616–628. https://doi.org/10.1016/j.ejor.2025.01.040

Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124–136. https://doi.org/10.1016/j.ejor.2015.05.030

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In NeurIPS 30 (pp. 4765–4774).

Medina-Olivares, V., Calabrese, R., Crook, J., & Lindgren, F. (2023). Joint models for longitudinal and discrete survival data in credit scoring. European Journal of Operational Research, 307(3), 1457–1473. https://doi.org/10.1016/j.ejor.2022.10.022

Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. In Proceedings of the AAAI Conference on Artificial Intelligence, 29(1), 2901–2907. https://doi.org/10.1609/aaai.v29i1.9602

Siddiqi, N. (2017). Intelligent credit scoring: Building and implementing better credit risk scorecards (2nd ed.). Wiley. https://doi.org/10.1002/9781119282396

Singer, J. D., & Willett, J. B. (1993). It’s about time: Using discrete-time survival analysis to study duration and the timing of events. Journal of Educational Statistics, 18(2), 155–195. https://doi.org/10.3102/10769986018002155

Stiglitz, J. E., & Weiss, A. (1981). Credit rationing in markets with imperfect information. American Economic Review, 71(3), 393–410.

U.S. Small Business Administration. (2021). 7(a) and 504 Section 1112 payment extension (Procedural Notice 5000-20079). https://www.sba.gov/document/procedural-notice-5000-20079-7a-504-section-1112-payment-extension

U.S. Small Business Administration. (2026). 7(a) & 504 FOIA [Data set]. https://data.sba.gov/dataset/7a-504-foia

Wang, H., Bellotti, T., Qu, R., & Bai, R. (2024). Discrete-time survival models with neural networks for age–period–cohort analysis of credit risk. Risks, 12(2), 31. https://doi.org/10.3390/risks12020031

Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the Eighth ACM SIGKDD (pp. 694–699). https://doi.org/10.1145/775047.775151

Downloads

Posted

2026-10-04

Categories