Artificial Intelligence and Machine Learning in Banking: A Critical Narrative Synthesis of Application Domains, Methods and Emerging Research Gaps (2015–2024)
Omar Al-Kasasbeh *
Faculty of Humanities, Midocean University, Moroni, Comoros.
Mohammad Khaled Alkfaween
Faculty of Business, Al-Balqaʼ Applied University, As-Salt, Jordan.
*Author to whom correspondence should be addressed.
Abstract
Learning systems have moved from experimental pilots to embedded components of banking operations over the past decade, yet the scholarly evidence supporting that transition remains uneven in quality and unevenly distributed across the functions banks actually perform. This critical narrative review examines research published between 2015 and 2024 on the application of artificial intelligence and machine learning within banking, with attention to the methods employed, the strength of the evidence, and the questions that remain unresolved. The review synthesises peer-reviewed empirical and methodological work alongside supervisory and intergovernmental reporting, organising the literature around application domains, methodological practice, interpretability and fairness, and the governance interface rather than around individual studies. Five conclusions emerge. Predictive gains from complex learners over regularised parametric baselines are real but modest and highly sensitive to the evaluation criterion chosen; gains attributable to novel data sources appear larger and more durable than gains attributable to novel algorithms. Deep architectures have not established a consistent advantage over gradient-boosted trees on the tabular data that dominate retail banking. The evidence base is heavily concentrated in credit assessment and payment fraud detection, while treasury operations, supervisory analytics, internal control and customer relationship management are supported by markedly thinner empirical work. Recurrent methodological weaknesses, including reliance on small public benchmark datasets, temporally naive validation, informal handling of class imbalance, and evaluation detached from the economics of lending, limit confidence in reported performance differences and constrain their generalisability to deployed systems. Distributional and fairness consequences are demonstrated in specific mortgage and consumer credit markets but have not been established as general properties of algorithmic underwriting. Priority research needs include prospective evaluation of deployed systems, disclosure of temporal validation protocols, outcome measures aligned with portfolio economics and consumer welfare, and evidence from jurisdictions currently absent from the literature.
Keywords: Credit scoring, algorithmic decision-making, financial crime detection, model risk governance, explainable artificial intelligence, algorithmic fairness, banking supervision