Abstract
A central challenge in financial risk modelling arises from the heterogeneity of financial data. In many real world settings, the characteristics of borrowers, firms, or transactions differ across multiple dimensions such as product type, demographic attributes, industry sectors, time periods, and geographical locations. These differences may lead to distinct behavioural patterns and varying relationships between explanatory variables and financial outcomes. When such heterogeneity exists, models built on a single unified population may fail to adequately capture subgroup specific patterns, potentially limiting both predictive accuracy and interpretability. Understanding how to account for this heterogeneity is therefore an important issue in financial decision making, particularly in applications such as credit risk assessment and corporate fraud detection where classification models are widely used to support operational decisions. Segmentation offers a natural approach to address this challenge by partitioning the population into more homogeneous groups, allowing models to better reflect the underlying structure of the data.Despite its widespread use in practice, the methodological role and empirical effectiveness of segmentation in financial risk modelling remain insufficiently documented in the academic literature, motivating the need for a more systematic investigation. In response to this gap, this thesis examines how segmentation based modelling can be incorporated into different stages of financial risk modelling, including explanatory analysis, predictive analysis, and methodological development. By studying segmentation from these complementary perspectives, the research aims to evaluate how structured population partitioning can improve model performance, enhance interpretability, and provide a clearer understanding of heterogeneous risk patterns in applications such as credit risk assessment and corporate fraud detection.
We begin by studying the effects of segmentation in explanatory analysis, where we utilise a microfinance dataset obtained from the Chinese informal finance market. Segmentation is applied to separate business loan applicants from non-business loan applicants, where comparison is carried out to study their respective characteristics. In terms of model building, we incorporate interaction terms as a means of segmentation into logistic regression models to investigate the determinants of loan funding success and default. Our findings suggest that segmentation reveals distinct risk patterns for different borrower types, thus improving the interpretability and accuracy of explanatory models.
We then expand the application of segmentation in predictive analysis by implementing a multidimensional segmentation by fraud type, industry, and time on a Chinese corporate fraud dataset. We compare the predictive performance between segmented models and a general pooled model and investigate the differences in risk features for each fraud type and industry. In addition, we address the population drift problem by selecting the optimal amount of training data, which is a form of segmentation by time. Our results demonstrate that segmentation not only enhances the predictive performance, interpretability, and effectiveness of predictive models but also provides a viable strategy for addressing evolving data distributions. This emphasises the value of segmentation as both a practical and methodological tool.
In terms of methodological development, we focus deeper on time-based segmentation and develop a methodological innovation by incorporating time-dependent weights to the loss function of the neural network model to mitigate population drift. The proposed solution automates segmentation by time through the identification of the optimal breakpoint, which is carried out simultaneously with model training. This approach addresses the population drift problem of credit risk assessment and fraud detection models, while also improving their predictive performance and computational efficiency.
In summary, this thesis demonstrates that segmentation provides a valuable framework for improving financial risk modelling across explanatory, predictive, and methodological contexts. The results confirm that segmentation not only uncovers hidden risk patterns but also improves model stability, interpretability, and predictive effectiveness. The proposed framework therefore offers a systematic approach for integrating segmentation into financial risk modelling and demonstrates its broad applicability to credit risk assessment and corporate fraud detection.
| Date of Award | 15 Jun 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Anthony Graham Bellotti (Supervisor) & Xiuping Hua (Supervisor) |
Cite this
- Standard