Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Distribution-Free Error Control for Fraud Alert Triage: Conformal Anomaly P-Values Under Class Imbalance, Concept Drift, and Calibration Contamination

Olamiji Onafowokan1, Olawale Fadugba2, Deborah Okunola1, Alex Mendy1

Abstract

Operational fraud detection almost always terminates in a fixed-quantile threshold applied to a model score, a rule that fixes alert volume but leaves the statistical properties of the alert list undefined. We study an alternative in which the alert list is produced by a multiple-testing procedure applied to split-conformal anomaly p-values, so that the expected proportion of false alarms is controlled at a level chosen in advance. Using a fully synthetic transaction generator with a hard-to-detect mimicry subpopulation, we quantify the behaviour of this design under the three conditions that characterize real fraud environments: extreme class imbalance, covariate drift, and contamination of the calibration sample by undetected fraud. Across twenty replications and four prevalence levels spanning 20 to 200 basis points, the Benjamini–Hochberg procedure applied to conformal p-values held the realized false discovery rate near its nominal 10 percent level (0.067 to 0.139) and delivered stable alert precision between 0.86 and 0.93, whereas a conventional top-one-percent score rule with the same detector produced precision ranging from 0.28 to 0.80 as prevalence varied, with realized false discovery proportions as high as 0.72. Adaptive null-proportion estimation gave no measurable power gain at these prevalences. The procedure is, however, fragile in two specific and quantifiable ways. Modest covariate drift broke exchangeability and inflated the realized false discovery rate from 0.09 to 0.63 at a drift magnitude that shifted the log-amount distribution by one fifth of a standard deviation, while recalibration on a recent window restored validity at a severe cost in power. Contamination of the calibration sample at only 50 basis points eliminated detection entirely by inflating the reference distribution. We show that trimming the upper tail of the calibration scores recovers most of the lost power, but only when the trimming fraction is strictly smaller than the true contamination rate: trimming at exactly the contamination rate inflated the false discovery rate to 0.40, and trimming at three times that rate to 0.80. We also find that a Kolmogorov–Smirnov uniformity test on the null p-values is a sensitive monitor for drift but a weak monitor for contamination, and we translate these results into a deployment protocol. All results are generated from synthetic data; no real transaction records were used.

Keywords

conformal prediction; false discovery rate; anomaly detection; class imbalance; concept drift; calibration contamination; fraud analytics; alert triage IJASMT E- ISSN 2489-009X

References

Abdallah, A., Maarof, M. A., & Zainal, A. (2016). Fraud detection system: A survey. Journal of Network and Computer Applications, 68, 90–113. Angelopoulos, A. N., & Bates, S. (2023). Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16(4), 494–591. Bahnsen, A. C., Aouada, D., Stojanovic, A., & Ottersten, B. (2016). Feature engineering strategies for credit card fraud detection. Expert Systems with Applications, 51, 134–142. Barber, R. F., Candès, E. J., Ramdas, A., & Tibshirani, R. J. (2023). Conformal prediction beyond exchangeability. Annals of Statistics, 51(2), 816–845. Bates, S., Candès, E., Lei, L., Romano, Y., & Sesia, M. (2023). Testing for outliers with conformal p-values. Annals of Statistics, 51(1), 149–178. Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B, 57(1), 289–300. Benjamini, Y., & Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under dependency. Annals of Statistics, 29(4), 1165–1188. Bhattacharyya, S., Jha, S., Tharakunnel, K., & Westland, J. C. (2011). Data mining for credit card fraud: A comparative study. Decision Support Systems, 50(3), 602–613. Bolton, R. J., & Hand, D. J. (2002). Statistical fraud detection: A review. Statistical Science, 17(3), 235–255. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. Carcillo, F., Le Borgne, Y.-A., Caelen, O., Kessaci, Y., Oblé, F., & Bontempi, G. (2021). Combining unsupervised and supervised learning in credit card fraud detection. Information Sciences, 557, 317–331. Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit card fraud detection: A realistic modeling and a novel learning strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784–3797. Davis, J., & Goadrich, M. (2006). The relationship between precision-recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning (pp. 233–240). Efron, B. (2004). Large-scale simultaneous hypothesis testing: The choice of a null hypothesis. Journal of the American Statistical Association, 99(465), 96–104. Elkan, C. (2001). The foundations of cost-sensitive learning. In Proceedings of the 17th International Joint Conference on Artificial Intelligence (pp. 973–978). Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1–37. Hand, D. J. (2009). Measuring classifier performance: A coherent alternative to the area under the ROC curve. Machine Learning, 77(1), 103–123. Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., ... Oliphant, T. E. (2020). Array programming with NumPy. Nature, 585, 357–362. Hilal, W., Gadsden, S. A., & Yawney, J. (2022). Financial fraud: A review of anomaly detection techniques and recent advances. Expert Systems with Applications, 193, 116429. IJASMT E- ISSN 2489-009X , Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). Isolation forest. In Proceedings of the 8th IEEE International Conference on Data Mining (pp. 413–422). Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., ... Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE, 10(3), e0118432. Storey, J. D. (2002). A direct approach to false discovery rates. Journal of the Royal Statistical Society, Series B, 64(3), 479–498. Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., ... van Mulbregt, P. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17, 261–272. Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic learning in a random world. New York: Springer. Wedge, R., Kanter, J. M., Veeramachaneni, K., Rubio, S. M., & Perez, S. I. (2019). Solving the false positives problem in fraud prediction using automated feature engineering. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2018) (pp. 372– 388).