Distribution-Free Error Control for Fraud Alert Triage: Conformal Anomaly P-Values Under Class Imbalance, Concept Drift, and Calibration Contamination
Abstract
Operational fraud detection almost always terminates in a fixed-quantile threshold applied to a model score, a rule that fixes alert volume but leaves the statistical properties of the alert list undefined. We study an alternative in which the alert list is produced by a multiple-testing procedure applied to split-conformal anomaly p-values, so that the expected proportion of false alarms is controlled at a level chosen in advance. Using a fully synthetic transaction generator with a hard-to-detect mimicry subpopulation, we quantify the behaviour of this design under the three conditions that characterize real fraud environments: extreme class imbalance, covariate drift, and contamination of the calibration sample by undetected fraud. Across twenty replications and four prevalence levels spanning 20 to 200 basis points, the Benjamini–Hochberg procedure applied to conformal p-values held the realized false discovery rate near its nominal 10 percent level (0.067 to 0.139) and delivered stable alert precision between 0.86 and 0.93, whereas a conventional top-one-percent score rule with the same detector produced precision ranging from 0.28 to 0.80 as prevalence varied, with realized false discovery proportions as high as 0.72. Adaptive null-proportion estimation gave no measurable power gain at these prevalences. The procedure is, however, fragile in two specific and quantifiable ways. Modest covariate drift broke exchangeability and inflated the realized false discovery rate from 0.09 to 0.63 at a drift magnitude that shifted the log-amount distribution by one fifth of a standard deviation, while recalibration on a recent window restored validity at a severe cost in power. Contamination of the calibration sample at only 50 basis points eliminated detection entirely by inflating the reference distribution. We show that trimming the upper tail of the calibration scores recovers most of the lost power, but only when the trimming fraction is strictly smaller than the true contamination rate: trimming at exactly the contamination rate inflated the false discovery rate to 0.40, and trimming at three times that rate to 0.80. We also find that a Kolmogorov–Smirnov uniformity test on the null p-values is a sensitive monitor for drift but a weak monitor for contamination, and we translate these results into a deployment protocol. All results are generated from synthetic data; no real transaction records were used.
Keywords
References
More Articles from INTERNATIONAL JOURNAL OF APPLIED SCIENCES AND MATHEMATICAL THEORY
Author: Samson Yunusa, Imande Terdon Terlumun,, Etuk Emmanuel Dan, and Micah Michael
Author: Ejes, Valentine, Liberty Ebiwareme
Author: Uwakwe, J. I., Anyanwu, E., Inalegwu, N.
Author: Xi Yang, Wenzhuo Zhang
Author: Daewii, Promise Saro, Godwin Lebari Tuaneh
