References
for institutions considering implementation. Layer Primary Owner/Stakeholder Key Risk if Layer Fails 1. Ingestion Data engineering / IT operations Missed or delayed transactions; incomplete regulatory data trail 2. Feature Engineering Data science / analytics engineering Weak or missing relational signal; degraded downstream detection 3. Detection Engine Model development team Missed typologies; excessive unproductive alerts 4. Explainability Model validation / compliance analytics Undocumented decisions; report narratives lack basis; validation gaps 5. Human-in-the- Loop Compliance operations / case management Analyst overload or under-review of high- risk cases 6. Governance Model risk management / internal audit Undetected drift or disparate impact; loss of audit trail Table 5. Layer ownership and failure-risk summary. 6.5 Cost and Resourcing Considerations Although this paper does not attempt a quantified cost-benefit analysis, the structure of the framework has implications for how such an analysis should be organized. The layers are not equally expensive to build or to operate, and the ordering of investment affects when returns materialize. Layers 1 through 3, ingestion, feature engineering, and detection, typically represent the largest upfront capital investment, since they require streaming infrastructure and model development effort. Layers 4 through 6, explainability, human review, and governance, are comparatively less capital-intensive to build but more labor-intensive to operate on a continuing basis, since they depend on sustained analyst and validation staffing rather than one-time engineering work. A practical implication is that the return profile shifts over time. Early phases dominated by Layer 1 through 3 investment are likely to show returns primarily through improved detection of patterns that record-level rules miss. Later phases, once Layers 4 through 6 are mature, are more likely to show returns through reduced unproductive alert volume and improved analyst throughput, because those gains depend on explanation quality and feedback-loop maturity that only develop once the governance layer has been operating long enough to accumulate reliable disposition data. Institutions evaluating the framework on a short horizon should therefore expect WJIMT its benefits to be back-loaded relative to its costs, which is a material consideration for how a pilot is scoped and how long it must run before its results are informative. 6.6 Vendor and Build-versus-Buy Considerations Institutions adopting the framework face a further practical choice cutting across all six layers: whether to build each component internally, license it from a specialist vendor, or combine the two. Ingestion, feature engineering, and detection, Layers 1 through 3, are the layers most commonly available as mature third-party platforms, since the underlying techniques, streaming pipelines, entity resolution, and ensemble anomaly scoring, are broadly similar across institutions. Explainability and governance, Layers 4 and 6, are harder to purchase off the shelf in a form that satisfies a specific institution's validation practice, because conceptual-soundness review is inherently tied to how that institution's model risk function is organized (Federal Reserve Board of Governors, 2011). Layer 5 sits in between: workflow and routing logic is often vendor-supplied, but the risk-tier thresholds embedded in that workflow, and the analyst training needed to use it well, are institution-specific. A pragmatic implication is that institutions should expect to build or closely customize Layers 4 through 6 even where Layers 1 through 3 are substantially purchased, since those governance-facing layers are precisely where the integration gap identified in Section 2 leaves generic tooling least adequate. 7. Limitations and Future Work As a conceptual contribution, this framework has not been empirically validated against live transaction data. Its components are each grounded in prior literature describing their use individually, but the specific combination and sequencing proposed here has not been tested, and no performance claim should be attributed to the framework as a whole on the basis of results reported for its parts. The illustrative scenarios presented in Section 4 are hypothetical and constructed for expository purposes; they demonstrate how the layers are intended to interact, not that the framework performs as described when implemented. A related caution concerns how any future evaluation should be conducted. Performance estimates in fraud and financial-crime detection are highly sensitive to how verification latency and class imbalance are modeled, and figures obtained without accounting for those constraints systematically overstate operational performance (Dal Pozzolo et al., 2018). Any empirical test of this framework would need to adopt an evaluation protocol that reflects the delayed, partial, and non-random nature of analyst confirmation, rather than reporting accuracy against a fully labeled retrospective dataset. The framework also does not resolve, and does not claim to resolve, the underlying tension between model complexity and interpretability. Deep and graph-based components generally offer stronger detection performance than simpler models (Pang et al., 2021), but post-hoc explanation methods for these architectures approximate rather than reproduce the model's internal reasoning, and the question of how faithfully an attribution or subgraph explanation reflects the actual computation remains methodologically open (Ribeiro et al., 2016; Ying et al., 2019). An institution relying on such explanations to satisfy a regulatory obligation is relying on an approximation whose fidelity it should independently assess. Calibrating the risk tiers used for human-in-the-loop routing in Layer 5 requires institution-specific data on the relative cost of different error types and on available reviewer capacity, information a conceptual framework cannot supply; each implementing institution would need to determine its own thresholds. Cross-jurisdictional regulatory divergence is a further limitation: obligations under the Bank Secrecy Act framework, data-protection regimes, AI-specific regulation, and WJIMT supervisory model risk guidance are not fully harmonized, and institutions operating across borders may require jurisdiction-specific variants of the governance layer, a challenge made more concrete by the staged implementation of the travel rule across jurisdictions (FATF, 2021a) and by the classification questions discussed in Section 6.2. The comparative analysis in Section 5 is qualitative rather than quantitative, drawn from characteristics reported across separate bodies of literature rather than from a controlled evaluation of the approaches on a common dataset; Table 3 should be read as a positioning device rather than as evidence of measured superiority. Similarly, the federated configuration described for Layers 2 and 3 has been validated in the cited literature primarily on benchmark rather than production data, and its behavior under a live, regulated, multi- institution deployment, including its communication overhead and its sensitivity to heterogeneous data distributions across participants, remains untested in this setting (Kairouz et al., 2021). Finally, the framework's design assumptions suit institutions large enough to justify the specialized development, validation, and case management functions described across the six layers. The paper does not address how a smaller institution operating at the lower end of the volume range defined in Section 1 should scale the framework down. Such institutions might implement the same six layers with lighter-weight components, or combine some layers organizationally while keeping them conceptually distinct, but neither possibility has been examined here. Future research could pursue at least five directions: institution-level piloting of the framework against live or historical transaction data, using an evaluation protocol that models verification latency explicitly; comparative evaluation of the hybrid rule, statistical, and graph ensemble described in Layer 3 against single-method baselines under matched conditions; study of the feedback loop itself, including how frequently recalibration should occur, how that cadence affects drift and stability, and how it can be governed so as to remain auditable; empirical testing of the federated variant of Layers 2 and 3 in a genuine multi-institution consortium setting; and development of drift-monitoring practice calibrated specifically to adversarial financial-crime settings, where the general methodological literature on concept drift (Gama et al., 2014) has not been tailored to the particular tempo and regulatory obligations of compliance monitoring. 8. Conclusion High-volume transaction environments have outgrown the practical capacity of static, rule-based compliance monitoring, and the combination of sustained alert volumes, constrained review capacity, and persistently low interception rates has pushed institutions toward artificial intelligence as a complement to existing controls. This paper has developed a conceptual framework organizing that response into six interdependent layers, spanning data ingestion, feature and entity-graph construction, hybrid detection, explainability, calibrated human review, and governance with continuous feedback, built around the principle that detection performance and regulatory defensibility should be treated as a single design problem rather than as separate technical and compliance concerns. By mapping these layers onto existing model risk guidance and human-oversight requirements, extending them to digital-asset, insurance, and cross- institutional contexts, illustrating their operation through a hypothetical payments scenario, and positioning them against alternative approaches, the framework offers compliance functions a way to think about AI adoption architecturally rather than as a series of point solutions. Its central limitation, that it has not been tested empirically, is also its central invitation: the framework is offered as a structure for institutional piloting and academic evaluation, not as a finished or validated system. WJIMT References Adebayo, A., Adegbite, M. P., & Ahmed, M. O. (2023). AI augmented threat detection in industrial control systems: A systematic review of machine learning approaches for ICS anomaly detection. International Journal of Engineering and Modern Technology, 9(3), 287-340. https://doi.org/10.56201/ijcsmt.v9.no3.2023.pg287.340 Adegbite, M. P., Adebayo, A., & Ahmed, M. O. (2023). ICS and SCADA threat detection architectures in energy sector networks: A systematic review of SIEM, NDR, and anomaly detection approaches. , 7(2), 121- 182. https://doi.org/10.56201/wjimt.v7.no2.2023.pg121.182 Adelanwa, A., Basnet, A., & Anene, U. N. (2023b). Predictive analytics models for financial risk detection and fraud prevention in public systems. International Journal of Advanced Multidisciplinary Research and Studies, 3(6). https://doi.org/10.62225/2583049X.2023.3.6.5969 Adesuyi, M. O., Akomolafe, O., Olaogun, B. O., Ndukwe, V. U., & Sakyi, J. K. (2024). AI-driven risk scoring model for global cross-border trade payment transactions. International Journal of Advanced Multidisciplinary Research and Studies, 4(1), 1569-1581. https://doi.org/10.62225/2583049X.2024.4.1.5281 Atakpa, M. I. (2021). A review of predictive analytics models for anti-money laundering detection in commercial banking systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 7(2), 778-804. https://doi.org/10.32628/CSEIT2064813 Atakpa, M. I., & Abetoh, N. F. (2022). A systematic review of machine learning advances in financial fraud detection for banking systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 8(2), 771-801. https://doi.org/10.32628/CSEIT23906220 Atakpa, M. I., Abetoh, N. F., & Akeju, B. (2024). Advances in artificial intelligence for healthcare payment fraud detection: A review of NHS applications. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 10(4), 1161- 1194. https://doi.org/10.32628/CSEIT26123240 Atakpa, M. I., & Fobellah, A. N. (2023). Anomaly detection in financial time-series data: A conceptual model for healthcare and banking applications. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 9(2), 967-997. https://doi.org/10.32628/CSEIT2342441 Bank Secrecy Act, 31 U.S.C. §§ 5311-5336 (and implementing regulations at 31 C.F.R. Chapter X). Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32. Breunig, M. M., Kriegel, H.-P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, 93-104. Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785-794. WJIMT Consumer Financial Protection Bureau [CFPB]. (2022). Consumer Financial Protection Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms. Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit card fraud detection: A realistic modeling and a novel learning strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. Debener, J., Heinke, V., & Kriebel, J. (2023). Detecting insurance fraud using supervised and unsupervised machine learning. Journal of Risk and Insurance, 90(3), 743-768. Eboh, E. E., & Aliliele, C. (2024). AI-driven data analytics framework for risk assessment and detection of venture capital and private equity investment fraud in U.S. capital markets. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 10(6), 2624-2665. https://doi.org/10.32628/CSEIT2410787 European Parliament & Council of the European Union. (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22. Official Journal of the European Union, L 119. European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Article 14 and Annex III. Official Journal of the European Union. Federal Reserve Board of Governors & Office of the Comptroller of the Currency. (2011). SR 11- 7: Guidance on model risk management. Financial Action Task Force [FATF]. (2021a). Updated guidance for a risk-based approach to virtual assets and virtual asset service providers. FATF. Financial Action Task Force [FATF]. (2021b). Opportunities and challenges of new technologies for AML/CFT. FATF. Financial Crimes Enforcement Network [FinCEN], Board of Governors of the Federal Reserve System, Federal Deposit Insurance Corporation, National Credit Union Administration, & Office of the Comptroller of the Currency. (2018). Joint statement on innovative efforts to combat money laundering and terrorist financing. Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), Article 44. Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75-105. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., et al. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1-2), 1-210. Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). Isolation forest. Proceedings of the 8th IEEE International Conference on Data Mining, 413-422. Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765-4774. Mbonu, I. S., Aliliele, C., Iwuanyanwu, U., & Uzoka, E. (2022). A conceptual framework for AI enabled IT general controls and SOX audit automation processes. Gyanshauryam, International Scientific Refereed Research Journal, 5(5), 384-414. https://doi.org/10.32628/GISRRJ2256239 Mbonu, I. S., Iwuanyanwu, U., Aliliele, C., & Uzoka, E. (2021). A review of VoIP forensic analytics models for financial fraud detection and regulatory compliance monitoring. WJIMT International Journal of Multidisciplinary Research and Growth Evaluation, 2(6), 711- 730. https://doi.org/10.54660/.IJMRGE.2021.2.6.711-730 McMahan, B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. (2017). Communication- efficient learning of deep networks from decentralized data. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 1273-1282. Pang, G., Shen, C., Cao, L., & van den Hengel, A. (2021). Deep learning for anomaly detection: A review. ACM Computing Surveys, 54(2), Article 38. Pourhabibi, T., Ong, K.-L., Kam, B. H., & Boo, Y. L. (2020). Fraud detection: A systematic literature review of graph-based anomaly detection approaches. Decision Support Systems, 133, 113303. Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). 'Why should I trust you?': Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135-1144. United Nations Office on Drugs and Crime [UNODC]. (2011). Estimating illicit financial flows resulting from drug trafficking and other transnational organized crimes. UNODC. Weber, M., Domeniconi, G., Chen, J., Weidele, D. K. I., Bellei, C., Robinson, T., & Leiserson, C. E. (2019). Anti-money laundering in bitcoin: Experimenting with graph convolutional networks for financial forensics. arXiv:1908.02591. Ying, R., Bourgeois, D., You, J., Zitnik, M., & Leskovec, J. (2019). GNNExplainer: Generating explanations for graph neural networks. Advances in Neural Information Processing Systems, 32, 9244-9255.