Reducing False-Positive Alert Burden in Transaction Monitoring: Calibrated Risk Scoring for Resource-Constrained Community Financial Institutions
Keywords:
False-Positive Alert, Transaction Monitoring, Risk Scoring, Financial InstitutionsAbstract
Rule-based transaction monitoring systems remain the default anti-money-laundering (AML) control at most deposit-taking institutions, yet they are widely reported to generate false-positive alert rates in excess of 95 percent. For large banks this inefficiency is absorbed by sizeable compliance teams; for community banks, credit unions, and other resource-constrained financial institutions it translates directly into alert backlogs, burnout among a small compliance staff, and a persistent gap between the volume of alerts generated and the capacity available to review them. This article reframes findings from a broader programme of transaction-monitoring research around a single, practical question: which of the machine-learning approaches currently proposed for AML detection can meaningfully lower the false-positive burden without requiring the computational budget, data-science headcount, or infrastructure that only the largest institutions can afford. Using a systematic review of the transaction-monitoring literature, structured interviews with anti-money-laundering specialists, a purpose-built synthetic dataset (SAML-D), and an evaluation on an anonymised real-world case dataset from a partner bank, the underlying research compared gradient-boosted trees, random forests, and three transformer-based deep learning architectures (Tab-AML, TabTransformer, and TabNet) on their ability to separate suspicious from non-suspicious activity. On the real-world, case-aggregated dataset, XGBoost, a comparatively lightweight tree-ensemble method, achieved the strongest discrimination (ROC-AUC of 88.73 percent) and, when calibrated to a 98 percent true-positive threshold, produced a false-positive rate 6.94 percentage points lower than the best-performing deep learning alternative, TabNet. A follow-up feature-importance analysis using SHAP values showed that a compact subset of roughly 40 features, drawn overwhelmingly from transaction-behaviour and risk-indicator categories, matched or exceeded the performance of the full 90-feature model while substantially reducing computational overhead. These results carry a direct implication for resource-constrained institutions: the model family that is easiest to deploy, tune, and explain on modest hardware is also the strongest performer once transactions have been aggregated into investigable cases, and further gains in efficiency are available by calibrating decision thresholds and pruning the feature set rather than by adopting deep learning infrastructure. The article closes with a practical framework for calibrated risk scoring that community institutions can use to reduce alert volumes while preserving detection sensitivity, along with a discussion of the governance, data-quality, and explainability considerations that must accompany any move away from static, rule-based thresholds.