Newer AI models missed more payment fraud in Coinbase’s benchmark

Newer AI models missed more payment fraud in Coinbase’s benchmark

Summary

Coinbase replayed 16,140 Onramp transactions, including 813 confirmed fraud cases, to compare newer AI models with earlier versions under the same screening policy. All three newer models caught a smaller share of fraud cases and fraud value, and had lower F1 scores. GPT’s precision improved even as its recall and dollar-weighted recall fell. The test did not establish that customers suffered losses or explain why performance regressed. In a separate evaluation, Coinbase said a fraud-specialized Qwen model outperformed Opus 4.5 on four metrics and had lower median request latency. Coinbase advises payment providers to test model upgrades in their own screening setup.