Coinbase says newer versions of three major artificial-intelligence model families performed worse at detecting payment fraud in a fixed historical test for its Onramp service. The result challenges a common assumption in financial technology: that upgrading to a newer general-purpose model automatically improves a specialized risk system.
A controlled replay of known outcomes
The company evaluated 16,140 transactions from 7,293 users across nine weeks of production data collected before its AI risk agent was deployed. The sample retained 813 confirmed fraudulent transactions and included sampled legitimate activity. Each candidate model received the same recent transaction context and operated under the same guidance and risk-to-decision policy, allowing Coinbase to isolate how the model itself changed within that setup.
Related reporting: HSBC and Ant Digital Test AI-Agent Payments With Tokenised Deposits
Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Across all three pairs, the newer version produced lower recall, F1 and dollar-weighted recall. Recall measures the share of known fraudulent transactions caught, while dollar-weighted recall measures the share of total fraudulent value detected. F1 combines precision and recall into one score.
Better precision did not mean better coverage
The size and pattern of the regression varied. Sonnet's recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus recall declined 0.8 points. Both newer versions also recorded lower precision, meaning a smaller proportion of their fraud flags were correct.
GPT showed a different trade-off. Its precision improved by 11.5 percentage points, but recall fell 20.7 points and dollar-weighted recall fell 21.8 points. In practical terms, its alerts were more likely to identify real fraud, but the system allowed a larger share of fraudulent cases and value to pass unflagged in the replay. That is why payment teams cannot judge an upgrade from one favorable metric alone.
What the benchmark does not prove
The evaluation was retrospective, not a live comparison of customer outcomes. Coinbase said the replay revealed regressions but did not identify their cause, and the results do not show that customers suffered losses because those newer models were deployed. The proprietary transaction dataset also cannot be released, limiting independent replication and making it risky to generalize the findings to other payment systems.
Coinbase separately reported that a post-trained open-weight Qwen3.5-9B model outperformed Opus 4.5 on its fraud benchmark after domain-specific training. That was a different evaluation and should not be read as proof that a smaller model will outperform every frontier system. It does reinforce the central lesson: model selection for crypto payments needs reproducible tests against verified fraud outcomes, alongside latency, reliability and operating cost.
Why it matters for crypto payments
Onramp services connect conventional payment methods with digital assets, so missed fraud can create losses, chargebacks and compliance risk. False positives matter too because they can block legitimate users. For exchanges, wallets and payment providers, a model upgrade is therefore a policy change as well as a software change. The safest deployment process is to test each candidate on the institution's own risk objectives, preserve rollback options and monitor several detection metrics after release. Teams also need mature labels because a benchmark is only as reliable as the fraud outcomes used to score it, especially when disputes and chargebacks can surface weeks after a transaction.
Sources
- Coinbase: How Coinbase Built, Evaluated, and Post-Trained a Fraud Agent, Part 2
- CryptoSlate: Newer AI models missed more payment fraud in Coinbase's benchmark
- SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection
AI-generated editorial image; not a photograph of the reported event. Prepared with AI assistance and source verification.
