OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77
OpenRouter published a head-to-head classification benchmark comparing its Jev 1.13 model with Claude Opus 5 on the Banking77 intent dataset. Researchers routed 3,080 utterances through both systems via the Decisions API under identical intent criteria. Opus reached 84.4% accuracy while Jev landed at 81.0%, a modest but meaningful gap on this banking-domain taxonomy. The trade-off profile differs sharply on economics and speed: Jev posted median latency near 175 milliseconds and roughly eleven cents per thousand requests, versus about 2.3 seconds and $2.42 per thousand for Opus. The write-up positions Opus ahead on spend-weighted classification rankings while framing Jev as a cost-efficient near-frontier option. Readers should treat scores as specific to Banking77 and OpenRouter's test harness rather than universal capability.
OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77
OpenRouter sent 3,080 Banking77 utterances through Jev 1.13 and Claude Opus 5 using the same intent criteria for both models. Jev scored 81.0% against 84.4% for Opus, with median latency 175 ms and about $0.11 per thousand requests compared with roughly 2.3 seconds and $2.42 for Opus.
Key takeaway
Jev 1.13 trails Claude Opus 5 by 3.4 points on Banking77 but costs far less per request and responds much faster in OpenRouter's Decisions API test.
What happened
OpenRouter reports it sent the same 3,080 Banking77 utterances to Jev 1.13 and Claude Opus 5 through the Decisions API, applying the same intent criteria to both models in a controlled classification comparison.
On that run, Jev 1.13 scored 81.0% accuracy against 84.4% for Opus. Jev showed median latency of 175 ms and about $0.11 per thousand requests, compared with roughly 2.3 seconds and $2.42 per thousand for Opus.
Evidence
OpenRouter compared Jev 1.13 and Claude Opus 5 on 3,080 Banking77 utterances via the Decisions API with shared intent criteria.
OpenRouter · attributed
We sent the same 3,080 Banking77 utterances to it and to Jev 1.13 through the Decisions API.
Claude Opus 5 outscored Jev 1.13 on Banking77 classification accuracy in OpenRouter's benchmark.
OpenRouter · attributed
Opus scored 84.4% to Jev's 81.0%
Jev 1.13 showed lower median latency and lower per-thousand request cost than Opus in the same test.
OpenRouter · attributed
Jev answered in 175 ms at $0.11 per thousand requests against 2.3 seconds and $2.42 for Opus.
OpenRouter states Claude Opus 5 leads its classification task ranking by spend.
OpenRouter · attributed
Claude Opus 5 leads OpenRouter's classification task ranking by spend.
Why it matters
High-volume intent routing on OpenRouter can trade a few accuracy points for large latency and cost savings when workloads resemble Banking77-style classification.
Limits and uncertainties
Results apply only to Banking77 and OpenRouter's Decisions API setup; the packet does not report broader task coverage or independent replication.
The comparison covers Jev 1.13 and Claude Opus 5 only; other models and deployment conditions are outside this evidence.
Practical implications
Operators choosing models for banking-intent classification should weigh the 81.0% versus 84.4% gap against 175 ms versus 2.3 s median latency and $0.11 versus $2.42 per thousand requests.
Builders using OpenRouter's classification ranking by spend should note Opus leads that ranking while Jev is positioned as a cheaper, faster alternative on this benchmark.