Skip to main content
LLMgram · AI News · 2026-09-22

OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77

OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77

OpenRouter published a head-to-head classification benchmark comparing its Jev 1.13 model with Claude Opus 5 on the Banking77 intent dataset. Researchers routed 3,080 utterances through both systems via the Decisions API under identical intent criteria. Opus reached 84.4% accuracy while Jev landed at 81.0%, a modest but meaningful gap on this banking-domain taxonomy. The trade-off profile differs sharply on economics and speed: Jev posted median latency near 175 milliseconds and roughly eleven cents per thousand requests, versus about 2.3 seconds and $2.42 per thousand for Opus. The write-up positions Opus ahead on spend-weighted classification rankings while framing Jev as a cost-efficient near-frontier option. Readers should treat scores as specific to Banking77 and OpenRouter's test harness rather than universal capability.

Sources

OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77

OpenRouter benchmarks Jev 1.13 at 81.0% vs Claude Opus 5 at 84.4% on Banking77

OpenRouter sent 3,080 Banking77 utterances through Jev 1.13 and Claude Opus 5 using the same intent criteria for both models. Jev scored 81.0% against 84.4% for Opus, with median latency 175 ms and about $0.11 per thousand requests compared with roughly 2.3 seconds and $2.42 for Opus.

Key takeaway

Jev 1.13 trails Claude Opus 5 by 3.4 points on Banking77 but costs far less per request and responds much faster in OpenRouter's Decisions API test.

What happened

OpenRouter reports it sent the same 3,080 Banking77 utterances to Jev 1.13 and Claude Opus 5 through the Decisions API, applying the same intent criteria to both models in a controlled classification comparison.

On that run, Jev 1.13 scored 81.0% accuracy against 84.4% for Opus. Jev showed median latency of 175 ms and about $0.11 per thousand requests, compared with roughly 2.3 seconds and $2.42 per thousand for Opus.

Evidence

  • OpenRouter compared Jev 1.13 and Claude Opus 5 on 3,080 Banking77 utterances via the Decisions API with shared intent criteria.

    OpenRouter · attributed

    We sent the same 3,080 Banking77 utterances to it and to Jev 1.13 through the Decisions API.

  • Claude Opus 5 outscored Jev 1.13 on Banking77 classification accuracy in OpenRouter's benchmark.

    OpenRouter · attributed

    Opus scored 84.4% to Jev's 81.0%

  • Jev 1.13 showed lower median latency and lower per-thousand request cost than Opus in the same test.

    OpenRouter · attributed

    Jev answered in 175 ms at $0.11 per thousand requests against 2.3 seconds and $2.42 for Opus.

  • OpenRouter states Claude Opus 5 leads its classification task ranking by spend.

    OpenRouter · attributed

    Claude Opus 5 leads OpenRouter's classification task ranking by spend.

Why it matters

High-volume intent routing on OpenRouter can trade a few accuracy points for large latency and cost savings when workloads resemble Banking77-style classification.

Limits and uncertainties

Results apply only to Banking77 and OpenRouter's Decisions API setup; the packet does not report broader task coverage or independent replication.

The comparison covers Jev 1.13 and Claude Opus 5 only; other models and deployment conditions are outside this evidence.

Practical implications

Operators choosing models for banking-intent classification should weigh the 81.0% versus 84.4% gap against 175 ms versus 2.3 s median latency and $0.11 versus $2.42 per thousand requests.

Builders using OpenRouter's classification ranking by spend should note Opus leads that ranking while Jev is positioned as a cheaper, faster alternative on this benchmark.

What to watch

Whether OpenRouter publishes additional classification benchmarks beyond Banking77 and updated Jev versions.

How spend-weighted classification rankings shift if accuracy thresholds or traffic mix change.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Is Jev as Accurate as Frontier Models at Classification?