Skip to main content
LLMgram · AI News · 2026-09-30

Comet ML: Jev beat gpt-4o-mini judge on cost and speed over 1,000 production turns

Comet ML: Jev beat gpt-4o-mini judge on cost and speed over 1,000 production turns

Comet ML published benchmark results comparing Jev, a yes-or-no evaluation model, against using gpt-4o-mini as an LLM judge on one thousand real production turns. The company reports Jev delivered roughly 3.5 times lower cost and 3.8 times faster inference than the mini judge on that workload. Agreement between the two approaches was high but not total: the yes-no model matched the mini judge on 897 turns, leaving 103 disagreements that Comet ML says it analyzes in the full post. The comparison matters for teams running continuous evals where judge spend and latency scale with traffic. Readers should treat the numbers as a single vendor-run slice on Comet's own production traffic rather than a universal benchmark across models or domains.

Sources

Comet ML: Jev beat gpt-4o-mini judge on cost and speed over 1,000 production turns

Comet ML: Jev beat gpt-4o-mini judge on cost and speed over 1,000 production turns

Comet ML reports Jev came out 3.5× cheaper and 3.8× faster than its gpt-4o-mini judge on 1,000 real production turns. On the same run, the yes/no judge agreed with the mini judge on 897 of those turns.

Key takeaway

A dedicated yes/no judge can cut eval cost and latency versus gpt-4o-mini while matching on most production turns.

What happened

Comet ML reports that on one thousand real production turns, Jev was 3.5 times cheaper and 3.8 times faster than its gpt-4o-mini judge, framing Jev as a model that answers yes/no questions instead of generating free-form judge text.

On the same run, Comet ML says the yes/no judge agreed with the gpt-4o-mini judge on 897 of those turns, implying 103 disagreements; the published post describes the test setup, patterns in those disagreements, and steps to reproduce the comparison.

Evidence

  • Jev was 3.5 times cheaper and 3.8 times faster than gpt-4o-mini on 1,000 production turns.

    Comet ML · attributed

    Comet ML reports Jev came out 3.5× cheaper and 3.8× faster than its gpt-4o-mini judge on 1,000 real production turns.

  • The yes/no judge agreed with the mini judge on 897 of 1,000 turns.

    Comet ML · attributed

    On the same run, the yes/no judge agreed with the mini judge on 897 of those turns.

  • Jev answers yes/no questions rather than writing judge prose.

    Comet ML · attributed

    A model that answers yes/no questions instead of writing text came out 3.5× cheaper and 3.8× faster than our gpt-4o-mini judge on 1,000 real production turns, and agreed with it on 897 of them.

Why it matters

If similar cost and speed gaps hold on your traffic, LLM-as-judge budgets and p95 eval latency may shrink, but the 103 mismatches are where quality risk lives.

Limits and uncertainties

Results come from Comet ML's own reporting on one thousand turns of its production traffic, not an independent or multi-vendor benchmark.

The packet does not define rubrics, domains, or error rates beyond the 897-of-1,000 agreement count.

Practical implications

Teams running high-volume judges should compare specialized yes/no eval models against their current mini-model judge on real turns before switching.

Operators should read Comet ML's disagreement analysis and reproduce the test on their own workloads rather than assuming identical speedup or alignment.

What to watch

Whether Comet ML or third parties publish replication on non-Comet production evals and how the 103 disagreements cluster by criterion or task type.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Jev vs. LLM-as-a-Judge for AI Evals