An evaluator in Rubrist asks one narrow question about a trace: did the reply stay within the refund policy, did the agent confirm before changing the order. We started with an LLM judge answering it. Depending on the model and trace, that takes seconds and can cost cents per trace, which makes scoring every trace expensive. Jev, TypeSafe’s System One model, answers a yes-or-no question with a probability in about a quarter of a second, for a small fraction of a cent. If it judged as well, every trace could be scored instead of a sample. We tested 459 cases, then added follow-ups and 1,000 fresh cases to see where the result holds.

Can a cheap probability model do a judge’s job?

The question is narrower than “is Jev as good as Claude”. A judge in an evaluator has to agree with people on one criterion, reliably, on the traces that criterion covers. So we measured agreement with labels on three kinds of criterion: a short language judgment, a comparison between two answers, and a verdict on a whole agent run. We also measured how the judges ranked cases, how much they cost and how long they took. Five judges saw the first 459 cases. Only Jev and Opus saw the 1,000 fresh cases. Comparisons are paired on cases both judges answered.

What we ran

Three public datasets supplied the cases and the labels:

  • ChaosMNLI (tasksource/chaos-mnli-ambiguity, from ChaosNLI by Nie et al., 2020, CC BY-NC 4.0): a hundred annotators judged each premise and hypothesis. The criterion was whether the hypothesis is definitely true given the premise, and it passed when at least half the annotators chose entailment. 200 cases drawn at random.
  • MT-Bench human judgments (lmsys/mt_bench_human_judgments, CC BY 4.0): experts compared two chatbot answers. The criterion was whether answer A is better than answer B at the judged turn. 150 pairs with a clear majority, drawn at random and run in both orders. 103 of them rest on a single expert’s vote.
  • tau-bench retail (sierra-research/tau-bench, MIT): recorded customer-service runs by a GPT-4o agent, with tool calls and results. The criterion was whether the agent resolved the request correctly under the retail policy. The label is the benchmark’s own database check, not a person. 109 runs, one per task.

Jev (jev-1.13.0, pinned) got one yes-or-no question per case, built from the criterion, and returned a probability. The four Claude models got the request Rubrist sends when it calibrates an evaluator: the rubric in its default prompt, a structured verdict of pass, fail or unsure with a score and a reason, with one planned call per case. Three verdicts came back without a reason and were sent again once. The models were Haiku 4.5; Sonnet 4.6, Rubrist’s default judge at the time; Sonnet 5; and Opus 5.5. Two of them needed changes, covered below. In the initial comparison, the 459 distinct cases became 609 presentations because MT-Bench ran in both orders. Across five judges that was 3,045 requests, plus the three missing-reason retries, for about $29 in API spend, or $34.90 including an earlier run. The follow-ups below added about $56; those costs are separate from the initial comparison. The code and fixtures are in the Rubrist repository, and the full per-case report is next to them.

On short judgments, the first sample showed no difference

Figure 1 shows each judge’s accuracy with its 95 percent interval, next to the accuracy of always giving the dataset’s most common answer. On the two short datasets the five judges sit on top of each other. The MT-Bench panel shows the original answer order; the section on order covers the swapped run.

Figure 01

On short judgments the five judges overlap. On agent traces only Opus 5.5 clears the baseline.

ChaosMNLI, 200 entailment judgmentsJev 1.13: 77.5% accuracy, interval 71.2% to 82.7%; Haiku 4.5: 76.4% accuracy, interval 69.9% to 81.9%; Sonnet 4.6: 75% accuracy, interval 68.6% to 80.5%; Sonnet 5: 78% accuracy, interval 71.8% to 83.2%; Opus 5.5: 78.5% accuracy, interval 72.3% to 83.6%. Majority-answer baseline (always fail): 59%.ChaosMNLI200 entailment judgmentsJev 1.13Haiku 4.5Sonnet 4.6Sonnet 5Opus 5.5MT-Bench, 150 pairwise preferences, original orderJev 1.13: 86% accuracy, interval 79.5% to 90.7%; Haiku 4.5: 79.3% accuracy, interval 72.2% to 85%; Sonnet 4.6: 83.9% accuracy, interval 77.2% to 88.9%; Sonnet 5: 82.3% accuracy, interval 75.3% to 87.6%; Opus 5.5: 83.3% accuracy, interval 76.6% to 88.4%. Majority-answer baseline (always A): 55.3%.MT-Bench150 pairwise preferences, original orderJev 1.13Haiku 4.5Sonnet 4.6Sonnet 5Opus 5.5tau-bench, 109 agent trajectoriesJev 1.13: 50.5% accuracy, interval 41.2% to 59.7%; Haiku 4.5: 41.3% accuracy, interval 32.5% to 50.7%; Sonnet 4.6: 52.3% accuracy, interval 43% to 61.4%; Sonnet 5: 58.7% accuracy, interval 49.3% to 67.5%; Opus 5.5: 74.3% accuracy, interval 65.4% to 81.6%. Majority-answer baseline (always pass): 57.8%.tau-bench109 agent trajectoriesJev 1.13Haiku 4.5Sonnet 4.6Sonnet 5Opus 5.5

Accuracy from 0.3 to 1.0Jev 1.13Claude judgedashed line: majority-answer baseline

One recorded run on 24 September 2026, not a benchmark. Each dot is a judge’s accuracy on the cases it decided; the bar is its Wilson 95% interval. The dashed line is the accuracy of always giving the dataset’s majority answer. Jev is drawn as a filled square and every Claude judge as an open circle; the tables in the text list every value.
Accuracy and ranking quality on the short datasets, one run
JudgeChaosMNLI accuracyChaosMNLI AUCMT-Bench accuracyMT-Bench AUC
Jev 1.130.7750.9440.8600.938
Haiku 4.50.7640.3380.7930.773
Sonnet 4.60.7500.7000.8390.907
Sonnet 50.7800.8340.8230.887
Opus 5.50.7850.9480.8330.932

A paired McNemar test, which compares two judges on the cases where exactly one of them was right, found no accuracy difference between Jev and Sonnet 4.6, Sonnet 5 or Opus 5.5 on either dataset. The lowest p-value was 0.065. Against Opus the rankings agree too. AUC is the chance that a judge scores a randomly chosen passing case above a failing one. The two AUCs differed by 0.004 on ChaosMNLI and 0.006 on MT-Bench, with intervals that straddle zero. On ChaosMNLI, Jev’s probability also tracked how many of the hundred annotators chose entailment, with a rank correlation of 0.85. Opus’s score tracked it about as closely, at 0.84.

No difference showing up is not the same as no difference. With 150 to 200 cases, only a gap of roughly five to seven points of accuracy would reliably show. On 500 more MT-Bench pairs, a smaller gap did show. On ChaosMNLI, Haiku’s and Sonnet 4.6’s low AUCs came mostly from our own prompt, which the section on our judge’s problems explains. After the fix they rose to 0.90 and 0.92. Haiku also declined to decide 9 ChaosMNLI cases. Counting those as wrong, its accuracy there is 0.730. All five judges were also stricter than the annotators on ChaosMNLI: each passed only about half of the cases the majority had called entailed. They shared that strictness even on cases where most annotators agreed. So it reflects the criterion’s wording, “definitely true”, at least as much as any one model.

What each decision cost

Cost per 1,000 cases and median latency, one run at concurrency 4
JudgeChaosMNLIMT-Benchtau-benchMedian latency
Jev 1.13$0.017$0.044$0.200.25–0.27 s
Haiku 4.5$2.12$3.16$7.332.2–4.3 s
Sonnet 4.6$6.09$9.24$22.723.5–8.2 s
Sonnet 5$5.04$8.85$19.133.0–5.7 s
Opus 5.5$13.47$20.25$52.804.0–8.6 s

Costs use list prices on 24 September 2026. Jev charges $0.042 per million input tokens and nothing for output. The Claude models charge per million input and output tokens: Haiku 4.5 $1 and $5, Sonnet 4.6 $3 and $15, Sonnet 5 $2 and $10, Opus 5.5 $4 and $20. On the short datasets Jev cost about 0.1 to 0.2 percent of what Opus did, and about 1 percent of Haiku. It answered in under a tenth of Opus’s time. At these prices, scoring every trace with Jev costs less than scoring 1 to 3 percent of them with Haiku, or 0.1 to 0.4 percent with Opus.

One judge preferred whichever answer came first

A pairwise judge should pick the same answer whichever one it is shown first. We ran every MT-Bench pair twice, with the answers swapped, and counted how often each judge picked the same underlying answer.

Same pick in both orders, MT-Bench
JudgeConsistentAccuracy, original orderAccuracy, swapped
Opus 5.5148 of 1500.8330.847
Jev 1.13143 of 1500.8600.867
Sonnet 4.6130 of 1480.8390.826
Sonnet 5128 of 1470.8230.847
Haiku 4.581 of 1500.7930.667

Haiku kept its pick only 54 percent of the time. In 58 of its 69 inconsistent pairs it picked whichever answer came first. Answer A was the more common winner in the original order, so that habit raised Haiku’s original-order accuracy and lowered its swapped one, from 0.79 to 0.67. Against Jev on the same swapped pairs, Haiku was clearly worse (p < 0.001). Opus and Jev barely moved.

On agent traces, only Opus held up

tau-bench is the dataset closest to how evaluators are used on agents: a full run with user turns, tool calls and results, averaging about five thousand tokens. Here the picture changes.

tau-bench retail, 109 runs; always answering pass scores 0.578
JudgeAccuracy95% intervalAUC
Opus 5.50.7430.654–0.8160.824
Sonnet 50.5870.493–0.6750.614
Sonnet 4.60.5230.430–0.6140.519
Jev 1.130.5050.412–0.5970.547
Haiku 4.50.4130.325–0.5070.469

Opus is the only judge clearly better than always answering pass. On the same runs it beat Jev by 0.24 in accuracy, with an interval of 0.14 to 0.34 and a McNemar p-value below 0.001. Jev, Haiku and Sonnet 4.6 were at chance, and Sonnet 5 was only slightly above.

We don’t know why. Opus 5.5 reasons before it answers by default, and in our checks its responses carried a thinking block. The other judges answered straight away. Haiku 4.5 and Sonnet 4.6 don’t reason unless asked, and a spot check of Sonnet 5 under the same forced verdict call came back with no thinking. Opus is also the most capable model of the five. This run didn’t isolate which of those matters. A follow-up suggests the model matters more than the reasoning. None of these runs can separate the length of these traces from the kind of task. The criterion asks whether the agent reached a goal that the benchmark checks against its database. A judge has to reconstruct that goal from the conversation, and a one-shot probability over five thousand tokens may be the wrong tool for it. Outcomes varied across repeated agent runs: on 46 of the 109 tasks the same agent passed in some trials and failed in others. That does not by itself mean a particular run’s label is wrong. Opus reached 0.74 overall and 0.85 on those 46 tasks.

Two problems the run exposed in our own judge

Both apply to anyone running LLM judges, not only to Rubrist.

The score had no stated direction. Rubrist’s instructions asked for “a confidence-weighted score in [0,1]”, and only the output schema said that 1 means a strong pass. Many judges read it as confidence in their own verdict, so a confident fail came back as 0.9. On ChaosMNLI, Haiku’s score contradicted its own label in 81 of 191 verdicts, Sonnet 4.6’s in 64 and Sonnet 5’s in 38. Opus’s never did. Read as confidence in the chosen label, Haiku’s AUC there rises from 0.34 to 0.81. We now state the direction in the instructions. The Claude runs above predate that fix. We reran them afterwards: the ChaosMNLI contradictions went to zero and its AUCs rose. If your judge returns a number, check which way it points against its verdict.

Newer models reject some usual judge settings. At the time of the first run, Rubrist sent temperature 0 and forced the verdict tool. Sonnet 5, Opus 5.5 and Fable 5.1 rejected that temperature, and Opus also refused the forced tool call, so the application could not use them as configured. For this experiment Sonnet 5 and Opus ran without temperature; Opus was offered the verdict tool without being forced to use it. A later check on 27 September clarified that these models reject temperatures 0 and 0.5 but accept 1. The issue is which settings they accept, not the presence of the parameter alone. Rubrist now checks the configured model’s capabilities and records the settings actually sent, rather than silently dropping unsupported ones. It also now supports TypeSafe’s Jev as a typed-question evaluator.

What we checked next

The first run left several questions open. We checked them the next day, 25 September, on the same public datasets, for about $56 more in API spend. The reports are next to the first run’s. Like the rest of this post, these are diagnostics on public labels, not governed truth.

The Claude judges, rerun after the score fix

We sent the four Claude models the same cases again, this time with the score’s direction stated in the instructions. Jev’s request hadn’t changed, so its answers are the recorded ones. The rerun cost about $29.

Scores that contradicted their own label on ChaosMNLI, and AUC before and after the fix
JudgeChaosMNLI contradictionsChaosMNLI AUCMT-Bench AUCMT-Bench swapped AUCtau-bench AUC
Jev 1.13 (unchanged)00.9440.9380.9400.547
Haiku 4.581 → 00.338 → 0.8960.773 → 0.7930.713 → 0.7940.469 → 0.414
Sonnet 4.664 → 00.700 → 0.9160.907 → 0.9040.901 → 0.8930.519 → 0.561
Sonnet 538 → 00.834 → 0.9290.887 → 0.8920.906 → 0.9140.614 → 0.632
Opus 5.50 → 00.948 → 0.9500.932 → 0.9350.933 → 0.9340.824 → 0.818

The fix did what it should. No ChaosMNLI score contradicted its label any more, and across all four sets, counting MT-Bench once per answer order, only two Sonnet 5 scores still did. The low ChaosMNLI AUCs were mostly our prompt, not the models. Accuracy moved by at most five points, and between 2 and 12 verdicts per judge and set changed. We have no second Claude run with the same request, so we can’t tell how much of that movement is the fix and how much is ordinary run-to-run variation.

The comparison with Jev mostly held.

  • Opus 5.5. Still no accuracy difference detected on ChaosMNLI (p = 0.83) or MT-Bench (p = 0.55), and the AUCs within 0.006 of each other. On the swapped order Jev was ahead by four points, which is at the edge of detectable (p = 0.07).
  • The Sonnets. On ChaosMNLI no accuracy difference from Jev showed for either. On MT-Bench in both orders, Jev ranked cases better than both Sonnets, by 0.03 to 0.05 of AUC. Those intervals exclude zero or just touch it.
  • Haiku. It trailed Jev on AUC everywhere. In the swapped order it also trailed on accuracy, by 15 points (p < 0.001).
  • tau-bench. Unchanged: Opus beat Jev by 0.24 in accuracy (p < 0.001) and by 0.27 in AUC.

Order consistency stayed in the same ranks. Opus kept its pick in 97 percent of pairs, Jev in 95, Sonnet 4.6 in 90, Sonnet 5 in 87, and Haiku in 59. The rerun also turned up a smaller problem: Sonnet 5 left out the required reason in 8 of its 300 MT-Bench verdicts. This time we didn’t resend them, so they count as errors. A verdict tool whose schema the provider enforces, rather than one the model is merely asked to follow, would prevent that.

Less reasoning didn’t cost Opus its lead on agent traces

Opus 5.5 reasons by default, and the other judges didn’t. So we turned the reasoning down for Opus and up for Sonnet 5, on the same 109 tau-bench runs. Opus 5.5 refuses to switch thinking off, so “low effort” is as close to no reasoning as it allows.

tau-bench retail, 109 runs, after the score fix
JudgeAccuracy95% intervalAUCCost per 1,000Median / p95 latency
Opus 5.5, default reasoning0.7430.654–0.8160.818$52.928.8 / 11.5 s
Opus 5.5, low effort0.7250.634–0.8000.783$47.546.2 / 8.3 s
Sonnet 5, no thinking0.5960.502–0.6840.632$19.025.1 / 10.4 s
Sonnet 5, adaptive thinking0.6330.539–0.7180.669$43.4125.7 / 72.1 s

At low effort, Opus lost no accuracy we could detect (p = 0.63). Its AUC fell by 0.03, with an interval of 0.006 to 0.066, and the saving was only a tenth of the cost. With thinking on, Sonnet 5 wasn’t detectably better (p = 0.48). It cost 2.3 times as much, and its median answer took five times as long. It still trailed default Opus by 11 points of accuracy (p = 0.05) and 0.15 of AUC. On these runs the model, more than its reasoning effort, is the likelier explanation for Opus’s lead. With 109 labels that is as far as it goes: for Sonnet 5, thinking’s effect on accuracy has an interval from −4 to +11 points, so a moderate reasoning effect isn’t excluded.

Six short questions instead of one

Jev is built for narrow questions, and “did the agent resolve the request under the policy” is a broad one. So we split it into six yes-or-no questions from the retail policy, written before we saw any results for them. We asked all six in one Jev call, and a run passed only if every answer was yes.

tau-bench retail, 109 runs
JudgeAccuracy95% intervalAUC
Jev, one question0.5050.412–0.5970.547
Jev, six questions0.5600.466–0.6490.617
Opus 5.5 after the score fix, for reference0.7430.654–0.8160.818

The accuracy gain wasn’t significant (p = 0.21). The observed ranking gain was 0.07 of AUC, with an interval of 0.003 to 0.14 that narrowly excludes zero. Most of the signal came from whether the agent’s actions matched what the customer asked for (AUC 0.68), and some from whether the request was resolved (0.61). The questions about procedure, such as verifying identity or confirming before a change, carried none. That fits a label that checks the final database state rather than how the agent got there. Splitting the question narrows the gap to Opus but doesn’t close it. The same person who saw Jev fail on the single question also wrote the six, so this was development on familiar cases, not an untouched validation of the method.

On 1,000 more cases, a small gap showed on MT-Bench

We drew the next 500 ChaosMNLI items and the next 500 MT-Bench pairs, in the same random order and disjoint from the first samples, and asked Jev and Opus 5.5 about each. MT-Bench ran in the original answer order only. This cost about $17.

500 new cases per dataset, Opus with the score fix
DatasetJev accuracyOpus accuracyJev AUCOpus AUCPaired
ChaosMNLI0.764 (0.725–0.799)0.742 (0.702–0.778)0.9150.917No difference detected (p = 0.16)
MT-Bench0.842 (0.807–0.871)0.878 (0.846–0.904)0.9260.938Opus ahead by 3.2 points (0.8–5.7, p = 0.01)

On ChaosMNLI the result held: the two judges ranked cases about equally well, and Jev was two points ahead, which isn’t significant. On MT-Bench, Opus was about three points more accurate. The first run’s 150 pairs had little power to detect a gap that size. The table gives Jev’s accuracy over all 500 pairs, but Opus’s over the 493 it answered. On those same 493 pairs, Jev scored 0.846, which is the basis for the paired 3.2-point gap. Counting Opus’s five abstentions and two errors as wrong gives it 0.866 over all 500. The ranking gap was small, 0.013 of AUC, with an interval that touches zero. Jev’s probabilities were slightly better calibrated on both datasets, with an expected calibration error of 0.04 against 0.05 on MT-Bench, and 0.15 against 0.19 on ChaosMNLI.

Jev first, Opus when Jev is unsure

If Jev is nearly as good and several hundred times cheaper, one option is to let Jev answer every case and pass a case to Opus only when Jev’s probability is close to 0.5. We tested this as an offline replay of recorded answers, not a deployed routing feature in Rubrist; costs below are estimated from the calls that routing would select. The question is how close. On the first samples we used two folds: choose the band on one half, score the other half, then swap and pool the held-out predictions. On the short datasets the cascade’s observed accuracy was at least as high as either judge alone, at 6 to 29 percent of Opus’s cost. That does not establish equivalence or superiority. On tau-bench it only matched Opus by sending it 91 percent of the runs, because Jev’s probability there didn’t say which runs it got wrong.

That result didn’t survive new cases. We chose the band on the whole of each first sample, using the rerun’s Opus answers, and applied it unchanged to the 500 new cases.

Offline cascade replay: 500 new ChaosMNLI cases and 493 answered MT-Bench pairs
Dataset and bandSent to OpusCascade accuracyJev aloneOpus aloneCost per 1,000
ChaosMNLI, band from the first sample32%0.7480.7640.742$4.41
MT-Bench, band from the first sample3%0.8520.8460.878$0.76
MT-Bench, two-fold band selection on new cases30%0.8740.8460.878$6.28

On ChaosMNLI the old band sent a third of the cases to a judge that was no better there, and ended below Jev alone. On MT-Bench it sent almost nothing, because on the first 150 pairs Jev had been ahead, so it stayed close to Jev’s accuracy. Repeating the two-fold procedure on the 493 answered new pairs selected bands of 0.25 and 0.35 around 0.5. Across the pooled held-out predictions, the cascade came close to Opus’s observed accuracy: 0.874 versus 0.878, at about 30 percent of its cost. This does not show the two methods are equivalent, and it leaves out the 7 pairs Opus didn’t decide.

This replay suggests a cascade can recover much of a stronger judge’s edge for a fraction of its price, but the useful band depends on the criterion and the sample. Choose it on representative development cases, measure it on held-out cases, and check again when the traffic changes. This experiment does not establish a universal minimum sample size. On 150 to 200 cases, even the sign of the gap between the two judges wasn’t stable.

What to do with this

  1. Try a cheap probability model for short, self-contained criteria like the entailment and pairwise judgments here. On 700 entailment cases no difference from Opus showed. On 650 pairwise comparisons it was level on the first sample and about three points behind on the second. At its price it can score every trace rather than a sample, so decide whether a few points matter for your criterion, and measure it.
  2. If a stronger judge is better by a few points, test whether the cheap model’s uncertainty identifies cases worth sending to it. Choose the band on development cases and verify it on held-out cases for that criterion. A band chosen on 150 didn’t carry over here.
  3. Keep the strongest judge you can afford for criteria that need the whole agent run. Here that was Opus 5.5. Its lead held at low effort, and turning thinking on didn’t detectably help Sonnet 5, so the model looks like the bigger factor. Don’t assume a judge that works on replies scales to trajectories.
  4. Run pairwise judges in both orders and check that a score points the way you think. Both problems were invisible until we looked.
  5. Measure each criterion against human labels before you trust any judge on it. Public benchmarks tell you what to try, not what to ship. Rubrist keeps an independently labelled sample and compares the evaluator against it. It now supports TypeSafe’s Jev alongside LLM judges through Anthropic and OpenAI-compatible APIs. The evaluator’s question, model and settings are versioned so you can compare a change against the human reference. The cascade above remains an experiment, not an automatic routing mode in the app.

What this does not establish

These are exploratory runs on public data, not a comprehensive benchmark or governed human truth. We generally made one call per case and configuration, with the retries and reruns described above; this is not repeated-trial validation. Most MT-Bench pairs rest on one expert’s vote, and the tau-bench label is an automated check. Any of these models may have seen these datasets in training. The samples left out long cases: ChaosMNLI premises over 600 characters, MT-Bench pairs over 9,000 characters, and tau-bench runs over 24,000 characters, which dropped 6 of 115 tasks. The main tables predate our fix to the score’s direction; the rerun doesn’t change their conclusions. Sonnet 5 and Opus 5.5 ran at their default temperature, and Opus with its default thinking and without a forced tool call; Haiku and Sonnet 4.6 ran at temperature 0. The models are not deterministic. Between two runs on the same cases, Jev’s probabilities moved by 0.01 to 0.02 on average, and 2 of its 109 tau-bench verdicts flipped. With a few hundred cases per dataset, small differences stay invisible. Latency was measured at a concurrency of four. Prompt injection was not tested. Treat the numbers as a reason to measure your own criteria, and as a hint about where to start.