Can Jev replace LLM as a Judge? We compared them...

Why we need evals
Our system generates spoken lessons with a large language model. Generation is not the hard part. Deciding whether a generated lesson is good enough to put in front of a student is.
We automate that decision with judges: nineteen of them, each an LLM call with a rubric and a structured output. A judge reads a lesson and returns a pass/fail verdict per question. Seven of the judges are gates. A gate that fails blocks the lesson.
That makes the judges the most important models in the pipeline, and it means they need their own quality bar. A judge that passes everything looks great in production and is worthless. A judge that fails good lessons wastes retries. A judge that passes bad lessons ships them.
So we grade the judges. We keep a gold set: a few hundred lesson questions that people read and labeled by hand, pass or fail, on every rubric. We run the judges on the gold and compare each verdict to the human label.
Two numbers come out per judge:
- Agreement: the share of verdicts matching the human. Bar: 0.90.
- Cohen's κ: agreement corrected for chance. If 90 percent of items pass, a judge that always says pass gets 0.90 agreement for free. κ removes that. Bar: 0.6.
We also record the confusion matrix, because κ alone does not say which way a judge is wrong. That turned out to be the whole story.
Why we looked at Jev
Cost and speed. Every lesson goes through nineteen judges, and every judge is an LLM call that reads the whole lesson and writes a paragraph per question. At the decision level, each of those calls is answering a yes/no question. We were paying text-generation prices and text-generation latency for a bit.
Jev is roughly a hundred times cheaper and a hundred times faster per decision. In our runs it answered in 0.29 seconds, median, with no rate limits and no declines. An LLM judge call takes seconds, has to be throttled to a few judges at a time, and sometimes returns nothing. If Jev could match the LLM on accuracy, we could afford to run every judge on every lesson, every time, instead of sampling.
What Jev is
Jev is TypeSafe's decision model. You hand it a document and a typed question, yes/no or pick-one, and it returns a probability for each answer. It never generates text.
Our judges already reduce to that. Each verdict has one decision field, a boolean or a small enum. The paragraph of reasoning the LLM writes alongside is useful to a person reading the report, but it is not the score.
How we made the comparison fair
The rule was: change nothing about the judges.
- Each judge renders its normal prompt: the rubric, the lesson, the list of questions to grade.
- Instead of that prompt going to the LLM, it goes to Jev as the document.
- For every question in the list, we ask Jev one typed question built from the verdict field's own description. "Does the stem give away the answer?" becomes a yes/no. A three-way enum becomes a pick-one.
- A yes/no passes at P ≥ 0.5. A pick-one takes the highest-probability option.
- The answers are assembled into the same verdict structure the LLM wouldhave produced, and the rest of the pipeline, coverage checks, thresholds, score aggregation, runs unchanged.
Same rubric text. Same gold. Same κ arithmetic. The only variable is who answers. That is what makes the two columns in the table below directly comparable.
We ran three gold sets: the main 164-case set covering 13 judges, a 41-case set for the grounded-faithfulness judge, and a 43-case set with 90 labeled follow-ups per judge for the four follow-up judges. 3,965 Jev requests, zero errors, median 0.29 seconds each.
The results
Jev clears 7 of 8 gates; the LLM clears 6. Jev wins 7 judges, the LLM wins 6, five are ties.
Here is a breakdown of Jev and LLM performance against our human labeled golden dataset.
How sure was Jev when it was right, and when it was wrong?
Every Jev verdict comes with a probability. We kept all of them, 2,896 per-item verdicts, and split them by whether Jev agreed with the human label.
When Jev is right, it is nearly certain: more than half of its correct verdicts sit at 0.95 or above. When it is wrong, its confidence drops into the 0.50 to 0.65 band. The two distributions barely overlap on the main judges.
That is the thing an LLM judge never gave us. A text model returns "pass" with a paragraph of reasoning that reads equally confident whether it is right or wrong. Jev returns "pass, 0.58", and that 0.58 is a flag. It says: take a second look. Route this one to a stricter check, or to a person, or to the LLM judge for a second opinion. Low confidence is not noise to be thresholded away. It is the model telling you where its errors live.
What Jev cannot do
Jev gives no reasons. Its explanation is the probability and nothing else. Our lesson generator improves itself by reading why a judge failed a draft, and a judge that only says "no, 0.73" cannot feed that loop. So Jev can own the ship/no-ship decision on the gates where it is strong. It cannot replace the judge whose reasoning teaches the generator.
What is the verdict?
Jev cannot be a drop-in replacement for existing LLM-as-a-judge. Its performance needs to be tuned and aligned against human judgment. Its confidence score is a good basis for further routing and validation, a low confidence means another checking needed. This still cuts cost and processing time down.