SCOREAgents & automation
Jev: A Stable and Cost-Effective Judge for Agent Evaluation
This project explores Jev, a TypeSafe model, as an alternative to traditional LLM-based judges for evaluating AI agents, demonstrating significantly improved stability, reduced cost, and lower latency in scoring agent performance.
TakeIt reads an agent state and returns typed answers plus probabilities — pass/fail, a 1–5 score, or a failure category.
Pattern✦ Agent trace evaluation with typed judges
Agent evals still sit on two stools. Code checks are fast and stable, but they only catch what you can write as rules. LLM-as-judge can score open-ended traces, but it is slow, expensive, and noisy enough that the judge itself becomes the unstable part of the system.
LangChain tests a third option: TypeSafe’s System One model, Jev, as the judge. It does not write a rationale. It reads an agent state and returns typed answers plus probabilities — pass/fail, a 1–5 score, or a failure category.
The setup is deliberately small. Same weather agent. Five frozen traces. Each judge scores every run 100 times.
The numbers are the point:
- Binary pass/fail vs a human oracle: Jev 500/500. GPT-5.6 Terra 99.8%, Luna 96.4%, Claude Sonnet 4.6 80%.
- Variance on continuous quality scores: Jev is 92–913× tighter than the LLM judges.
- Cost and latency: $0.00035 / 0.44s per call. The full judgment set is $0.34 on Jev, $28.17 on Claude.
The useful claim is not “Jev won a bake-off.” It is that online evals can stop being a sample and start looking like coverage.
At 10k traces a day, an LLM judge is a budget decision. A judge this cheap and this stable is infrastructure: denser feedback, earlier regressions, a loop you can actually leave on.
The authors flag the limits. Tiny corpus. One human labeler. Do not promote this into production truth. But the direction is clear. Agent engineering is short on judges that write essays. It is short on decision models that software can call, trust, and afford.
