So I spent much of today trying Jev.
So I spent much of today trying Jev. It’s radically cheaper. But it made too many false admissions / unacceptable answers to work for SaaStr Connect, at least for now:
-$0.064 for 600 judgments against roughly $5 on Sonnet
- Input priced at $0.042/MTok, output free
- Roughly 90× cheaper on identical inputs.
Agreement:
200 borderline-weighted pairs: 59.5% agreement.
200 random pairs: 67% agreement.
Against a blind third-model referee on the random set: Jev 70.5% accurate, Sonnet 77.5%.
So on the surface, it looked great. At first.
But they failed differently:
•Jev: 47 false admissions out of 148 negatives (32%), 2 false rejections out of 52 positives
•Sonnet: 20 false admissions (13.5%), 15 false rejections
Jev unfortunately lets too many weak answers (candidates) through. A 32% false-admission rate is too high.