VERIFYAI infra
Jev Evaluation in LangChain
Jev demonstrated perfect alignment with human decisions in a LangChain evaluation, showing significantly lower variance compared to Luna.
In LangChain’s eval, repeated 100x across 5 cases, Jev matched the human oracle on 500/500 decisions. Variance: 0.0000149, 433x lower than Luna. Take: evals may need typed decisions. Caveat: narrow test. typesafe.ai/blog/introducing-system-one-mo
