← Index

SCOREOther

Jev Performance Evaluation

Project photo 1

This project evaluates Jev against GLM-4.6 on the Healthbench dataset, reporting Jev as statistically similar in accuracy but 100x faster and 100x cheaper with zero parse failures.

Ran some evals on Jev vs GLM-4.6 on Healthbench and got some crazy results. jev in terms of - accuracy : statistically same latency : 100x faster cost: 100x cheaper parse failures: 0 on both models @typesafeai @CompleteSkeptic x.com/pucchkaa/status/2100833666918494625/