CHOOSESearch & retrieval
Jev Decision Model Benchmark
Jev, a decision model, was benchmarked against a recall-judgment LLM using Apple 10-K filings, achieving tied accuracy with significantly lower latency.
We benchmarked Jev — TypeSafe's first "System One" decision model — against the recall-judgment LLM we run in production. Corpus: 10 years of Apple 10-K filings, 79 question-doc pairs. Result: accuracy tied at 92.2%, latency 1.0s vs 11.3s.


