VERIFYSearch & retrieval
Jev Recall Judging Bake-off
This project compares the recall judging accuracy of official Jev against DeepSeek Flash and two open-source System One reproductions on an Apple 10-K corpus.
We ran a 6-way bake-off for recall judging on the same Apple 10-K corpus: official Jev tied our production DeepSeek Flash at 92.2% accuracy — while two open-source System One reproductions (openjev, semif) only reached 75.9%/79.7%.



