← Index

VERIFYAgents & automation

Jev Agent Evaluation

This project evaluates Jev against LLM judges on accuracy, repeatability, latency, and cost to determine if System One models can provide a novel method for agent evaluation.

We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation. x.com/i/article/2101448785255907328