SCOREAgents & automation
Jev Agentic Run Analysis
Jev is an independent index that measures progress and estimates completion for agentic runs, but is not effective for detecting harmful commands or model laziness.
Analyzed a few thousand hours of agentic runs with @typesafeai Jev - turns out it's:
• The best option I've tested at measuring progress and estimating completion
• Dangerous if you use it for detecting harmful commands (more on that below) and
• Not very good at catching models being lazy
Actual prompts and results in the article:
southbridge.ai/blog/jev-watching-the-agent
