← Index

VERIFYAgents & automation

Cultivar Agent Skills and Docs Eval Library

Cultivar is a library for evaluating how well agents use documentation and skills against specified tasks, now incorporating the Jev model for binary yes/no decisions.

A new release of our Agent Skills and Docs eval library, Cultivar is available! Now, use @typesafeai's Jev model to evaluate how well your agents use docs and Skill against specified tasks in Modal sandboxes. This works great, because Jev allows for binary yes/no decisions calibrated with respect to criteria, exactly the format Cultivar works best with! Preliminary results benchmarking grading agent runs with Jev over Claude produces a 19x speedup and is 38x cheaper.