ROUTEOther
Carpooling Optimization Benchmark
This project benchmarks decision models (JAV) against LLMs using carpooling optimization.
I'm working on a new type of benchmark (carpooling optimization), which measures the performance of decision model (JAV) against LLMs. Here is the quick summary @typesafeai:
Claude Opus: 74.1%, 2.1sec, $60
GPT-5.6 Sol: 74.9%, 2.2sec, $27
Gemini: 72.9%, 1.5sec, $9
JEV: 74.3%, 0.13sec, $0.34