ROUTEOther
Carpooling Optimization Benchmark
This project introduces a new benchmark for carpooling optimization, comparing the performance of a decision model (JAV) against large language models (LLMs).
I'm working on a new type of benchmark (carpooling optimization), which measures the performance of decision model (JAV) against LLMs. Here is the quick summary @typesafeai:
Claude Opus: 74.1%, 2.1sec, $60
GPT-5.6 Sol: 74.9%, 2.2sec, $27
Gemini 3.2 flash: 72.9%, 1.5sec, $9
JEV: 74.3%, 0.13sec, $0.34