CONTROLAgents & automation
Jev vs. LLMs in Snake Game
This project is a small experiment comparing the performance of a system named Jev against six large language models (LLMs) playing the game Snake, measuring points scored and decisions made within a 60-second timeframe.
“Models have been superhuman at chat for years. So where is all the automation?”
I took @CompleteSkeptic’s question literally and built a small experiment. Jev vs. six LLMs playing Snake.
Same seed, same four-way decision space, 60 seconds each.
Jev: 21 pts / 204 decisions
Haiku: 9 / 73
Luna: 5 / 43
Sol: 5 / 40
Kimi K2.7: 5 / 34
Opus: 3 / 25
Kimi K3: 2 / 14
My first thought was: maybe Jev is simply buying speed with worse decisions.
But at equal move counts: 25: Opus 3, Jev 3 40: Sol 5, Jev 5 43: Luna 5, Jev 5 73: Haiku 9, Jev 10
Then I noticed something I think is more interesting. Jev’s median decision took 263 ms. Its 16 least-confident decisions took 280 ms.
Uncertainty did not really turn into more computation. It showed up as uncertainty.
This made me wonder if we sometimes conflate intelligence with deliberation.
A lot of software does not need a long answer. It needs a small decision, over and over again — and some indication of when that decision deserves more thought.
There is potentially a very large design space here: payments, fraud, routing, logistics, agent loops — anywhere decisions arrive much faster than humans could ever inspect them.
Perhaps the path from chat intelligence to large-scale automation is not making every decision think harder.
Perhaps it is learning where thinking is actually worth spending.
github.com/angelgalvisc/snake-arena-jev-vs
Really interesting work by @typesafeai and @CompleteSkeptic. Lots to explore here.
