CHOOSEAgents & automation
๐ ๐ฟ๐ฎ๐ฐ๐ฒ๐ฑ ๐๐ฒ๐ ๐ฎ๐ด๐ฎ๐ถ๐ป๐๐ ๐๐น๐ฎ๐๐ฑ๐ฒ ๐ข๐ฝ๐๐, ๐๐ฎ๐ถ๐ธ๐ ๐ฐ.๐ฑ ๐ฎ๐ป๐ฑ ๐๐ฃ๐ง-๐ฑ.๐ฐ
๐ ๐ฟ๐ฎ๐ฐ๐ฒ๐ฑ ๐๐ฒ๐ ๐ฎ๐ด๐ฎ๐ถ๐ป๐๐ ๐๐น๐ฎ๐๐ฑ๐ฒ ๐ข๐ฝ๐๐, ๐๐ฎ๐ถ๐ธ๐ ๐ฐ.๐ฑ ๐ฎ๐ป๐ฑ ๐๐ฃ๐ง-๐ฑ.๐ฐ ๐ ๐ถ๐ป๐ถ, ๐ฎ๐ป๐ฑ ๐ฏ๐๐ถ๐น๐ ๐๐ต๐ฒ ๐ท๐๐ฑ๐ด๐บ๐ฒ๐ป๐ ๐ฎ๐ฟ๐ฒ๐ป๐ฎ ๐๐ผ ๐ฑ๐ผ ๐ถ๐.
Jev is @typesafeai's System One model, and it does something no chatbot does: it never talks to you. JSON in, JSON out, a probability on every answer.
I drew System 1 and System 2 out on the whiteboard first, step by step, then put all four models on the same 15 human-labelled questions inside my own dashboard.
๐ Github Repo (1.8k): github.com/ShenSeanChen/waku-agent
๐ป Portable memory: waku.one
โ Join our community: seanchen.io
Opus scored highest. It also took 10x longer and cost 146x more than Jev to get one extra answer right out of 15.
Accuracy stops being the deciding number somewhere, and for fast judgment calls I think we're already past it.
Picking a side is easy. Ranking how much is hard. Turns out that's true of models for the same reason it's true of us.
Watch it, save it, let me know what you think. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
