
In March this year I started working on an idea I believe changes how we do RAG.
In March this year I started working on an idea I believe changes how we do RAG. Vectorless RAG. The idea has been around for a while; I designed my own process, which borrows a lot from the llms.txt convention: a document becomes a table of contents with real page ranges, and retrieval navigates it the way a reader would, instead of embedding chunks and hoping. I think vectorless is the best way to do this. The catch was always latency. It depends entirely on the model, and on the language you build the pipeline in. I built mine in Go for concurrent document processing, and the language was never the problem. The model round-trips were. I tried to cut them for months and couldn't. Two days ago @typesafeai released Jev. I integrated Jev into Vectorless and rebuilt the pipeline around one rule: a generative model only where text must actually be written. Everything else, every "is this a contents page", "does this section begin here", "is the answer on this page", goes to Jev as a batch of yes/no questions, hundreds per request. Jev was the magic wand I had been waiting for. Here is what happened, benchmarked on FinanceBench 10-K filings, 68 to 549 pages each. The PepsiCo filing is 549 pages. Vectorless with Jev now builds its full table of contents, every section placed on its page, in 8.1 seconds, for $0.0006. The same stage on the chat model took between 100 seconds and 14 minutes per filing, and nine of the 21 filings didn't finish inside a 5-minute limit at all. How each number came to be: How each benchmark was run ๐ง๐ฎ๐ฏ๐น๐ฒ ๐ผ๐ณ ๐ฐ๐ผ๐ป๐๐ฒ๐ป๐๐ โข 21 FinanceBench 10-K filings, 68 to 549 pages. โข The chat client was swapped for one that fails on any call. Zero generative calls is proof, not a claim. โข Checked against FinanceBench's gold evidence pages: 47 of 47 fall inside a section of the tree. โข The section holding a gold page is 36 pages at the median. The old path gave 183, because it lost every section's page on long filings. โข Trees compared title by title with the chat model's trees: 0.975 recall, 0.966 precision. ๐ฆ๐ฝ๐ฒ๐ฒ๐ฑ โข Whole table-of-contents stage: 9.5 seconds per filing on average, 7 to 24 seconds, $0.0003 to $0.0015. โข Same stage on the chat model: 103 to 840 seconds. 9 of 21 filings never finished inside 5 minutes. โข With an adaptive limiter: 21 filings in 122 seconds wall instead of 199, zero rate limits. ๐ฅ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น โข 40 FinanceBench questions, each with a gold evidence page. โข Jev ranks sections, then page heads, then full pages. Four requests, $0.003, 9 to 20 seconds per question. โข Right section chosen: 37 of 40. Every gold page in the evidence: 34 of 40. โข The chat model is asked once, at the end, over the pages Jev chose. It never navigates. I have not yet benchmarked this against a chunk-and-embed pipeline on the same questions. I will, and I'll post it either way. Engine: https://t.co/Den4Zbffuh Every benchmark, with the numbers I'm less proud of: https://t.co/8t24aZkwO2 Website: https://t.co/UOt704tU9t Leave a star as you check this out.