back to index
AutomationRANK #01
ranked by jev

In March this year I started working on an idea I believe changes how we do RAG.

In March this year I started working on an idea I believe changes how we do RAG. Vectorless RAG. The idea has been around for a while; I designed my own process, which borrows a lot from the llms.txt convention: a document becomes a table of contents with real page ranges, and retrieval navigates it the way a reader would, instead of embedding chunks and hoping. I think vectorless is the best way to do this. The catch was always latency. It depends entirely on the model, and on the language you build the pipeline in. I built mine in Go for concurrent document processing, and the language was never the problem. The model round-trips were. I tried to cut them for months and couldn't. Two days ago @typesafeai released Jev. I integrated Jev into Vectorless and rebuilt the pipeline around one rule: a generative model only where text must actually be written. Everything else, every "is this a contents page", "does this section begin here", "is the answer on this page", goes to Jev as a batch of yes/no questions, hundreds per request. Jev was the magic wand I had been waiting for. Here is what happened, benchmarked on FinanceBench 10-K filings, 68 to 549 pages each. The PepsiCo filing is 549 pages. Vectorless with Jev now builds its full table of contents, every section placed on its page, in 8.1 seconds, for $0.0006. The same stage on the chat model took between 100 seconds and 14 minutes per filing, and nine of the 21 filings didn't finish inside a 5-minute limit at all. How each number came to be: How each benchmark was run ๐—ง๐—ฎ๐—ฏ๐—น๐—ฒ ๐—ผ๐—ณ ๐—ฐ๐—ผ๐—ป๐˜๐—ฒ๐—ป๐˜๐˜€ โ€ข 21 FinanceBench 10-K filings, 68 to 549 pages. โ€ข The chat client was swapped for one that fails on any call. Zero generative calls is proof, not a claim. โ€ข Checked against FinanceBench's gold evidence pages: 47 of 47 fall inside a section of the tree. โ€ข The section holding a gold page is 36 pages at the median. The old path gave 183, because it lost every section's page on long filings. โ€ข Trees compared title by title with the chat model's trees: 0.975 recall, 0.966 precision. ๐—ฆ๐—ฝ๐—ฒ๐—ฒ๐—ฑ โ€ข Whole table-of-contents stage: 9.5 seconds per filing on average, 7 to 24 seconds, $0.0003 to $0.0015. โ€ข Same stage on the chat model: 103 to 840 seconds. 9 of 21 filings never finished inside 5 minutes. โ€ข With an adaptive limiter: 21 filings in 122 seconds wall instead of 199, zero rate limits. ๐—ฅ๐—ฒ๐˜๐—ฟ๐—ถ๐—ฒ๐˜ƒ๐—ฎ๐—น โ€ข 40 FinanceBench questions, each with a gold evidence page. โ€ข Jev ranks sections, then page heads, then full pages. Four requests, $0.003, 9 to 20 seconds per question. โ€ข Right section chosen: 37 of 40. Every gold page in the evidence: 34 of 40. โ€ข The chat model is asked once, at the end, over the pages Jev chose. It never navigates. I have not yet benchmarked this against a chunk-and-embed pipeline on the same questions. I will, and I'll post it either way. Engine: https://t.co/Den4Zbffuh Every benchmark, with the numbers I'm less proud of: https://t.co/8t24aZkwO2 Website: https://t.co/UOt704tU9t Leave a star as you check this out.

1 likes 0 reposts 0 replies 3 views
View on X
Oludele Halleluyah

Posted as @hallelx2

Original post