Developer toolsRANK #333
View on X
@typesafeai Before shipping we replayed real user questions through it: Jev put
@typesafeai Before shipping we replayed real user questions through it: Jev put the right docs higher (nDCG@5 0.65 vs 0.56, graded by an LLM judge).
@typesafeai Before shipping we replayed real user questions through it: Jev put the right docs higher (nDCG@5 0.65 vs 0.56, graded by an LLM judge).