NEEDLE: The live benchmark your search engine can't memorize

Needle: The benchmark your search engine can't memorize

NEEDLE: The live benchmark your search engine can't memorize

Static benchmarks let search engines cheat—by memorizing answers or even downloading them from HuggingFace mid-test. Keenable's new live benchmark, NEEDLE, refreshes tasks hourly and daily across news, finance, scholar, rare-entity, and legal queries, making overfitting impossible. Early results reveal that engines with independent indexes, like Keenable itself, outperform those that merely federate others, and that traditional engines optimized for human clicks fail at agentic search. The benchmark tracks improvement over time, showing who's learning fastest.

Two students with the same right answer studied. Two students with the same wrong answer sat next to each other.
  1. daft_pink

    so is this an independent search engine benchmark or a blog post from a search engine provider showing their search engine at the top of the benchmark?

    i'm a little bit confused as at first when I was reading it i thought it was a search engine benchmark, but it seems that keenable is at the top which i assume is related to the web domain owner? i've never heard of keeenable

  2. MarkusQ

    I wonder if search engines linked to human-use-case engines (e.g. google/bing) start at a disadvantage because they have been historically incentivized to break themselves to support their business models? It seems reasonable to suppose that "good at selling ads" ≠ "good at finding results".

  3. terno

    do you somehow control how non-trivial the queries are? The LLM generates them, right?

    what if every engine returns garbage, or on the other hand, handles them too well?

    building a benchmark like this in a genuinely fair way seems extremely hard to me. I’m very curious about the details, of course within what you can share.

More from this day

2026-08-27