RAG Is Simpler Than You Think: Start with BM25

RAG Is Simpler Than You Think: Start with BM25

Most teams over-engineer their RAG stacks with embeddings and vector databases when a simple full-text search might suffice. This article breaks down six retrieval architectures, from plain BM25 to agentic decomposition, and provides a decision framework based on data freshness, corpus size, query patterns, and team expertise. It emphasizes starting with the simplest approach and only adding complexity when data proves it's necessary, with concrete cost and latency comparisons.

Most 'semantic search' problems are actually query formulation problems.
  1. usernametaken29

    I worked on large scale RAG systems before and can say people vastly underestimate full text search and vastly overestimate embeddings. FTS is really easy, portable and scalable and gets you very far, the 80/20 rule applies. Embeddings appear to be nice and magic but when you really get into them you notice: semantic similarity isn’t as good as you think and certainly it won’t make everyone happy. You will inevitably end up having to re-embed more or different chunks of your text to accommodate more and more precise embedding search - at which point you’ll go the last mile and do reranking etc etc all the while having to support the operational burden of vector search.

    Then you turn around and build a search query with 500 keywords and sure it’s painful but it just works, accommodates all use cases, scales and is overall less annoying to maintain.

  2. alansaber

    I have built systems using all of these approaches (all in tandem). For the most part, the juice is not worth the squeeze (in building a highly optimised corpus-specific information retrieval strategy) outside of a very few fringe cases. The amount of technical discussion far outstrips the use case for RAG.

  3. jillesvangurp

    RAG is basically good old information retrieval with LLMs doing the querying. This can include vector search but it works without that as well. Treating vector search as magic pixie dust that makes search great without effort is not necessarily going to work that well. Also, it can add a lot of cost and complexity to the equation. And if not tuned properly, you don't necessarily get good results.

    The key thing with RAG is to get the right information in the context with as few queries as possible. That requires good recall (ensuring that if it is there it can be found with a reasonable query) and precision (ensuring the best stuff is on top and minimizing false positives).

    With search, and by extension RAG, the principle of shit in, shit out applies. Most of what search teams did before AI and RAG is still the best way to optimize the experience with RAG. And if you mess that up, search is not going to be working that well and no amount of AI can compensate for that or only at great cost in tokens and time. So, having an ETL pipeline to pre-process what you index, testing & benchmarking search quality, etc. are all helpful.

    The good news is that you don't need that much skills with agentic coding to build something half decent for this. This code almost writes itself. And even a little bit of effort on extracting structure before indexing can make a big difference.

  4. jrochkind1

    More LLM-generated text about LLMs.

    Is anyone else actually finding it harder and harder to read LLM generated text? I find it quite tiring, my brain just does not want to get through it.

  5. Angostura

    I have a particular antipathy for articles too lazy to spell out acronyms on first use.

    So: https://en.wikipedia.org/wiki/Retrieval-augmented_generation

  6. yipinwong

    Only those who mastered the craft makes their work look simple.

    The AI that wrote this might be the master not the writer, as this looks written by AIs.

    I will use the author's agents, not read his articles or use him for the job.

  7. refactor_master

    Here’s an even simpler take: just embed everything the first time, then track what was changed. Use a cheap model to summarize and clean up the documents/chats with summary and keywords. Unless you have entire libraries of books to embed it’s going to be a few hundred dollars of API calls.

    Then, throw it all in BigQuery. Handles all the vector stuff natively.

    Sprinkle an agentic bot UI thing on top to make it appear all-knowing and magical.

    I assume other vendors than Google have a similar batteries-included approach you can just plug in.

  8. klm127

    RAG stands for Retrieval Augmented Generation. The purpose is to search a corpus of text by meaning rather than exact match.

    I had to look it up.

More from this day

2026-08-26