|
At Haystack I spoke about autoresearch: Code generation to optimize search rankers. Can we use it to improve on BM25? This article represents my lab notes. My agent starts with a BM25 implementation, proposes changes, and accepts those that improve NDCG. We’ll zero-in on passage retrieval dataset MSMarco. I won’t claim I’ve found a “better BM25” but I’ve iterated towards a decent tuning regime. All while learning valuable lessons about how validation data can leak. Let’s walk through what happened. More in my blog article: https://softwaredoug.com/blog/2026/05/17/autoresearching-a-better-msmarco-bm25 -Doug Slack Community * Events · Consulting · Training (use code search-tips) You're subscribed to Doug Turnbull's daily search tips where I share tips, blog articles, events, and more. You can always manage your profile: |
I share search tips, blog articles, and free events I'm hosting about the search+retreval industry, vector databases, information retrieval and more.
My work is split between mature search teams and new AI teams. Search teams are often farther along and manage a mature, traditional search product. AI teams, however, often don't know what they don't know yet. They've just been cobbled together, have built a few agent demos, and are often in the process of discovering the three big mistakes I blog here: Evals+measurement need to dominate a lot of your product thinking Retrieval isn't "one thing" (ie classic RAG) - its extremely custom to...
I'm writing about the weak spots in vector databases. Where you should prod and poke when selecting a vendor. Today: Updating Vector Databases If you think about the old vector search regime, it involved Indexing everything up front Never updating the index Search-only We overindexed on this paradigm, creating data structures focused on good search performance that couldn't tolerate updates. I wrote about how sensitive graph-based vector DBs in particular are to updates, picking on Lucene...
Hey all, I wrote a new article about a technique that has come up over and over in my work, especially in my Cheat at Search training, for doing effective query understanding into a large vocabulary. Instead of asking an LLM to classify into a vocabulary. Ask it to hallucinate fake entities, then resolve those to real ones client side. Save yourself a lot of tokens and use cheaper models. https://softwaredoug.com/blog/2026/08/10/hypothetical-classifications -Doug PS - A reminder that Vectors...