The unreasonable effectiveness of BM25 for agentic search

Session Abstract

GPT-5 is shockingly good at search, and that changes the “BM25 as a baseline” story. Using GPT-5 search trajectories from BrowseComp-Plus, I’ll show how default BM25 parameters and evaluation harnesses can make lexical retrieval look weak, while real agent queries often play directly to BM25’s strengths.

Session Description

GPT-5 is shockingly good at search, and that changes the “BM25 as a baseline” story. Using GPT-5 search trajectories from BrowseComp-Plus, I’ll show how default BM25 parameters and evaluation harnesses can make lexical retrieval look weak, while real agent queries often play directly to BM25’s strengths. Much like grep became a core retrieval primitive for coding agents, BM25 is re-emerging as a powerful primitive for agentic search.

Main Stage
15.Sep 2026
14:15pm - 15:00pm
Talk