Our Journey to Multi-Tenant Learned Re-Ranking for Commerce
Session Abstract
Search is essential to how shoppers discover products, and we asked whether machine-learned re-ranking could improve search effectiveness across hundreds of commerce clients. This is that journey: from a baseline and a null A/B test, through modelling, architecture, and evaluation, to a client revenue win, and toward mass generalisation.
Session Description
Commerce search at Coveo serves hundreds of enterprise clients from one platform – and no two are alike. A fashion retailer and a B2B parts distributor have different catalogues, shoppers, and ideas of what a good result is. Maintaining a hand-tuned ranker per client doesn’t scale. AI-powered search techniques like learning to rank (LTR) have demonstrated compelling improvements to search effectiveness, but building and validating multiple LTR models for each of our clients by hand would have been intractable.
To integrate LTR in this multi-tenant setting, we chose to frame it as producing a recipe for LTR: a shared paradigm that takes a client’s data as input, produces a model tuned to their search behaviour, and is served within the client’s commerce search pipeline. In this talk, we cover our journey of building and validating the first successful version of this recipe, including:
- How getting online with a low-risk, unsophisticated baseline mattered more than getting it perfect. We cut a thin “tracer bullet” through the whole stack – real infrastructure, real observability – to reach a trustworthy production signal quickly, weighing the “blast radius” of each step so shoppers were never harmed. Our first A/B test moved click rank but not revenue. That null result wasn’t a failure; it was essential online feedback to begin progressing from.
- From that null result we climbed to our first “base camp”. We re-grounded the evaluation – simplifying rather than adding, and choosing metrics that best tracked online impact – then made deliberate bets on features and training objectives. Successive architectural decisions each served what we needed to learn at that stage, rather than prematurely investing in parts of the system that did not take us toward our next required insight.
- The last stretch is our ongoing push toward generalisation: proving and refining the recipe across the whole customer base, not just a few wins. A recurring lesson keeps surfacing: there is no one-size-fits-all LTR.
By the end of this session, you’ll have learnt:
- What “multi-tenant” means in our particular setting, and how it differs from single-tenant LTR.
- How we integrated business rules alongside machine-learned re-ranking.
- How we approached this greenfield problem: start narrow with a thin vertical slice, expand and generalise progressively.
- How we progressed through model iterations, from failure to success.
- How we architected for generalisable LTR – serving hundreds of enterprise commerce clients, and millions of end-users.
- About the role of empirically testing generalisation through mass A/B testing in a multi-tenant setting.
Matt Williams
Coveo