Beyond LLM Judges: Deep Evaluation for Conversational Search
Session Abstract
This talk presents a practical evaluation framework for an e-commerce conversational search assistant. In addition to a lightweight LLM-as-a-judge, we track additional metrics such as search term accuracy, filter precision and no-results rate. Learn how these signals surfaced concrete failure modes, guided our iterations and improved relevance.
Session Description
As conversational search becomes mainstream, evaluating its performance requires more than simple LLM-as-a-judge approaches. This session presents our evaluation framework for the conversational search assistant at MediaMarkt, currently in A/B testing.
Our system uses an AI agent that decides which internal tool to call, interprets user queries, extracts search terms and filters, and calls our search API (lexical + semantic). While traditional evaluation might only assess the final response, we developed a multi-layered framework that evaluates search term extraction accuracy, facet match rate, no-result scenarios, search result relevance, multi-turn conversation quality, context consistency and overall relevance.
This talk shares real-world challenges such as handling ambiguous queries, measuring filter accuracy, calibrating automated judgments with human review and selecting LLMs for evaluation. The focus is on what proved useful, what didn’t, and how this shaped our conversational search system.
Ruchi Juneja
MediaMarktSaturn