Training 2: LLMs as Judges for Search Result Quality

Session Abstract

Large Language Models (LLMs) transform how we build and evaluate search systems, it’s crucial to understand how to use them effectively as “judges.” This condensed, hands-on training introduces the principles and practical techniques for implementing “LLM as a Judge” to evaluate search result quality.

Session Description

We’ll start with the fundamentals of search evaluation: How can search result quality be measured? How do LLM-based judgments differ from human ratings and behavioral signals? You’ll learn how to design evaluation frameworks, craft effective prompts, and define output structures that make LLM-based judgments robust, interpretable, and aligned with your goals. The session closes with a hands-on application exercise where you’ll put these techniques into practice.

The class will cover these areas:

  • Introduction to search evaluation in the age of LLMs: how LLMs compare to human judgments and behavioral data for evaluation
  • Designing LLM evaluation frameworks
  • Pointwise vs. comparative (pairwise) judgments
  • Evaluation of LLM judges
  • Adding reasoning to LLM-based evaluation
  • Context engineering: improving evaluation quality by adding contextual information
  • Using critique models to improve evaluation quality
  • Hands-on application session: applying LLM-as-judge techniques to a real evaluation task

Who should attend this training?

Suitable for everyone with beginner to intermediate expertise in search. The course gives a kickstart into using LLMs as Judges for search result quality in real-world applications.

The class will use Python Jupyter Notebooks for the hands-on lab. The code will be explained step-by-step. Some basic knowledge of Python will be beneficial.

Main Stage
14.Sep 2026
13:45pm - 17:45pm
Training
René Kriegler
René Kriegler

OpenSource Connections

Daniel Wrigley
Daniel Wrigley

OpenSource Connections