Module 2

LLM-as-judge

Using a model to grade a model: when it works, how to build judge prompts, and how to keep the judge honest by calibrating against human labels.

4 lessons

  1. 2.1When a model should grade a modelAnd when code should instead25 min
  2. 2.2Writing judge promptsCriteria, examples, output shape22 min
  3. 2.3Calibrating against human labelsTrust, but verify the judge16 min
  4. 2.4Common judge failure modesPosition bias, leniency, drift22 min