Using a model to grade a model: when it works, how to build judge prompts, and how to keep the judge honest by calibrating against human labels.
4 lessons