Most agents read before they act, and retrieval is where they fail silently: the reply sounds confident while the wrong documents sit underneath it. This module makes every leg of the triangle of question, context, and answer measurable.
Hi, I'm the buildevals tutor. Ask me anything about building and running agent evals: fundamentals, LLM-as-judge, trajectories, rubrics, RAG, production, red-teaming, or benchmarks.