Every module so far assumed a user trying to succeed. Some inputs are designed to make your agent fail. This module is about testing your own agent against them before someone else does.
Hi, I'm the buildevals tutor. Ask me anything about building and running agent evals: fundamentals, LLM-as-judge, trajectories, rubrics, RAG, production, red-teaming, or benchmarks.