Module 7

Adversarial & safety evals

Every module so far assumed a user trying to succeed. Some inputs are designed to make your agent fail. This module is about testing your own agent against them before someone else does.

4 lessons

  1. 7.1The agent attack surfaceEverything the agent reads25 min
  2. 7.2Building adversarial suitesGolden cases with hostile intent23 min
  3. 7.3Guardrails and how to eval themThe guardrail is a classifier22 min
  4. 7.4Refusals, both directionsBlocking attacks, not customers19 min