RED-04

Red-Teaming & Evals

A reproducible lab for breaking agents — test harnesses, tooling (Garak, PyRIT, Promptfoo), security benchmarks and CTF-style challenges.

The red-team evaluation loop: plan an attack, run it against the agent, observe the outcome, score it, then iterateplan01attack02observe03score04findings feed the next round — coverage grows until the agent stops failing
Pillar

Offensive Red-Team Frameworks for AI Agents: Garak, PyRIT, and Promptfoo in the Lab

A hands-on offensive comparison of Garak, PyRIT, and Promptfoo for red-teaming AI agents in the lab: what attack techniques each tool automates, how they differ, and when to reach for which.

Read

Other domains