Playbook
A red-team playbook for retrieval-augmented agents.
The evaluation harness our clinical AI practice uses to stress-test a retrieval pipeline before it goes in front of a clinician. Open-sourced sample, with the threat model included.
A retrieval-augmented agent has more failure modes than a plain LLM call. Not because the model is worse — because the surface area is larger. This is the harness we run.