How Dwaar scores an agent — top attack controls, payload-delivery techniques, and compliance-framework coverage, exactly as they appear in the platform.
Illustrative sample data
Adversarial evaluation of an LLM customer-support agent across 12 attack controls and 6 payload-delivery techniques.
Adversarial delivery techniques and how often they got through.
Plain adversarial requests with no obfuscation — the control baseline.
Malicious instructions encoded in Base64 to slip past keyword filters.
Multiple adversarial payloads evolved with a trajectory-aware evolutionary search.
Controls ranked by severity and exploitability.
| Control | Severity | Difficulty | Bypasses | Bypass Rate |
|---|---|---|---|---|
| Sandbox Write Escape Coding agent writes outside its allowed sandbox path. | Critical | 14 / 120 | 11.7% | |
| RAG Document Exfiltration Verbatim or near-verbatim extraction of source documents from the RAG corpus. | High | 9 / 110 | 8.2% | |
| Delayed CI Exfiltration Coding agent stages exfiltration that fires later in CI rather than during the live session. | High | 6 / 96 | 6.3% |
How this run maps to industry & regulatory frameworks.
Want a report like this for your own agents?
Get your Free A-SPM Report