Published agent-security research gives teams useful designs to test. It does not provide a universal scoreboard. The model, task, attacker budget, permitted tools and evaluation harness all affect the result.

This article compares what several papers actually report and where their conclusions stop. The studies are not Oktsec evaluations, and their results should not be presented as measurements of our product.

Start with the evaluation setting

AgentDojo provides 97 tasks and 629 security test cases for evaluating agents under indirect prompt injection. Sharing a benchmark helps researchers compare approaches, but it does not make results directly comparable when models, attacks or configurations differ.

Attack success rate measures how often an attacker achieves the objective defined by a particular experiment. Utility measures legitimate task completion. Report both, together with the tested configuration. A system that rejects every task can have a low attack success rate and still be unusable.

Detection and sanitization results

PromptArmor, a July 2025 preprint, uses an LLM to detect and remove injected instructions before the agent processes the input. Its authors report false-positive, false-negative and attack-success rates below 1% in their AgentDojo evaluation. They also evaluate adaptive attacks. Those results do not establish resistance to every future attack, but describing the work as testing only a fixed attack set would be incorrect.

CommandSans, an October 2025 preprint, removes instructions from tool outputs at token level. Its authors report a reduction from 34% to 3% attack success on AgentDojo without reducing utility in their tested settings, along with evaluations on other benchmarks.

What adaptive attacks establish

Adaptive Attacks Break Defenses, published in NAACL 2025 Findings, evaluates eight defenses and reports attack success above 50% against each after adapting the attacks.

The chronology matters. The paper was submitted in February and revised in March 2025. PromptArmor appeared in July and CommandSans in October. The earlier study cannot be used as evidence that those later defenses failed. It supports a narrower, useful requirement: evaluate attacks tailored to the actual defense being deployed.

What structural defenses add

CaMeL separates control flow derived from a trusted query from untrusted data and enforces capabilities when tools are called. Its authors report solving 77% of AgentDojo tasks with their security guarantees, compared with 84% for the undefended system. Those guarantees depend on the system design, policy and threat model described in the paper; they are not a guarantee for arbitrary agent deployments.

Before the Tool Call, a March 2026 single-author preprint, evaluates a different authorization system in a live adversarial testbed. It reports 0% success across 879 attempts under a restrictive policy and 74.6% against the model under a permissive policy, using comparable attacker populations. It also reports 53 ms median authorization latency. This is not an AgentDojo result, an independent replication or proof of zero risk in production.

A successful experiment supports a claim about its tested system. Extending that claim requires another experiment.

Compare what each study measured

Swipe or scroll horizontally to compare all columns.

Selected results reported by the authors; configurations and denominators differ.
StudyReported resultLimit when interpreting it
PromptArmorBelow 1% attack success in its AgentDojo evaluation.Evaluate the chosen model and adaptive attack settings.
CommandSans34% to 3% on AgentDojo.A result for the paper’s sanitization setup, not every workflow.
Adaptive AttacksAbove 50% against eight tested defenses.Does not evaluate the later PromptArmor or CommandSans papers.
CaMeL77% task completion under its security model.Task utility and a conditional security guarantee, not an ASR ranking.
Before the Tool CallNo successful attacks in 879 restrictive-policy attempts.A separate testbed reported in a single-author preprint.

What to test in your own workflow

  1. Threat coverage. Include poisoned documents and tool metadata, credential misuse and actions that bypass the inspected path.
  2. Policy quality. Test allowed actions that become harmful through excessive scope or unsafe combinations.
  3. Adaptive attacks. Let testers inspect the defense and adapt within a recorded budget.
  4. Operational behavior. Measure task completion, latency, false positives, timeouts and review queues.
  5. Evidence. Preserve enough detail to reproduce a finding while protecting sensitive inputs.

Our engineering preference is to combine explicit action permissions with inspection, isolation and monitoring. The literature provides reasons to evaluate each layer. It does not justify treating detection as useless or authorization as infallible.

To evaluate these choices in your environment, our AI red teaming assessment defines a workflow, threat model and test scope with your team. The prompt injection protection guide explains how content inspection, permissions and isolation fit together.

Editorial review, September 5, 2026: corrected the chronology of the adaptive-attack comparison and distinguished the authorization testbed from AgentDojo. Reported figures remain attributed to their authors.