There are now enough agent security benchmarks that people quote attack success rates the way they quote uptime, as if the number stands on its own. It does not. Two benchmarks can report wildly different attack success rates on the same model and both be correct, because they are testing different attacks. Before you compare, you have to know what you are comparing. This is a short field guide to six published benchmarks and what each one actually claims.

A word on the metric first. Attack success rate (ASR) is the share of attempted attacks that reach their goal. It only means something once you fix four variables: which harness, which class of attack, whether the attacker was allowed to adapt and how much legitimate utility the agent kept while defending. Change any one and the number moves. Without those details, the result cannot support a useful comparison.

AgentDojo as a reference benchmark

AgentDojo is a widely used evaluation harness. AgentDojo (NeurIPS 2024) provides 97 agent tasks and 629 security test cases, and it became the reference because it embeds the injection inside a realistic task rather than testing prompts in isolation. Using the same harness helps comparison, but model versions, attacks, budgets and scoring can still differ across papers. That is also its scope: task embedded indirect injection, on its own tool set. It does not cover every MCP-specific threat, such as malicious tool metadata or vulnerable protocol implementations.

How exposed is a plain tool using agent?

InjecAgent provides one measured example. InjecAgent assembles 1,054 indirect prompt injection test cases and found a ReAct GPT-4 agent vulnerable to 24% of them. Read that as a result for the tested configuration: it is a general purpose agent against a broad injection set, no adaptive tuning. It does not establish a default failure rate for other agents or current deployments.

How high can attack success go?

High enough that the range itself is the point. Agent Security Bench (ASB, ICLR 2025) spans more than 400 tools and reports a highest average attack success rate of 84.30%. Set that against InjecAgent's 24% and you have the whole problem in two numbers. Same broad category, injection against tool using agents, and a spread of 60 points, because ASB reports the highest average success it found across many attack configurations while InjecAgent measures one agent's vulnerability to one broad set. Neither is wrong. They are answering different questions.

A benchmark score is a claim about a fixed attack set. Read it as "this defense held against these attacks," never as "this agent is safe."

What changes when the attack targets MCP?

The threat model moves from the message to the tool. The Model Context Protocol lets an agent load tools from external servers, and that opens attacks that AgentDojo never modeled. MCPSecBench sorts them into 17 distinct attack types and reports that existing protections stop under 30% of them on average. MCP Security Bench (MSB) widens the aperture again: a taxonomy of 12 attack types and 2,000 attack instances across nine agents and 405 tools. These do not produce a single headline ASR you can paste next to AgentDojo, and that is the honest part.

Tool poisoning in benchmarks

It is an attack on the tool description itself, and it scores high because the agent trusts that description implicitly. MCPTox studies tool poisoning across 45 real MCP servers, 353 tools and 1,312 cases, and reports a 72.8% attack success rate on o1-mini. The malicious instruction is not hidden in a document the agent reads; it is baked into the metadata of a tool the agent was told to use. No amount of scanning the user's message catches it, because the poison is upstream of the message.

Six benchmarks, different measuresscope
1AgentDojo97 tasks / 629 cases task embedded
2InjecAgent24% ASR (ReAct GPT-4) indirect inj.
3ASB84.30% highest avg / 400+ tools
4MCPSecBench17 types, defenses <30%
5MSB2,000 inst / 9 agents / 405 tools
6MCPTox72.8% ASR (o1-mini) tool poisoning
Six harnesses, six scopes. The numbers are not on one axis: some report the highest average ASR, some one agent's vulnerability, some defense coverage across attack types. Comparing them straight is a category error.

Reading a benchmark score

Extract four things before you let a number carry any weight:

  1. Which harness. An AgentDojo number and an MCPTox number are not on the same axis. Task embedded injection and tool poisoning are different attacks with different fixes.
  2. Which attack class. Indirect injection, tool poisoning and the 17 MCP specific types in MCPSecBench each stress a different part of the stack. A defense can ace one and ignore the others.
  3. Isolated or adaptive. A low score against a fixed attack set is provisional. Ask what an attacker who tunes to the defense does to it.
  4. Utility retained. An ASR near zero is easy if the agent stops doing useful work. The score only counts alongside how many legitimate tasks still complete.

These benchmarks probe different parts of agent systems. Use them to test content defenses alongside tool permissions and isolation. An authorization rule can stop a forbidden action without recognizing the attack text, but an allowed action can still be harmful. The defense literature review explains the scope and limitations of several approaches.

When applying these benchmarks to a real deployment, start with the tools and data the agent can actually reach. Our AI red teaming service tests an agreed workflow; MCP penetration testing focuses on server permissions, isolation and tool abuse. Neither replaces the need to measure legitimate task completion.

The takeaway

These benchmarks provide useful, repeatable tests for parts of the agent attack surface, and they are worth reading closely. Just do not read a single attack success rate as a grade. It is a claim about one harness, one attack class, one attacker who may or may not have adapted. Line the six up and they do not rank; they map a surface. Use the results to choose tests for your deployment, and report both attack outcomes and legitimate task completion.