Authorization telemetry is the record of the authorization decision on every tool call an agent makes: which identity acted, which tool it invoked, with what arguments, under which policy version and whether the call was allowed, sent for review or blocked. It is separate from model and tool-latency traces, and it helps reviewers connect requested actions to the rules that governed them.

The gap is easy to miss because agent observability looks healthy. You have spans for the model turn, the retrieval, the tool latency, the token spend. What you usually do not have is a span that answers the security question, and that question does not live in any of the others. OWASP's MCP Top 10 names the gap directly as MCP08:2025, Lack of Audit and Telemetry, and its remedy is the one this piece describes: detailed logs of tool invocations with immutable audit trails.

Why isn't tool call latency enough?

Because latency measures how long a request took. It does not establish whether the request was authorized or whether its intended effect occurred. A tool call can be fast, successful and completely unauthorized. The security event is not the duration; it is the decision that let the call through.

The common fallback is to log the chat: the prompt, the model's reasoning, the response. That is useful for debugging behavior and insufficient on its own to establish an authorization decision. The chat is upstream of the action and easy to manipulate. The authorization decision sits right at the boundary between intent and effect, which is exactly where you want your evidence.

Trace the decision, not the conversation. The conversation is why the agent wanted to act; the decision is whether it was allowed to.

What belongs in an authorization span?

These five attributes are a starting point; also record timestamps, correlation identifiers, execution outcomes and any required reviewer decision. Keep them stable across services so the span is queryable as one thing.

  1. Identity: the authenticated agent instance, not the connection or the shared account.
  2. On behalf of: the human or process that delegated the work.
  3. The call: the tool name and the arguments that matter (redact secrets, keep shape).
  4. Policy version: the exact ruleset in force when the decision was made.
  5. Verdict: allowed, review or blocked, plus the rule that decided.

OpenTelemetry setup in Python

Emit one span per tool call, named consistently, before the tool runs. The name matters: pick one identifier and use it everywhere, so the whole decision stream is one query in any backend. trace_authorization is the example convention used here, not a standard span name or a promise about an installed product version. The OpenTelemetry GenAI semantic conventions, still marked Development, define spans for invoking agents and executing tools (invoke_agent, execute_tool) and none for the authorization decision, so treat the span below as a sibling of execute_tool, not a replacement. A minimal shape:

trace_authorization · one span per tool callpython
1from opentelemetry import trace
2tracer = trace.get_tracer("agent.authorization")
3with tracer.start_as_current_span("trace_authorization") as span:
4  span.set_attribute("agent.identity", agent_id)
5  span.set_attribute("agent.on_behalf_of", principal)
6  span.set_attribute("tool.name", tool)
7  span.set_attribute("policy.version", policy_ver)
8  verdict = policy.decide(agent_id, tool, args)
9  span.set_attribute("authz.verdict", verdict.result)  # allow · review · block
10  if verdict.result != "allow": raise Denied(verdict.rule)
One span, one stable name, emitted before the call runs. Everything downstream queries it as a single stream.

A few decisions make this hold up in production. Emit the span before the tool executes, so a blocked call still leaves a record. Attach the verdict rather than inferring it from an exception later. Use a bearer or workload token to establish agent.identity; positional identity ("it came in on this connection") does not survive shared gateways or retries. And export through your existing OpenTelemetry pipeline: the point is that this lands in the backend your team already watches, not a separate console.

What questions does this answer later?

"Which agent was authorized to modify that record, on whose behalf, under what policy?" can be answered from trace_authorization spans. To establish whether the record changed, correlate the decision with the destination system’s records. To find unresolved reviews, join decision spans with approval records. Compare policy versions when investigating a change in behavior, alongside changes to tools, credentials and task inputs. Session logs cannot answer these once connections are shared or the protocol goes stateless, which MCP now has: the 2026-07-28 specification retired the initialize handshake and the Mcp-Session-Id header, so each request travels on its own. Per-call decision spans preserve the authorization context across those boundaries.

Update, September 3, 2026: this paragraph originally pointed at the May release candidate as the direction MCP was heading. The 2026-07-28 specification shipped on July 28 with the stateless core; the link and wording are updated.

This is the instrumentation view of a larger idea: the authorization decision, not the model output, is the security event worth recording. We wrote about why that record is the product in Agent work needs evidence, not trust, and why the decision has to be deterministic in Detection versus authorization.

The takeaway

If your agents call tools with real credentials, add one span to every call: the authorization decision, named consistently, exportable to your observability stack. Correlate those spans with durable audit records and execution results. Sampling, dropped exports and redaction must be accounted for before treating traces as a complete audit trail.