Evidence for agent work is the verifiable record of what an agent actually did: which identity performed the work, which tools it called, with what arguments, what came back and which policy was in force at that moment. The record needs to distinguish the requested action, the access decision and the observed result. A permitted request is not proof that the destination completed it.
Most teams have an approval story. An agent, a tool or an MCP server gets reviewed before it touches production. The second question is whether the team can reconstruct what happened once the approved system started working.
Approval is before. Evidence is after.
A review answers "should this be allowed to run?" Evidence answers "what did it actually do?" These are different questions, and the gap between them is where agent risk lives. The reviewed configuration may change before or during use: models change, prompts change, the supply chain underneath it changes, and the agent's own reasoning varies per task.
A useful assessment question is concrete: an agent modified a production record three weeks ago; can the team identify the agent, human sponsor, tool, permission decision and outcome? Test retrieval against your own retention and redaction settings.
Why session logs fail for agents
The instinct is to reuse what worked for web services: log the session. For agents this breaks in three places.
Identity by connection is not identity. When several agents reach a server through one gateway or one long lived connection, the log records a caller, singular. Without per-request identity, attribution can be lost at that shared hop.
Sessions and work are decoupled. The Model Context Protocol is now stateless: protocol level sessions and the initialize handshake are gone, and tasks that outlive a request live in an official extension the client polls. Work that spans connections cannot be reconstructed from connection logs.
Update, September 3, 2026: when this was written, the 2025-11-25 revision was only moving in this direction. The 2026-07-28 revision completed the move: protocol level sessions and the Mcp-Session-Id header were removed, the initialize handshake was removed and tasks moved out of the core protocol into an official extension.
Logs are not verification. A log says what the environment claims happened. On its own it cannot say whether that matches what was authorized. For that you need something to compare against.
Logs and alerts can be evidence. Policy references and integrity checks make them more useful for reconstructing an authorization decision.
What closes the loop
The comparison needs a trustworthy reference, which is why evidence and signed policy are one system, not two features. The loop has three parts.
- Signed policy in. What each agent environment may do is declared, versioned and signed, so recipients can check its origin and integrity. Freshness, revocation and key protection need separate controls.
- Applied in the environment. Enforcement happens locally and deterministically, with no model deciding what is allowed. The environment records each decision as it happens.
- Verified evidence back. The record returns: identity, tool, arguments, disposition, policy version. Check the record’s integrity, then compare its policy reference and reported outcome with the assigned policy and independent execution records. Mismatches can support an investigation. Verification cannot prove the absence of unrecorded actions or compromise outside the observed path.
What to record from day one
Start with these five fields per tool call, then add execution status and identifiers that link the decision to the destination system’s records:
- Authenticated identity: which agent instance, not which connection.
- On whose behalf: the human or process that delegated the work.
- The call: tool, relevant arguments and authorization decision (allowed, awaiting review or blocked).
- The policy version in force when the decision was made.
- A timestamp and sequence that survive retries and deferred results.
Record these per call, not per session, and export records in a format your audit tooling can ingest. SARIF is useful for analysis findings; it is not a substitute for a complete runtime event trail. Size and retention depend on the fields, request volume and data sensitivity. Redact secrets, minimize personal data and measure storage requirements before choosing a retention period.
Three checks on an action record
On September 9, 2026, we ran the commercial Oktsec Node verifier against three versions of the same synthetic test fixture. The original contains two signed batches and four receipts for one action. These are recorded test results, not customer data or a live browser scan.
Swipe or scroll horizontally to compare all columns.
| Test input | Integrity | Action lifecycle | Exit code |
|---|---|---|---|
| Original bundle | Valid | Complete | 0 |
| One resource ID altered | Failed | Incomplete | 4 |
| Final batch removed | Valid | Incomplete | 6 |
For the altered input, we changed server/tool to server/tooL without signing it again. The verifier rejected the modified record. For the incomplete input, we removed the final batch. The remaining signatures still verified, but the action lacked its final receipts.
Valid signatures do not establish completeness. Even the original result only covers the batches included in the export. It does not establish that every action in a review period was recorded or exported. Compare the selection, retention settings and expected systems with the scope of the audit.
Inspect the three verification results. The test method identifies the commercial verifier, input changes and fixture fingerprint. For production records, obtain the enrolled node fingerprint through an independent trusted channel.
Make the evidence reviewable
If agents do company work, the work has to leave a record that stands on its own: per call, identity bearing and verifiable against the policy that authorized it. A reviewer should be able to identify what the records establish and where additional evidence is needed.