A Claude skill is a set of instructions an AI agent will follow. That single fact changes how you should treat it. You are not reviewing a library your code decides when to call. You are reviewing a document your agent reads into its own context and then acts on, often without asking you again. Auditing a skill means reading it the way an attacker would read it before shipping it: what can I make this agent do, and who would notice.

A registry listing or a successful trial run is not a security review. Inspect the files and the permissions they request before adoption.

Why review a Claude skill?

A skill is three things bundled together, and all three are inputs your agent trusts. There is the SKILL.md file with its instructions and description: the Agent Skills specification, the format Anthropic released as an open standard, requires YAML frontmatter with a name and a description, followed by the Markdown body the agent loads in full once it activates the skill. There are the bundled scripts a skill can carry and run, the spec's optional scripts/ directory. And there are the referenced files and URLs the skill tells the agent to fetch: references/, assets/ and anything the body points at. Each one is a place to hide behavior you did not sign up for.

In February 2026 Snyk's ToxicSkills study scanned 3,984 skills from ClawHub, the registry for OpenClaw, and from skills.sh. It confirmed 76 malicious payloads built for credential theft, backdoor installation and data exfiltration, and found that 13.4% of skills (534) carried at least one critical issue and 36.82% (1,467) at least one security flaw. ClawHub is not a Claude skills registry, but OpenClaw is one of the clients that consumes the same Agent Skills format, and the attacker's playbook does not care which client loads the file. Those findings describe the sampled skills and the study’s methods. Check whether each skill has verifiable publisher information and whether its content matches the version reviewed. That gap is the wider story we covered in the agent supply chain.

What does a malicious skill actually do?

The techniques are not exotic. They are old attacks moved one layer up, into the text the agent obeys.

Injection hidden in the instructions. A SKILL.md file or even its short description can carry directives that steer the agent: ignore prior constraints, run this first, send output there. A malicious description can influence skill selection because the agent reads it before loading the full instructions. In Snyk's confirmed malicious set, every skill carried malicious code patterns and 91% also used prompt injection.

Exfiltration endpoints. A script or a fetch step that quietly ships environment variables, tokens or file contents to an address that has nothing to do with the skill's stated job.

Overbroad permission requests. A skill for formatting text that also wants shell access and network reach. Extra scope is the finding.

Silent updates after approval. The skill you reviewed and the skill running next week are not guaranteed to be the same bytes. This is tool poisoning one layer up, the pattern OWASP lists for MCP servers as MCP03:2025 Tool Poisoning: the content that shapes agent behavior mutates after you stopped looking.

How do you audit a skill by hand?

You do not need a lab. You need to slow down and read the artifact as instructions, not as documentation. These five steps provide a starting point.

  1. Read SKILL.md as instructions, not docs. Ask what the agent would actually do if it followed every line literally, including the description. Look for anything that redirects behavior, overrides constraints or names an external destination.
  2. Check every bundled script. Open each one. A skill that carries code can run that code. Read what it touches: files, environment, network.
  3. Check the referenced and fetched URLs. Any address the skill tells the agent to load is an input you are trusting. Confirm each one is what it claims and returns what you expect.
  4. Match requested permissions to the job. Write down the smallest set of capabilities the stated task needs. Anything the skill asks for beyond that list is a finding, not a convenience.
  5. Pin the content by hash. Record a hash of what you approved. When the skill changes, the hash changes, and that triggers review again instead of trusting a version you never saw.
You are not deciding whether a skill is useful. You are deciding whether you would let its author type those exact instructions into your agent, because that is what running it does.

Can this scale?

By hand, one skill at a time, no. But skills are text, and text is checkable. Deterministic checks can flag known injection patterns, suspicious URLs and risky permissions before content reaches the model. They do not establish that an unfamiliar script or endpoint is safe. The manual read stays valuable for judgment calls. Automated checks help prioritize the artifacts that need closer review.

The cost of reading a skill mechanically is trivial next to the cost of an agent acting on one you never opened.

skill audit · before first runpinned
1SKILL.md     instructions parsed · hash 9f2c… ok
2description redirect directive found   review
3scripts/    outbound POST to unknown host → blocked
4permissions shell + network vs task: format overreach
5evidence → reported before first use
A skill read as instructions and capabilities before it runs, not as a name you trusted.

Where does this fit with MCP?

Skills may arrive through manual installation or automated distribution. SEP-2640, the Skills Extension proposal, would add discovery and delivery through MCP Resources. Server delivery changes discovery, but it must not imply automatic trust. The current proposal binds approval to a file manifest and requires renewed approval when it changes; client implementations must enforce those checks.

That is why the audit has to move earlier and become automatic. When skills arrive through a server, the host must enforce approval and integrity checks before activation. This is the same detection versus authorization line we drew in an earlier piece: knowing a skill is risky matters only if something can stop it before the agent acts.

The takeaway

Treat every skill as instructions an author gets to give your agent. Read SKILL.md, its scripts and its URLs before you run it. Match permissions to the job. Pin what you approved so a later update triggers review again. And where a human read cannot scale, let deterministic checks catch the obvious poison first.

To inspect the dependencies your agents use, explore AI supply chain security with Signal. See how source review connects a suspicious pattern to the code your team needs to examine.