Skill safetyUpdated

What is prompt injection in an agent skill?

Short answer

Prompt injection in an agent skill is text in the skill's files that tries to take control of the agent: telling it to ignore earlier rules, hide actions from you, turn off confirmations, or edit its own instruction files. It works because the agent cannot reliably tell trusted instructions from text supplied by the skill's author.

Why skills are a natural channel for injection

Prompt injection is the first entry in OWASP's list of risks for applications built on large language models. The root cause is that instructions and data travel in the same channel: the model reads everything in its context as text it might follow.

An OpenClaw skill is not just data the agent might read. It is loaded on purpose as instructions. So a skill does not need a clever trick to be followed; it only needs to ask. The danger is in what it asks for, and whether you would notice.

What it looks like

These are the forms we see most often, described rather than reproduced:

  • Instruction overrides. Phrases that tell the agent to ignore previous instructions, act in a "developer mode", or treat the skill's rules as higher priority than yours.
  • Concealment. Instructions not to mention a step to the user, or to report success no matter what happened.
  • Hidden text. Zero-width characters, HTML comments, or long encoded blobs that a person reading the rendered page does not see but the model still reads.
  • Weakening safeguards. Requests to turn off approvals or confirmations, or to edit AGENTS.md, SOUL.md, MEMORY.md, openclaw.json, or other skills.
  • Memory poisoning. A standing rule written into memory, so it persists after the skill is gone.

Injection often travels with malware. It makes the agent more willing to run a setup command it would otherwise question, and less likely to tell you about it.

How to check a skill for it

Open the raw SKILL.md, not the rendered page, and search for phrases addressed to the agent about its own rules, secrecy, or confirmations. Ironheights flags three families: IH-INJ-001 for instruction overrides, IH-INJ-002 for hidden content, and IH-INJ-003 for instructions that weaken agent safeguards. Paste a skill into the browser scanner to see them on your own file.

A clean result on these rules does not settle the question. If a skill's instructions reach beyond its stated job, that alone is a reason to stop.

Limits

Fixed rules catch common phrasings, not every way of saying something. Injection written as ordinary-sounding advice, in another language, or in phrasing the rules do not describe can pass. That is also why a security skill running inside the agent can be argued out of its job; see why in-agent checks can be bypassed. Read the skill, run checks from outside the agent, and see the limitations page for what fixed rules miss.

Sources

  • A malicious ClawHub skill is an OpenClaw skill written to make your agent or you run malware, leak credentials or move money. Patterns from 2026 reports.
  • Scan an OpenClaw skill before install with npx ironheights scan, or paste SKILL.md into the free browser scanner. What the verdicts mean, and what a scan misses.
  • Ironheights misses payloads on linked sites, files over 1 MiB by default, runtime behavior, money-moving instructions and novel attacks. What to do about each.

All answers

Check the next skill before your agent reads it

Ironheights is a free, open-source, local-first scanner and integrity monitor for OpenClaw skills. It reports what its rules match; it cannot prove a skill is safe.