Ironheights 0.2.0: protection beyond scanning, and where it stops
Ironheights 0.2.0 is on npm. Until now the tool did one job well: you pointed it at a skill folder and it told you which rules matched. This release moves the check earlier, before a skill reaches your machine, and later, while an agent runs and after files change. This post covers what is new in plain language, how to try it, and, for each piece, what it does not do. None of it makes a skill or an agent safe, and a clean result still means only that no rule matched.
What is new in one minute
fetchandsafe-installdownload a ClawHub skill, scan it without running anything, and install it only when the scan is clean.- Advisory feed support lets the CLI match a skill against a signed list of reported skills as IH-ADV-001. The feed itself is not published yet.
- Signed baselines let
verifynotice when the saved baseline was edited. - A guard plugin for OpenClaw logs risky tool calls. It is not a sandbox.
- Every scan prints an A to F grade, and a scan that skipped files says so instead of passing.
- A bug that cut off piped
--jsonoutput is fixed. - New rules cover MCP server configuration and your OpenClaw config, and the scanner now runs on Windows.
Everything is on the features page, with example commands for each.
Check a skill before it reaches your agent
The old routine was: download a skill, put it somewhere, scan it, then copy it in. Every manual step was a place to skip the check. Now one command does it:
npx ironheights safe-install <owner>/<slug>
The command downloads the skill from clawhub.ai over HTTPS into a staging folder, never runs a file, never extracts an archive, and checks that each file's size and sha256 match what ClawHub lists. It then copies the skill into your skills folder only when the verdict is no-findings. A block or incomplete verdict is never installed, and a review verdict installs only if you add --accept-review. If a copy is already installed, it is backed up and restored if the new copy fails.
If you only want the verdict, npx ironheights fetch <owner>/<slug> stops after the scan.
What it does not do. A clean verdict means the rules did not match. It is not a statement that the skill is safe, and a payload on a linked website is still a link finding at most. These are among the few commands that use the network, and they do so only when you run them; the CLI prints each URL first.
Advisory feed support
Pattern rules read text. An advisory is different: it says someone has already reported this skill, even if its text looks harmless. The 0.2.0 CLI can download a signed feed with npx ironheights advisories update, verify an Ed25519 signature against a key built into the release, and compare a skill's name, file hashes and indicator hosts with the cached copy, offline. A match is reported as a critical IH-ADV-001 finding and blocks safe-install.
What it does not do. The feed is not published yet. Until it is, advisories update has nothing to download, and scans report nothing from it. We are not going to describe contents we do not have. When it goes live, a skill missing from the feed will still not be evidence that it is harmless.
Signed baselines
A baseline is a saved list of hashes for your skills and agent files such as AGENTS.md and SOUL.md. If an attacker can edit the list as well as the files, verify has nothing honest to compare with. In 0.2.0 you can sign the baseline with a key file you keep:
npx ironheights baseline create --key ~/.ironheights/baseline.key
npx ironheights verify --key ~/.ironheights/baseline.key
Ironheights does not create or store the key for you, and on macOS and Linux it refuses a key file that other users can read. A baseline that was rewritten, even one whose hashes still agree with each other, fails the signature check and is reported as a tampered baseline with exit code 2. We wrote about why this matters in agent file baselines.
What it does not do. Anyone who can write your home directory and also has the key can sign a new baseline. Keep the key private, and keep a copy of it off the machine.
The guard plugin
Scanning reads files. The guard watches what the agent is about to do. It is an OpenClaw plugin that checks each tool call for four risky behaviors: reading credential locations such as SSH, cloud and wallet folders and .env files; piping a download into a shell or running a file it just downloaded; contacting a host that is not on your allowlist, a paste site, a tunnel or a raw IP address; and writing to the agent's identity and memory files or to skill folders. Each hit is a redacted line in a local log.
npx ironheights guard status
npx ironheights guard log
It starts in monitor mode, which logs and blocks nothing. Enforce mode is opt-in and blocks a matching call until your policy file has an allow entry with a reason.
What it does not do. The guard is not a sandbox. It runs inside the OpenClaw process and sees only the tool name and parameters OpenClaw hands it. A compromised skill that can edit your OpenClaw config or the policy file can switch it off, and a quiet log does not prove nothing happened. We explained the underlying problem in why in-agent security skills can be bypassed. Run agents with the least access they need regardless.
An A to F grade, and an honest incomplete
Every scan now prints a trust grade from 0 to 100 points: A is 90 or more, B is 80 to 89, C is 70 to 79, D is 60 to 69 and F is below 60. The number is 100 minus the risk score the rules already produced. It is a summary, not a safety rating, and the line before it still says that absence of findings is not proof of safety.
The grade also respects a rule we care about more: a scan that could not read everything must not look clean. A file over the size limit, or a .git or node_modules directory that was not entered, makes the verdict incomplete, exits with code 3, names what was skipped, and hides the number. You can acknowledge a directory in ignoreDirs with a written reason. See the limitations page for what a scan cannot see.
The piped JSON fix
Before 0.2.0, npx ironheights scan --all --json | jq could receive a report cut off at 64 KiB when the pipe was slow. The process now waits for the pipe to accept the whole report, for --json, --format html and every other command that writes to stdout. This is a bug fix and does not change what a scan finds.
More in the release
audit-configreads your local OpenClaw config and reports nine kinds of risky setting asIH-CFG-001toIH-CFG-009.scanreportsIH-MCP-001toIH-MCP-003for risky MCP server commands, literal secrets in a server's environment and filesystem roots that cover a whole disk. The in-browser scanner does not run these.scan --since-baselinereports only new findings, so CI can fail on what changed.scan --stdinandscan --textcheck a piece of text, such as an email, with the content rules.- Suppressions need a written reason.
- A GitHub Action, a pre-commit hook and Windows CI on Node.js 20.0.0 and 24.
- An optional model second opinion adds advisory notes that never change a verdict.
Try it
npx ironheights --version
npx ironheights scan <path-to-skill>
Or install it once with npm install -g ironheights. If you would rather not install anything, the browser scanner runs the content rules on a pasted SKILL.md; it uses the 0.1.5 engine, so grade, MCP, advisory and config checks need the CLI. The full rule list is in the rules reference, and the public cases we map our rules against are in the tracker.
Numbers we did not change
The benchmark was measured on version 0.1.0 with a 20-skill synthetic corpus that we wrote, so it is a regression check and not a detection rate. We did not re-run it for 0.2.0, and we have not changed its figures. The new features are not part of those numbers.
FAQ
Does safe-install guarantee a skill is safe?
No. It installs a skill only when no rule matched, and it never installs a block or incomplete verdict. A skill can still hide behavior the rules do not describe, so read what you install.
Is the advisory feed live?
Not yet. The 0.2.0 CLI contains the signed feed support, but the feed is not published, so scans report nothing from it today.
Is the guard plugin a sandbox?
No. It logs risky tool calls from inside the OpenClaw process, and a compromised skill that can edit your config can turn it off. It is a monitor, not a boundary.
Sources
- Ironheights 0.2.0 on npm, npm.
- Ironheights v0.2.0 release, GitHub, 10 October 2026.
- Ironheights changelog, GitHub.
- Fetch and safe-install, GitHub.
- Advisory feed, GitHub.
- Baseline signing, GitHub.
- Guard plugin and threat model, GitHub.