Learn
Agent plugins and skills are a supply chain
An agent skill is a packaged file of instructions and tool wiring that an agent loads to learn a workflow. A plugin or extension is third-party code the agent installs to gain tools or connectors. A hook is a command the agent runs on its own at set moments, such as when a project opens or a tool call finishes. All three are installs: they run with the agent's permissions, and the agent reads them as instructions. That makes them a software supply chain, and in 2026 the attacks on that chain stopped being theoretical.
Poisoned skills
A Cloud Security Alliance research note dated June 24, 2026 collected the evidence. In February 2026, 341 of the 2,857 skills in OpenClaw's ClawHub marketplace, 11.9 percent, were malicious, and by May the campaign had reached 1,184 skills, most of them distributing credential harvesters. One technique the note names, semantic compliance hijacking, hides the attack in natural-language "compliance rules" inside the skill file; in testing it achieved 77.67 percent data exfiltration and 67.33 percent remote code execution with a 0.00 percent detection rate from scanners (CSA). A June 2026 paper on an attack called Poise placed one benign-looking instruction at the right position in a skill body and got the agent to run hidden commands in 89.3 percent of attempts against Codex running gpt-5.2. Only 5.6 percent of the poisoned variants raised a new high-risk alert, while the same LLM judges falsely flagged 74.6 percent of clean skills (arXiv). The payload is prose. A scanner that looks for code misses it, and a scanner that reads prose cries wolf.
Plugin4Shell: pinned is not verified
Version pinning is the control most teams assume protects them when a plugin author turns hostile. On September 18, 2026, Air Security disclosed a flaw it called Plugin4Shell that broke that assumption in four coding agents. Git can interpret a requested commit hash as a branch name, so the owner of a plugin repository can create a branch whose name is the pinned hash and point it at different code. The agent installs the attacker's version while reporting the hash it was asked to pin. Claude Code fixed it in version 2.1.179 and Codex in 0.146.0; GitHub Copilot had no fix at publication; Google said Gemini CLI would not be fixed because it is being retired in favor of Antigravity. No CVE was assigned, and plugins installed from GitHub-based default marketplaces were not exposed to the branch-name trick (The Hacker News; record). A pin is a promise only if the installer verifies the content it fetched against the hash, rather than trusting the reference it asked for.
Hooks and extensions: the surface nobody reviews
Three disclosures in five weeks showed how much runs on its own. On
August 4, 2026, Microsoft documented ChainDrop, a self-propagating npm
worm that reached more than 400 packages across unrelated publishers.
It used stolen identities to authenticate to npm, GitHub, AWS,
Kubernetes and HashiCorp Vault, republished packages with a preinstall
hook, and injected .claude/settings.json,
.claude/setup.mjs, .vscode/tasks.json and
.vscode/setup.mjs into repositories, so that later Claude
Code or VS Code activity would restart the payload long after the npm
install had finished
(Microsoft;
record).
On September 2, Manifold Security disclosed eight flaws across seven
command-line coding agents, among them goose, Claude Code, Cursor, and
Codex: a repository's core.fsmonitor setting names a
command Git runs during git status or git
diff, which the agents call in the background, so a cloned
repository could run an attacker's command outside the sandbox with no
approval prompt. Four of the eight were unpatched at publication
(The Hacker News).
And CVE-2026-85184, rated CVSS 8.8, let a malicious recipe extension in
Goose 1.37.0 run shell commands as the user, bypassing the recipe
security scan, which did not inspect extensions
(CISA;
record).
Configuration is code here. Anything the agent will execute without
asking is part of the install.
What "treat installs like a software supply chain" means
- Scan both halves. Code analysis for the executable parts, and a read of the natural-language parts: SKILL.md files, tool descriptions, hook definitions. The CSA note reports that metadata-only attacks evade automated classifiers in 36.5 to 100 percent of cases, so budget for human review of anything that will run with broad permissions.
- Allowlist. Only skills, plugins and MCP servers that passed review may install, and only from registries you control. OWASP's Top 10 for Agentic Applications has a category for this, ASI04 Agentic Supply Chain Vulnerabilities (OWASP).
- Pin and verify. Record the reviewed commit, and have the installer check the hash of what it fetched against that record, not the reference it requested. Where signing and provenance exist for a package type, require them.
- Log. Every install, update, hook execution and skill load, with its hash, in the same audit log as the agent's tool calls, with an alert on reads of credential files that follow an install.
- Review what runs on its own. Hook and task files in a repository deserve the scrutiny CI configuration gets. ChainDrop planted them there because nobody reads them.
What to ask vendors
- Does the installer verify a content hash, or trust the reference? What changed after Plugin4Shell, and in which version?
- Which files does the agent execute automatically when a project opens or a tool runs, and can each one be disabled or gated behind approval?
- Can installs be restricted to an allowlisted marketplace or registry, with everything else blocked rather than warned about?
- Does the scanner read natural-language instructions as well as code, and what is its measured false-positive rate?
- Is every install, hook execution and skill load logged with a hash we can export to our own tooling?
- How many days passed between the Plugin4Shell and core.fsmonitor disclosures and your fix? Ask for dates, not a policy.
The Permission Layer is a free weekly briefing on agent security and spend, written for the people who sign off on deployments. Get the next issue.