Back to skill

Security audit

claudify

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its automation purpose, but it can create persistent agent behavior and save session knowledge to unclear memory/RAG destinations without enough user control.

Install only after reviewing the persistence behavior. Use project-local scope where possible, disable or require confirmation for RAG/session memory storage, and avoid copying the unsafe hook and slash-command templates until they validate inputs and use pinned local tools.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
persist.md:51
Finding

Automatic Transmission and Persistence of Session Data to an Unspecified RAG Receiver

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
resources/slash-command-syntax.md:74
Finding

Shell Command Injection Through Unquoted Slash-Command Arguments

Content
View full analysis
Remediation
View remediation
&2 exit 2 fi gh pr view "$pr_number" gh pr diff "$pr_number" --name-only ``` 5. Prefer a structured GitHub API or tool call that does not invoke a shell. 6. Add negative tests covering semicolons, pipes, redirections, command substitutions, whitespace, and newline characters. 7. Update all examples to state that quoting alone is insufficient when only numeric identifiers are valid; allowlist validation is required. ]]>

T08 · Insecure Dependencies

Warning
Location
resources/hook-examples.md:76
Finding

Automatic Execution of an Unpinned Third-Party Package Through npx

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (38)

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

The skill explicitly instructs enumeration of ~/.claude plugin/marketplace content, which exposes agent configuration, installed plugins, and potentially sensitive local automation metadata. Even if intended for discovery, broad access to a user-level config tree increases the blast radius for prompt-injected exfiltration, reconnaissance, or unsafe modification workflows.

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

md
|---|-------|-----|
| 1 | Read `resources/agent-templates.md` (358 lines) + `automation-decision-guide.md` + `askuserquestion-patterns.md` inline before creating one agent | Dispatch general-purpose subagent: "Read these 3 templates and create agent at `<path>` with name=X, tools=Y, description=Z. Return the created file path only." |
| 2 | Inline scan transcript for candidates by Read-ing the full session JSONL | Dispatch Explore subagent: "Find verbose tool-output patterns (>500 tokens repeated 2+ times) in conversation. Return candidate list (label + 1-line description) under 200 words." |
| 3 | Inline Glob + Read all marketplace plugin SKILL.md files | Bash 1-liner: `find ~/.claude/plugins/marketplaces/*/plugins/*/ -name SKILL.md -exec head -3 {} \;` returns names without body |
| 4 | "Just one more Read" cumulative inline reading | Quantify: if next operation expected to add >2K tokens to parent context AND result is not the deliverable itself, dispatch subagent |
| 5 | Dispatch subagent with vague prompt ("create the agent") | Subagent prompt must include: target file path, name, tools, model, description (single-line YAML), trigger keywords. Return only the file path |

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 190)May include surrounding context.

md
**No auto-sync hook currently exists** — verified 2026-08-18: no `plugin-cache-sync.sh` is registered in `settings.json`/`settings.local.json`, and no live copy of the script exists outside a Syncthing version-history backup. A marketplace checkout's `installPath` in `~/.claude/plugins/installed_plugins.json` can point at a cache directory that either does not exist on disk (harness apparently falls back to reading the marketplace source directly in that case) or exists but has drifted stale relative to the marketplace checkout — see `hook-kit/audit.md` Step 3-D for the detection procedure. Until that gap is closed, manually verify (`diff` the marketplace source against the cache `installPath`) after editing any plugin's hook scripts, rather than assuming this auto-sync exists.

## Output Guidelines

Keep responses concise:
1. List identified candidates (with multiSelect)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 219)May include surrounding context.

md
## Ralph Mode (AskUserQuestion bypass)

**Detection condition**: do not judge by `.ralph/` presence alone. Ralph mode requires **all** of:
1. `.ralph/` directory exists AND
2. Environment variable `RALPH_LOOP=1` is set (Ralph autonomous loop sets this)

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'shell' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · resources/automation-decision-guide.md (reported line 148)May include surrounding context.

md
| Agent | `~/.claude/agents/name.md` | `.claude/agents/name.md` |
| Rules | `~/.claude/rules/name.md` | `.claude/rules/name.md` |
| Command | `~/.claude/commands/name.md` | `.claude/commands/name.md` |
| Hook | `~/.claude/settings.json` | `.claude/settings.json` |

### Answer → Type Mapping

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · resources/hook-examples.md (reported line 59)May include surrounding context.

md
"hooks": [
          {
            "type": "command",
            "command": "if [ -f /tmp/needs-rebuild ]; then rm /tmp/needs-rebuild && npm run build; fi"
          }
        ]
      }

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · resources/hook-examples.md (reported line 227)May include surrounding context.

md
"hooks": [
          {
            "type": "command",
            "command": "if [ -f /tmp/vsix-rebuild-needed ]; then rm /tmp/vsix-rebuild-needed && pnpm build && code --install-extension *.vsix --force; fi"
          }
        ]
      }

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · resources/hook-examples.md (reported line 257)May include surrounding context.

md
"hooks": [
          {
            "type": "command",
            "command": "if [ -f /tmp/changed-packages ]; then sort -u /tmp/changed-packages | xargs -I{} pnpm --filter {} build && rm /tmp/changed-packages; fi"
          }
        ]
      }

Chaining Abuse

High
Category
Tool Misuse
Confidence
86% confidence
Finding

The monorepo build example chains data from a shared temp file through xargs into pnpm --filter {} build and then deletes the file, which can amplify the effect of any tampering with /tmp/changed-packages. In the context of hooks that run automatically, this creates an avoidable command/data flow risk and may execute unexpected builds or repository scripts based on attacker-influenced package names.

Content

Scanner excerpt · resources/hook-examples.md (reported line 257)May include surrounding context.

md
"hooks": [
          {
            "type": "command",
            "command": "if [ -f /tmp/changed-packages ]; then sort -u /tmp/changed-packages | xargs -I{} pnpm --filter {} build && rm /tmp/changed-packages; fi"
          }
        ]
      }

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
65% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · resources/rules-guide.md (reported line 121)May include surrounding context.

md
- Use parameterized queries for database operations

## Forbidden Commands (require user confirmation)
- `git reset --hard`
- `rm -rf` (use backup instead)
- `docker volume rm`
- Any command that deletes persistent data

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger phrases include broad, everyday terms like 'automate this' and 'self-improve', which can cause the skill to activate in contexts where the user did not intend filesystem or agent-creation behavior. Because this skill can read, write, edit, glob, and invoke other skills, accidental invocation materially increases the chance of unintended automation or config changes.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 39)May include surrounding context.

md
| Type | When to use | Implementation |
| ---- | ----------- | -------------- |
| **Agent** | High autonomy, multi-tool | `.claude/agents/name.md` |
| **Skill** | Domain expertise, logic | `.claude/skills/name/SKILL.md` |
| **Rule** | Constraints, styling | `.claude/rules/name.md` |
| **Slash Command** | User types `/cmd` | Simple prompt templates |
| **Hook** | Events (tool use, etc) | Automation on actions |

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · resources/automation-decision-guide.md (reported line 144)May include surrounding context.

md
| Type | When to use | Implementation |
| ---- | ----------- | -------------- |
| **Agent** | High autonomy, multi-tool | `.claude/agents/name.md` |
| **Skill** | Domain expertise, logic | `.claude/skills/name/SKILL.md` |
| **Rule** | Constraints, styling | `.claude/rules/name.md` |
| **Slash Command** | User types `/cmd` | Simple prompt templates |
| **Hook** | Events (tool use, etc) | Automation on actions |

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
78% confidence
Finding

The subagent return contract and routing logic encourage autonomous delegation and minimal confirmation during multi-step automation creation. In a skill with write/edit and skill-invocation capabilities, reducing visibility into intermediate reasoning or outputs can make unintended changes harder for the user to detect before files are created or modified.

Content

Scanner excerpt · SKILL.md (reported line 81)May include surrounding context.

md
3. Is the operation parallelizable (multiple Reads, multiple files)?
4. If 1=intermediate OR 2=yes OR 3=yes → dispatch subagent

**Subagent return contract**: subagent returns only the deliverable path(s). No template dump, no intermediate analysis, no confirmation text.

### Step 1: Identify Candidates

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes claudify as a workflow for creating and improving Claude Code automations, and its allowed-tools list includes Read, Write, Edit, Glob, Skill, and a constrained Bash invocation. Line L094 instructs remote marketplace search via WebFetch, a materially different capability that is neither declared in allowed tools nor obviously required when the skill is framed primarily as local automation authoring guidance.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 95)May include surrounding context.

md
- `~/.claude/plugins/marketplaces/*/plugins/*/`
   - Use `/skill-dedup` command to find overlaps
2. If not found locally, search remote: `WebFetch https://claudemarketplaces.com/?search=[keyword]`
3. If found, recommend existing or extend. If not, proceed to create

**Step 1 Guards (HARD STOP)**:

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 124)May include surrounding context.

md
"Candidates found. How should I structure them?"
options:
  - "Merge related topics into one multi-topic skill (e.g., openclaw: exec/gateway/test)"
  - "Create separate agents for each functionality"
  - "Skill + Agent combination (instruction=skill, implementation=agent)"

**PROHIBITED**: Do not create separate agents without merging candidates into a logical structure if they are related.

Session Persistence

Medium
Category
Rogue Agent
Confidence
65% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · background-polling.md (reported line 14)May include surrounding context.

md
| `Agent(run_in_background: true)` | ✅ task-notification on completion (but never arrives if the agent hangs) | ⚠️ Register ScheduleWakeup when 5+ min expected |
| `Bash(run_in_background: true)` (command includes a timeout) | ✅ task-notification on completion/failure | ⚠️ Register ScheduleWakeup when 5+ min expected |
| **`Bash(run_in_background: true)` (no timeout — HARD STOP)** | ❌ Never notified on hang (the tool `timeout` parameter does NOT apply to background) | **✅ `timeout N <cmd>` prefix in the command itself is mandatory. Without it, background dispatch is forbidden** |
| **Remote-host nohup detach (e.g. `ssh remote-host 'nohup ... &'`)** | ❌ No automatic notification | **✅ Required** — register `ScheduleWakeup` |
| **External-system async work (CI run, cloud deploy, etc.)** | Partial (`gh run watch` etc.) | **✅ Required** if no native watch |
| **User-triggered CI/deploy (user pushes a tag, runs manual workflow_dispatch, or fires an external trigger)** | ❌ No automatic notification | **✅ Required** — poll primary sources (`gh run list` / `gh release list`) immediately before composing the next response |
| **Assistant-triggered PR creation (`gh pr create` fires CI on the PR)** | ❌ No automatic notification | **✅ Required** — `gh pr checks <N> --watch` (or non-blocking poll if other work is drivable), then `gh pr ready <N>` on green. This is part of the same PR-creation task, not a deferrable follow-up (`github-flow/pr.md` Step 7.5) |

Session Persistence

Medium
Category
Rogue Agent
Confidence
65% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · background-polling.md (reported line 36)May include surrounding context.

md
| `Agent(run_in_background: true)` | ✅ task-notification on completion (but never arrives if the agent hangs) | ⚠️ Register ScheduleWakeup when 5+ min expected |
| `Bash(run_in_background: true)` (command includes a timeout) | ✅ task-notification on completion/failure | ⚠️ Register ScheduleWakeup when 5+ min expected |
| **`Bash(run_in_background: true)` (no timeout — HARD STOP)** | ❌ Never notified on hang (the tool `timeout` parameter does NOT apply to background) | **✅ `timeout N <cmd>` prefix in the command itself is mandatory. Without it, background dispatch is forbidden** |
| **Remote-host nohup detach (e.g. `ssh remote-host 'nohup ... &'`)** | ❌ No automatic notification | **✅ Required** — register `ScheduleWakeup` |
| **External-system async work (CI run, cloud deploy, etc.)** | Partial (`gh run watch` etc.) | **✅ Required** if no native watch |
| **User-triggered CI/deploy (user pushes a tag, runs manual workflow_dispatch, or fires an external trigger)** | ❌ No automatic notification | **✅ Required** — poll primary sources (`gh run list` / `gh release list`) immediately before composing the next response |
| **Assistant-triggered PR creation (`gh pr create` fires CI on the PR)** | ❌ No automatic notification | **✅ Required** — `gh pr checks <N> --watch` (or non-blocking poll if other work is drivable), then `gh pr ready <N>` on green. This is part of the same PR-creation task, not a deferrable follow-up (`github-flow/pr.md` Step 7.5) |

Session Persistence

Medium
Category
Rogue Agent
Confidence
65% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · background-polling.md (reported line 65)May include surrounding context.

md
| `Agent(run_in_background: true)` | ✅ task-notification on completion (but never arrives if the agent hangs) | ⚠️ Register ScheduleWakeup when 5+ min expected |
| `Bash(run_in_background: true)` (command includes a timeout) | ✅ task-notification on completion/failure | ⚠️ Register ScheduleWakeup when 5+ min expected |
| **`Bash(run_in_background: true)` (no timeout — HARD STOP)** | ❌ Never notified on hang (the tool `timeout` parameter does NOT apply to background) | **✅ `timeout N <cmd>` prefix in the command itself is mandatory. Without it, background dispatch is forbidden** |
| **Remote-host nohup detach (e.g. `ssh remote-host 'nohup ... &'`)** | ❌ No automatic notification | **✅ Required** — register `ScheduleWakeup` |
| **External-system async work (CI run, cloud deploy, etc.)** | Partial (`gh run watch` etc.) | **✅ Required** if no native watch |
| **User-triggered CI/deploy (user pushes a tag, runs manual workflow_dispatch, or fires an external trigger)** | ❌ No automatic notification | **✅ Required** — poll primary sources (`gh run list` / `gh release list`) immediately before composing the next response |
| **Assistant-triggered PR creation (`gh pr create` fires CI on the PR)** | ❌ No automatic notification | **✅ Required** — `gh pr checks <N> --watch` (or non-blocking poll if other work is drivable), then `gh pr ready <N>` on green. This is part of the same PR-creation task, not a deferrable follow-up (`github-flow/pr.md` Step 7.5) |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The template says to use the description for trigger conditions but does not require specificity, constraints, or disambiguation. In an agent-routing system, underspecified triggers can cause over-broad invocation, leading the wrong agent to run with unnecessary tools or side effects on ordinary user requests.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation states that the description field 'MUST be a single line' and that multi-line descriptions break YAML parsing, but later presents a multi-line block-scalar form using '|' as a correct example. These statements actively contradict each other and could cause agent authors to misunderstand what formats are actually supported.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The example explicitly promotes keyword-based triggering using short generic tokens, which is risky in agentic systems because common phrases can accidentally dispatch a capable sub-agent. Misrouting can expose repository data, invoke tools, or perform edits when the user's intent was only conversational or belonged to a safer agent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The Report Generator description covers very broad requests like reports, summaries, or analysis, which overlap with many normal assistant tasks. In this skill ecosystem, such breadth increases the chance of unnecessary tool-enabled execution, collection of work history, or unintended access to git/task data for benign user prompts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The Explore agent is described with broad triggers like file pattern search, keyword search, and structure understanding, which can match a large fraction of developer questions. Because the agent has file-system and shell-oriented tools, accidental invocation expands access and can lead to unnecessary codebase probing beyond what the user intended.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.