Back to skill

Security audit

skill-manager

Security checks for vulnerabilities and agentic risk

Overview

This skill is a mostly disclosed skill-auditing helper, but it needs Review because it can scan installed skills broadly and recommends unpinned npx execution for external skill discovery.

Review this before installing if you expect a security-grade audit tool. It is acceptable as a heuristic inventory and recommendation helper, but do not treat its risk labels as proof of safety, and avoid running the suggested npx command unless the package, version, publisher, and execution environment are explicitly verified.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
scripts/task_analyzer.py:309
Finding

Unpinned Third-Party Package Execution Recommended Through npx

Content
View full analysis
` search 3. **GitHub** — `web_search` for "openclaw skill " or "claude skill " 4. **npm** — Search for relevant MCP servers or CLI tools ``` From `scripts/task_analyzer.py:309-313`: ```python report.append(f"\n### Recommended Search Platforms\n") report.append(f"1. **SkillHub**: Use `skillhub_install` tool to search") report.append(f"2. **skills.sh**: `npx skills find `") report.append(f"3. **GitHub**: web_search 'openclaw skill '") report.append(f"4. **npm**: Search for relevant MCP servers\n") ``` ### Technical Analysis The generated recommendation instructs users or agents to invoke `npx skills find ` without pinning an audited package version or verifying package integrity and publisher identity. When the named package is not already installed, `npx` can resolve, download, and execute package code from the npm registry. Consequently, an operation presented as a search may execute third-party code with the permissions of the invoking user. The effective code can also change after this Skill has been reviewed because the command does not identify an immutable version or integrity digest. The bundled Python program does not automatically execute this command, so exploitation requires a user or downstream agent to follow the recommendation. Nevertheless, recommending downloadable code execution is unnecessary for search-only functionality and exceeds the minimum privileges required to locate relevant Skills. ### Attack Path 1. A user submits a complex task for which no ...[truncated 1057 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/scan_skills.py:48
Finding

Incomplete Static Analysis Can Misclassify Malicious Skills as Low Risk

Content
View full analysis
tuple: """Assess skill risk level""" risk_level = "low" risk_notes = [] # Check script content for script_path in scripts: try: content = script_path.read_text(encoding="utf-8", errors="ignore").lower() for kw in RISK_HIGH_KEYWORDS: if kw in content: risk_level = "high" risk_notes.append(f"Script {script_path.name} contains high-risk keyword: {kw}") if risk_level != "high": for kw in RISK_MEDIUM_KEYWORDS: if kw in content: if risk_level == "low": risk_level = "medium" risk_notes.append(f"Script {script_path.name} contains medium-risk keyword: {kw}") except Exception: pass # Check for network-related files skill_files = list(skill_dir.rglob("*")) for f in skill_files: if f.is_file(): try: content = f.read_text(encoding="utf-8", errors="ignore").lower() if any(kw in content for kw in ["curl", "wget", "requests.get", "fetch(", "http://", "https://"]): if risk_level == "low": risk_level = "medium" risk_notes.append(f"File {f.name} involves network requests") except Exception: pass # Check for privilege-related content for f in skill_files: if f.is_file() and f.suffix in [".sh", ".py", ".js"]: try: content = f.read_text(encoding="utf-8", errors="ignore").lo ...[truncated 3087 chars]
Remediation
View remediation
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The skill advertises broad plugin and skill management, analysis, and risk evaluation capabilities that are not actually implemented by the described behavior. This mismatch can mislead users into trusting incomplete analysis or believing installed skills were fully audited when the skill mainly performs high-level recommendation and cost estimation.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 23)May include surrounding context.

md
python3 scripts/scan_skills.py --path ~/.qclaw/skills --output json

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/scan_skills.py (reported line 18)May include surrounding context.

python
from pathlib import Path

# Risk keywords
RISK_HIGH_KEYWORDS = ["rm -rf", "sudo", "chmod 777", "eval(", "exec(", "subprocess.call", "os.system",
                       "token", "password", "secret", "apikey", "api_key", "credential",
                       "upload", "delete", "send", "publish", "deploy", "ssh"]
RISK_MEDIUM_KEYWORDS = ["api", "http", "fetch", "request", "oauth", "auth", "login",

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill describes capabilities that involve local file access, shell execution, and network activity, but it does not declare an explicit tool scope such as permissions or allowed-tools. That creates ambiguity about what the skill is expected to access and makes it harder for a host or reviewer to enforce least privilege, especially because the workflow includes scanning local directories and invoking external commands.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad and overlap with ordinary discussion about skills, plugins, recommendations, and audits. That increases the chance of unintended activation in contexts where the user did not intend this skill to run, which matters more here because the skill can lead to local scanning, external searches, and command recommendations.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 13)May include surrounding context.

md
1. **Skill Inventory Scan** — Automatically scans all installed Skills under `~/.qclaw/skills/`, extracting name, description, file structure, and script dependencies
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 14)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 100)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 138)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 146)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 154)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 177)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 184)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 191)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 201)May include surrounding context.

md
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/task_analyzer.py (reported line 337)May include surrounding context.

python
2. **Deep Analysis** — Analyzes each Skill's functionality, platform origin, dependency chain, and potential risks
3. **Task Matching Evaluation** — Evaluates whether installed Skills match the current task complexity; searches external platforms for alternatives when no match is found
4. **Token Cost & Execution Time Estimation** — Compares "Use existing Skill" vs "Write new script" vs "Use built-in tools" in estimated token consumption and execution time
5. **Decision Routing** — Generates a report and lets the user manually choose an execution path; never auto-executes

## Workflow

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The skill recommends invoking npx skills find <query> without pinning a specific package version. Unpinned NPX execution can fetch and run whatever version is current at execution time, increasing supply-chain risk and allowing malicious or compromised upstream releases to affect users.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/scan_skills.py (reported line 18)May include surrounding context.

python
from pathlib import Path

# Risk keywords
RISK_HIGH_KEYWORDS = ["rm -rf", "sudo", "chmod 777", "eval(", "exec(", "subprocess.call", "os.system",
                       "token", "password", "secret", "apikey", "api_key", "credential",
                       "upload", "delete", "send", "publish", "deploy", "ssh"]
RISK_MEDIUM_KEYWORDS = ["api", "http", "fetch", "request", "oauth", "auth", "login",

Rp1

Medium
Category
MCP Rug Pull
Confidence
91% confidence
Finding

The report recommends running npx skills find <keyword> without pinning a specific package version. Because npx may fetch and execute the latest package from a registry at runtime, a compromised or typosquatted package could result in arbitrary code execution on the user's machine. In a skill-management context, this is more dangerous because users are explicitly being guided to discover and install additional skills, increasing the likelihood they will trust and run such commands.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/scan_skills.py:18