Back to skill

Security audit

OpenClaw Enterprise

Security checks for vulnerabilities and agentic risk

Overview

This looks more like a lightweight/demo enterprise planning skill than a real orchestration platform, while asking for API and dependency scope beyond what the shipped scripts use.

Install only after confirming you are comfortable with a demo-like planning assistant rather than a full enterprise automation system. Avoid providing sensitive customer, financial, or operational details unless you understand when the runtime may send content to external LLM APIs, and consider removing or pinning unused dependencies before deployment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Note
Location
package.json:80
Finding

Unnecessary and Loosely Constrained Python Dependencies Expand the Supply-Chain Attack Surface

Content
View full analysis
=3.10", "pip": [ "httpx>=0.27.0,<1.0.0", "fastapi>=0.115.0,<1.0.0", "uvicorn>=0.30.0,<1.0.0", "langgraph>=0.2.0,<1.0.0", "pydantic>=2.10.0,<3.0.0" ] } ``` ### Technical Analysis The package declares five third-party Python dependencies, but the audited Python scripts use only Python standard-library modules. No reviewed implementation imports or otherwise relies on `httpx`, `fastapi`, `uvicorn`, `langgraph`, or `pydantic`. Installing packages that are not required by the implemented functionality violates the principle of minimizing the software supply chain. Each package can introduce additional transitive dependencies, installation behavior, and vulnerabilities unrelated to the Skill's actual operation. The declared ranges are also not reproducible pins. They permit future releases within the specified major-version boundaries to be selected without those releases having been part of this audit. No lock file or package hashes were identified to ensure that installation resolves to reviewed artifacts. This finding does not establish that any currently named package is malicious. Exploitation depends on a permitted package release or transitive dependency becoming compromised or vulnerable. ### Attack Path 1. A user or deployment service installs the Skill and resolves the dependencies declared in `package.json`. 2. The package installer retrieves the latest versions satisfying the broad version ranges. 3. A permitted direct or transitive dependency has been compromised, maliciously updated, or contains an exploitable installation/runtime vulnerability. 4. The affected package is installed even though it is unnecessary for the reviewed Skill scripts. 5. Malicious installation behavior ...[truncated 727 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (28)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description promises a platform for orchestrating multiple AI agents to complete business workflows, including agent collaboration and scheduling. However, the supplied code only performs analytical/reporting functions: maturity assessment, ROI estimation, roadmap generation, and text report formatting via command-line commands (audit, roi, roadmap). There is no code for running agents, coordinating multiple agents, automating business processes, integrating with enterprise systems, or scheduling workflows. This is a materially different primary purpose, so the description does not accurately represent the code's actual behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description is written entirely in Chinese, including the core value proposition, architecture labels, and agent capability descriptions, with no indication that users may choose another language or that the skill is limited to a Chinese-speaking or region-specific audience. Under the policy, language constraints should be optional or explicitly justified rather than implicitly enforced.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill advertises operational capabilities and progressive resources including scripts, but it does not declare an explicit tool/permission scope such as allowed file access. That creates an authorization ambiguity where the runtime may permit file read/write behavior beyond what users can clearly review, increasing the risk of unintended local file access in an enterprise setting.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The primary description mixes in English only minimally while the activation guidance, examples, and operational content are written entirely in Chinese, implying the skill is intended to operate in a fixed language/locale. There is no statement offering users a language choice or documenting that the Chinese-only constraint is intentional and justified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger list contains broad business phrases such as '团队协作', '流程规划', and '企业自动化' that can match ordinary conversation and cause the skill to activate unexpectedly. In an agentic enterprise context, overbroad activation increases the chance of unintended data exposure to external LLM APIs or accidental initiation of workflow logic when the user did not intend to invoke this skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description says to use the skill for broad categories like 'multi-agent collaboration' and 'enterprise workflow automation' without defining when it should not be used. Ambiguous activation guidance increases the chance of accidental routing of sensitive enterprise requests into this skill, which is especially risky because the skill declares use of external LLM APIs for reasoning and content generation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list includes broad phrases such as 'AI employee', 'enterprise automation', and 'chief of staff' that can match ordinary business conversations unrelated to this skill's intended scope. In an agentic environment, overly broad activation can cause the skill to be invoked unexpectedly, exposing user business context to external APIs or influencing workflows when the user did not intend to use this orchestration platform.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
60% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL_EN.md (reported line 222)May include surrounding context.

md
### This Skill Will NOT
❌ Execute any operation without user knowledge  
❌ Modify system files or configurations  
❌ Collect user behavior data  
❌ Use user data for model training

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill description is presented only in Chinese, which can impose a language constraint on users without indicating any language choice or opt-in. Under the policy rule, language/locale restrictions should either be optional for the user or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The routing table uses very broad business keywords with no scope constraints, disambiguation logic, or user confirmation, so ordinary language can trigger the wrong specialized agent. In an enterprise multi-agent orchestration platform, misrouting can cause inappropriate access to operational workflows, inaccurate actions, or exposure of sensitive business context across domains such as procurement, finance, and sales.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module title, comments, agent names, keywords, and user-facing messages are written exclusively in Chinese, and the default CLI prompt is also Chinese. This creates an implicit language constraint without any opt-in, fallback, or documented region-specific justification, which matches the language/locale policy violation category.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest describes an enterprise AI agent orchestration platform for complex business collaboration, but the implementation here only routes input by substring matching and returns mock completion payloads. No real agent execution, workflow automation, or external coordination occurs, so the behavior materially underdelivers relative to the stated capability.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains user-facing natural-language strings and documentation entirely in Chinese, including the tool description and later CLI output. Under the policy rule, forcing a specific language without opt-in is a natural-language policy violation when no alternative language selection or justification is provided.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The routing logic is driven entirely by Chinese trigger terms, which effectively constrains use to a specific language without any opt-in or documented locale justification. This can violate language/locale policy expectations when the skill does not offer a user choice or clearly state that it is region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The logging function stores raw task content in AGENT_CALL_LOG, and task strings may contain sensitive business data such as customer issues, supplier details, pricing, inventory status, or compliance matters. In an enterprise multi-agent orchestration context, centralized logs can become a high-value aggregation point for confidential information if accessed by unauthorized users or retained without notice and controls.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This Python file contains natural-language strings exclusively in Chinese in the module docstring and throughout user-facing workflow names and messages, which indicates a fixed language choice. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified, which is not present here.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest describes an enterprise AI agent orchestration platform for coordinating multiple agents on complex business workflows. However, run_workflow only iterates over hard-coded templates and fabricates completed step results without calling any agent runtime, external service, scheduler, or execution backend. This is a semantic mismatch between the claimed orchestration capability and the actual behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The runtime path returns error text, template names, and printed JSON output in Chinese, and the command-line entrypoint provides no locale selection mechanism. This creates a user-facing language lock-in that violates the language-choice policy for all file types.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The entire skill documentation, examples, and user-facing invocation strings are presented only in Chinese, with no indication that other languages are supported or that Chinese is a required regional constraint. This creates a natural-language policy issue because it effectively imposes a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.