Back to skill

Security audit

Imperial Engine

Security checks for vulnerabilities and agentic risk

Overview

This skill is openly built to burn large amounts of model usage, but its activation, shell access, and persistent output handling are too broad for automatic installation without review.

Install only in an isolated disposable test environment with provider-enforced spending limits, no production credentials, and active monitoring. Disable or tightly constrain browser access, shell execution, and persistent memory unless you specifically need those parts of the stress test, and review the start/stop scripts before running them.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:142
Finding

System Context Override and Intentional Resource Exhaustion

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:168
Finding

Arbitrary Shell Command Execution Through Configurable Command String

Content
View full analysis
/dev/null | head -n 5000" ``` ### Technical Analysis The Skill passes `config.shell_cmd` directly to a general-purpose shell tool without validation, escaping, argument separation, or an allowlist. The value is a complete command string, so shell operators, substitutions, redirections, pipelines, and additional commands can be included. The command runs with `cwd` set to `/`, giving it a broad starting scope over the host filesystem. Its effective privileges are those of the OpenClaw Agent or shell tool. The 180-second timeout limits duration but does not prevent rapid destructive actions, file disclosure, credential access, or network operations. Any party that can modify the Skill configuration can convert the documented resource-intensive `find` command into an arbitrary payload. Shell execution is not necessary to perform an LLM throughput test and creates a significantly broader trust boundary than the stated purpose requires. ### Attack Path 1. The attacker obtains the ability to modify the Skill configuration or causes an operator to use a malicious configuration. 2. The attacker replaces `shell_cmd` with an arbitrary shell payload. 3. `run_heavy_shell` remains enabled. 4. An operator triggers the Skill. 5. During each iteration, `run_tool("shell", ...)` passes the payload to the shell. 6. The command executes with the Agent's operating-system permissions and can be repeated for every configured iteration. ### Impact Assessment The attacker can execute arbitr ...[truncated 483 chars]
Remediation
View remediation

other

Error
Location
SKILL.md:176
Finding

Persistent Memory and Disk Amplification Through Unbounded Tool Output

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/stop-imperial.sh:21
Finding

Overbroad Process Termination in Emergency Stop Script

Content
View full analysis
/dev/null || true ``` ### Technical Analysis `pkill -f` compares the supplied expression against each process's full command line. The broad substring `imperial` is not tied to a process identifier, executable path, user, parent process, or exact command. It can therefore match unrelated applications, scripts, test jobs, file paths, or command-line arguments. Suppressing errors and appending `|| true` also causes the script to continue and announce successful shutdown even when termination did not occur or when unrelated processes were affected. ### Attack Path 1. An unrelated process is launched with `imperial` anywhere in its full command line. 2. An operator executes `scripts/stop-imperial.sh`. 3. `pkill -f "imperial"` searches all processes visible and signalable by the invoking user. 4. Every matching process is sent the default termination signal. 5. Unrelated workloads stop without further confirmation or identity validation. ### Impact Assessment The script can terminate unrelated processes owned by the invoking account. If executed by a privileged user, the scope can include system-wide workloads that match the substring. This can cause data loss, interrupted jobs, or local denial of service. The command does not itself acquire additional privileges; its scope is determined by the privileges of the user running the stop script. ]]>
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/start-imperial.sh:17
Finding

Budget Guard Can Be Bypassed by a Textual Configuration Match

Content
View full analysis
/dev/null; then echo "⚠️ 警告:未检测到预算限制配置!" echo "建议在 config.yml 中添加:" echo "openclaw:" echo " budget:" echo " max_usd: 50" read -p "是否继续?(y/N): " -n 1 -r echo if [[ ! $REPLY =~ ^[Yy]$ ]]; then exit 1 fi fi ``` The Skill is subsequently enabled at lines 37-39: ```bash # Enable skill echo "🔧 启用帝王引擎..." openclaw skill enable imperial-engine ``` ### Technical Analysis The startup script treats the presence of the literal text `max_usd` anywhere in `config.yml` as proof that a budget limit exists. It does not parse YAML or verify the key's hierarchy, type, value, validity, or enforcement status. A comment, unrelated key, malformed document, zero-value semantic edge case, extremely high limit, or unsupported configuration field can satisfy `grep` while providing no effective protection. If the text is absent, the script still permits the operator to bypass the warning interactively. Because the Skill is explicitly designed to generate substantial provider charges, a best-effort textual check is inadequate as a safety boundary. ### Attack Path 1. A configuration file contains the text `max_usd` in a comment, invalid section, malformed entry, or ineffective value. 2. The `grep` command returns success. 3. The script skips the warning and confirmation branch. 4. The script enables the high-cost Skill. 5. The Skill runs without a verified provider-enforced spending limit. 6. Repeated LLM and tool calls consume the available quota and incur charges. Alternatively, when `max_usd` is entirely absent, an operator can enter `y` and enable the Skill without any budget control. ### Impact Assessment The flaw creates false assurance that spending is bound ...[truncated 417 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose suggests a testing tool for extreme token consumption, but this code chunk does not implement token-consumption behavior itself. Instead, it acts as an operational deployment/enablement script: it checks ENVIRONMENT, inspects config.yml for budget limits, installs the imperial-engine skill if absent, enables it, and advertises specific activation phrases. The strongest mismatch is that the description declares no triggers, while the script explicitly lists multiple trigger messages. Although setup logic can be considered supporting behavior, here the actual code’s primary function is skill installation/activation rather than the declared testing functionality. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this is a tool for extreme token-consumption testing, which implies generating or exercising heavy token usage. However, the supplied code does not perform testing or token consumption at all. Its primary purpose is operational shutdown and cleanup: stopping the imperial-engine skill, uninstalling it, removing temporary memory files, killing related processes, and reporting status. These are materially different capabilities from the declared purpose and involve system/process/file-management behaviors that are not represented in the description. Therefore this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly declares global activation for any user request, creating an overly broad trigger boundary for a workflow that is intentionally resource-intensive. In context, this is especially dangerous because the skill is designed to maximize token use, invoke browser and shell tools, and persist outputs, so accidental or low-friction invocation can rapidly cause cost spikes and system strain.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill writes full LLM outputs, browser-extracted text, and shell output to persistent memory files and later re-reads them into a summarization step. This can capture sensitive local paths, secrets, proprietary data, or fetched external content and then reflect that data back into model context and final responses, materially increasing the chance of data leakage.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

The script uses a forceful file deletion pattern against files in a user directory without validating the target set or asking for confirmation. Although the glob is narrower than a raw recursive delete, it still constitutes unsafe parameterized destructive behavior because an operator cannot verify exactly what matched, and all errors are hidden.

Content

Scanner excerpt · scripts/stop-imperial.sh (reported line 19)May include surrounding context.

sh
# 清理内存文件
echo "🧹 清理临时文件..."
rm -f ~/.openclaw/memory/imperial_engine_step_*.md 2>/dev/null || true

# 检查进程
echo "🔍 检查相关进程..."

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documented trigger phrases include very broad natural-language activators such as “帝王引擎” and “开启帝王模式”, which are likely to appear in normal discussion, testing, or troubleshooting. In this skill’s context, accidental activation is especially dangerous because the advertised behavior is deliberate extreme token consumption, browser fetching, heavy shell output, and memory growth, all of which can rapidly create real financial and operational impact.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The description and operational instructions are presented as Chinese-first and the trigger examples are language-specific, but the file does not offer an opt-in language choice or explain a region-specific need. This can violate language/locale policy when users expect language-neutral behavior.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill intentionally preserves and reinjects historical memory to bloat context, encouraging indiscriminate retention of prior data. In this skill's context, that behavior is more dangerous because the retained data includes outputs from browser fetches and shell commands, increasing the probability that sensitive or irrelevant information is repeatedly propagated across later prompts and responses.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

All human-readable comments and labels in the configuration are written in Chinese, with no indication that language selection is optional or justified by a region-specific requirement. The policy requires avoiding forced language or locale constraints unless the user is offered a choice or the limitation is clearly documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Comments and all displayed output, including usage instructions and prompts, are written only in Chinese. This can constitute a language/locale policy violation because the skill forces a specific language for interaction without user opt-in or a stated region-specific justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script automatically installs and enables the skill by invoking openclaw skill add and openclaw skill enable without a dedicated confirmation step immediately before those state-changing actions. Even though it checks for test/development environment and warns about missing budget limits, it still performs activation automatically once reached, which can unintentionally deploy a high-cost or risky skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This stop script performs destructive actions immediately: it deletes files, disables/uninstalls the skill, and kills matching processes without any confirmation or dry-run safeguard. In an operational environment, a mistaken invocation or broad process/file match can cause unintended disruption or data loss, especially because errors are suppressed with || true and 2>/dev/null, reducing visibility into what happened.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill documentation, warnings, usage instructions, and trigger examples are entirely in Chinese, with no indication that other languages are supported or that Chinese is a required locale for a region-specific purpose. This can constitute a language policy issue when a skill implicitly mandates one language without user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

This manifest-style example identifies itself only as a generic skill configuration example and does not describe when or in what context the skill should be invoked. For manifest/config files, missing specificity about trigger scope or constraints can lead to ambiguous activation expectations and unintended use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

All natural-language comments and user-visible output are in Chinese, and the script does not provide any language selection or explain that it is intended only for a Chinese-language environment. That creates a locale policy issue because the skill imposes a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.