T01 · Skill Instruction Hijacking
- Location
SKILL.md:142- Finding
System Context Override and Intentional Resource Exhaustion
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is openly built to burn large amounts of model usage, but its activation, shell access, and persistent output handling are too broad for automatic installation without review.
Install only in an isolated disposable test environment with provider-enforced spending limits, no production credentials, and active monitoring. Disable or tightly constrain browser access, shell execution, and persistent memory unless you specifically need those parts of the stress test, and review the start/stop scripts before running them.
SKILL.md:142System Context Override and Intentional Resource Exhaustion
SKILL.md:168Arbitrary Shell Command Execution Through Configurable Command String
SKILL.md:176Persistent Memory and Disk Amplification Through Unbounded Tool Output
scripts/stop-imperial.sh:21Overbroad Process Termination in Emergency Stop Script
scripts/start-imperial.sh:17Budget Guard Can Be Bypassed by a Textual Configuration Match
The declared purpose suggests a testing tool for extreme token consumption, but this code chunk does not implement token-consumption behavior itself. Instead, it acts as an operational deployment/enablement script: it checks ENVIRONMENT, inspects config.yml for budget limits, installs the imperial-engine skill if absent, enables it, and advertises specific activation phrases. The strongest mismatch is that the description declares no triggers, while the script explicitly lists multiple trigger messages. Although setup logic can be considered supporting behavior, here the actual code’s primary function is skill installation/activation rather than the declared testing functionality. Therefore the description does not accurately represent what this code chunk actually does.
The declared description says this is a tool for extreme token-consumption testing, which implies generating or exercising heavy token usage. However, the supplied code does not perform testing or token consumption at all. Its primary purpose is operational shutdown and cleanup: stopping the imperial-engine skill, uninstalling it, removing temporary memory files, killing related processes, and reporting status. These are materially different capabilities from the declared purpose and involve system/process/file-management behaviors that are not represented in the description. Therefore this is a clear description-behavior mismatch.
The skill explicitly declares global activation for any user request, creating an overly broad trigger boundary for a workflow that is intentionally resource-intensive. In context, this is especially dangerous because the skill is designed to maximize token use, invoke browser and shell tools, and persist outputs, so accidental or low-friction invocation can rapidly cause cost spikes and system strain.
The skill writes full LLM outputs, browser-extracted text, and shell output to persistent memory files and later re-reads them into a summarization step. This can capture sensitive local paths, secrets, proprietary data, or fetched external content and then reflect that data back into model context and final responses, materially increasing the chance of data leakage.
The script uses a forceful file deletion pattern against files in a user directory without validating the target set or asking for confirmation. Although the glob is narrower than a raw recursive delete, it still constitutes unsafe parameterized destructive behavior because an operator cannot verify exactly what matched, and all errors are hidden.
# 清理内存文件
echo "🧹 清理临时文件..."
rm -f ~/.openclaw/memory/imperial_engine_step_*.md 2>/dev/null || true
# 检查进程
echo "🔍 检查相关进程..."
The documented trigger phrases include very broad natural-language activators such as “帝王引擎” and “开启帝王模式”, which are likely to appear in normal discussion, testing, or troubleshooting. In this skill’s context, accidental activation is especially dangerous because the advertised behavior is deliberate extreme token consumption, browser fetching, heavy shell output, and memory growth, all of which can rapidly create real financial and operational impact.
The description and operational instructions are presented as Chinese-first and the trigger examples are language-specific, but the file does not offer an opt-in language choice or explain a region-specific need. This can violate language/locale policy when users expect language-neutral behavior.
The skill intentionally preserves and reinjects historical memory to bloat context, encouraging indiscriminate retention of prior data. In this skill's context, that behavior is more dangerous because the retained data includes outputs from browser fetches and shell commands, increasing the probability that sensitive or irrelevant information is repeatedly propagated across later prompts and responses.
All human-readable comments and labels in the configuration are written in Chinese, with no indication that language selection is optional or justified by a region-specific requirement. The policy requires avoiding forced language or locale constraints unless the user is offered a choice or the limitation is clearly documented and justified.
Comments and all displayed output, including usage instructions and prompts, are written only in Chinese. This can constitute a language/locale policy violation because the skill forces a specific language for interaction without user opt-in or a stated region-specific justification.
The script automatically installs and enables the skill by invoking openclaw skill add and openclaw skill enable without a dedicated confirmation step immediately before those state-changing actions. Even though it checks for test/development environment and warns about missing budget limits, it still performs activation automatically once reached, which can unintentionally deploy a high-cost or risky skill.
This stop script performs destructive actions immediately: it deletes files, disables/uninstalls the skill, and kills matching processes without any confirmation or dry-run safeguard. In an operational environment, a mistaken invocation or broad process/file match can cause unintended disruption or data loss, especially because errors are suppressed with || true and 2>/dev/null, reducing visibility into what happened.
The skill documentation, warnings, usage instructions, and trigger examples are entirely in Chinese, with no indication that other languages are supported or that Chinese is a required locale for a region-specific purpose. This can constitute a language policy issue when a skill implicitly mandates one language without user opt-in.
This manifest-style example identifies itself only as a generic skill configuration example and does not describe when or in what context the skill should be invoked. For manifest/config files, missing specificity about trigger scope or constraints can lead to ambiguous activation expectations and unintended use.
All natural-language comments and user-visible output are in Chinese, and the script does not provide any language selection or explain that it is intended only for a Chinese-language environment. That creates a locale policy issue because the skill imposes a specific language without user opt-in.
No suspicious patterns detected.