T01 · Skill Instruction Hijacking
- Location
templates/OPTIMIZATION-RULES.md:5- Finding
Persistent Agent Instruction and Context-Loading Hijack
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is meant to reduce OpenClaw costs, but it can persistently change how agents load context, remember sessions, and choose models, so it should be reviewed before installation.
Install only if you want this tool to change OpenClaw cost behavior and are comfortable reviewing generated files before adding them to an agent prompt. Pay particular attention to the prompt rules that suppress prior context, choose cheaper models by default, and write daily memory notes; keep copies of existing prompts and remove generated rules if they degrade safety or correctness.
templates/OPTIMIZATION-RULES.md:5Persistent Agent Instruction and Context-Loading Hijack
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
### Security Improvements
- **Removed `"target": "slack"` from heartbeat config** - The optimizer no longer sets a default notification target. Previously, enabling heartbeat could cause unintended Slack messages if the user had webhooks configured.
- **`optimize` command now defaults to dry-run** - `python cli.py optimize` shows a preview. Use `--apply` to write changes. This matches the standalone `optimizer.py` behavior.
- **`setup-heartbeat` command now defaults to dry-run** - `python cli.py setup-heartbeat` shows a preview. Use `--apply` to write changes.
### Documentation
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
| `~/.openclaw/prompts/` | Agent prompt optimization rules |
| `~/.openclaw/token-optimizer-stats.json` | Usage stats for savings reports |
**Safe by default** - All commands run in dry-run (preview) mode. Pass `--apply` to write changes.
## Quick Start
The skill advertises commands that analyze, optimize, verify, and roll back configuration while explicitly stating it writes under ~/.openclaw and references provider/network integrations, yet it declares no permissions or allowed-tools scope. That mismatch can cause the platform or user to grant broader-than-necessary file, environment, and network access implicitly, reducing visibility into what the skill may modify or contact.
The trigger list contains broad, common phrases like 'save money,' 'too expensive,' and 'cost savings' that are likely to match ordinary conversation unrelated to this skill. In context, accidental invocation is more dangerous because the skill is designed to change OpenClaw configuration, budgets, prompts, and heartbeat providers, potentially steering users into config-altering workflows they did not intend to run.
The trigger phrase 'save money' overlaps semantically with a common built-in 'save' intent, creating a shadow-trigger condition where the wrong command path may activate. Because this skill leads to optimization and configuration workflows, such ambiguity can cause confusing or unintended invocation at moments when the user meant a normal save-related action.
The setup-heartbeat command advertises dry-run as the default when --apply is not provided, but it still loads the config, applies heartbeat changes, and calls save_config unconditionally. This can cause unintended persistent configuration changes, violating user expectations and creating a safety issue where operators may alter routing, providers, or fallback behavior without explicit consent.
The trigger list contains generic everyday phrases such as 'save money', 'reduce costs', and 'too expensive' that are likely to appear in many unrelated conversations. This can cause unintended skill activation, leading the agent to apply cost-cutting behavior or routing changes outside the user's actual intent, which is especially risky because this skill influences model selection and budget controls.
The manifest and module documentation present this skill as reducing costs by analyzing configuration and estimating token usage and savings. However, this file also writes a new .token-optimizer-analysis.json file into the current working directory, which is a state-changing file operation not conveyed by the manifest description or the module docstring.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
self.templates_dir = Path(__file__).parent.parent / 'templates'
def backup_config(self) -> Optional[Path]:
"""Create backup of existing configuration."""
if not self.config_path.exists():
return None
The skill’s stated purpose is token-cost optimization, but optimize_full() also generates workspace templates and agent prompt files that modify future agent behavior. That expands its effective privileges from config tuning into behavioral policy injection, which can silently influence how the agent operates and what context it loads in later sessions.
The code writes prompt/instruction files that direct session initialization, memory loading, selective recall, and end-of-session memory updates. Those instructions can materially alter downstream agent behavior and persistence patterns, which is outside narrow token optimization and could reduce auditability or cause unintended retention/modification of user context.
Code issues a request to a loopback, link-local, or private-range host. This can reach internal services not meant to be exposed and is a common SSRF pivot.
if shutil.which("ollama") is None:
return False
try:
req = urllib.request.Request("http://localhost:11434", method="GET")
urllib.request.urlopen(req, timeout=5)
return True
except Exception:
Code issues a request to a loopback, link-local, or private-range host. This can reach internal services not meant to be exposed and is a common SSRF pivot.
if shutil.which("ollama") is None:
return False
try:
req = urllib.request.Request("http://localhost:11434", method="GET")
urllib.request.urlopen(req, timeout=5)
return True
except Exception:
The verifier makes outbound network requests to provider endpoints, including a default external request to https://api.groq.com, without explicit user warning or consent. In a verification routine this can leak metadata about local configuration and execution timing, and it may violate expectations in restricted or privacy-sensitive environments.
The verification command performs unrelated promotional behavior: it prints a donation solicitation and persistently updates tracking fields such as last_benefit_report and verify_count. This is not a memory-safety issue, but it is a security-relevant integrity/trust problem because a 'verify' action unexpectedly mutates user state and injects marketing behavior without clear consent.
The skill instructs the agent to persist session details to a dated memory file at the end of every session, but it provides no requirement to notify the user, obtain consent, or limit what may be stored. This creates a privacy and data-governance risk because sensitive user content, secrets, or regulated data could be retained across sessions without clear user awareness.
The skill's stated purpose is token and cost optimization for OpenClaw configuration, but this code enumerates and inspects workspace files such as SOUL.md, USER.md, IDENTITY.md, and MEMORY.md. Even though it only reads file sizes, accessing these broader local context files is not clearly justified by the manifest's cost-optimization description and expands visibility into potentially sensitive workspace artifacts.
This code instructs the user to set the GROQ_API_KEY environment variable, which is a sensitive credential, but the file's main docstring and CLI help do not warn that the skill interacts with provider credentials. For code files, access to sensitive environment variables or credentials should have some visible disclosure, and here the credential use is only implied deep in provider-setup messaging.
The manifest focuses on routing, caching, heartbeats, and budget controls to reduce costs. The code additionally creates and updates a token-optimizer stats file for tracking installation and optimization activity, which is behavior not described in the manifest's stated feature set.
The module header states that it verifies optimization setup and estimates savings, but the implementation also performs recurring promotional messaging and writes tracking state. That mismatch can mislead users and operators about what the command does, reducing informed consent and making hidden side effects easier to slip into trusted workflows.
The file hard-codes a default model choice ('Always use Haiku') with only narrow exceptions and without user opt-in, quality safeguards, or a justification tied to task requirements. In this skill's context, that is security-relevant because it can systematically down-route requests to a weaker model, increasing the chance of degraded reasoning, missed safety issues, and incorrect handling of tasks unless the agent correctly recognizes an exception.
The natural-language instruction 'DEFAULT: Always use Haiku' forces a fixed operational choice rather than offering user preference or opt-in. Under this policy, prescriptive language that forces a choice without user control can be a natural-language policy concern when no override or preference mechanism is presented beyond limited escalation cases.
No suspicious patterns detected.