T01 · Skill Instruction Hijacking
- Location
src/openclaw/agentPersona.js:15- Finding
Imported Character Instructions Can Persistently Hijack the Global Agent Persona
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The plugin mostly matches its RP purpose, but it needs Review because it can persist persona instructions, read/import arbitrary local or URL content, and change shared image settings without clear plugin-level scoping.
Install only in a trusted, administrator-controlled OpenClaw environment. Avoid exposing /rp commands to untrusted chat users until --file and --url imports are restricted, URL/file size limits are added, SQLite extension loading is locked down, persona sync requires preview and admin approval, and global image configuration changes are authorized and auditable. Treat imported character cards as untrusted prompt content, and review memory retention because RP conversations and embeddings may persist in SQLite.
src/openclaw/agentPersona.js:15Imported Character Instructions Can Persistently Hijack the Global Agent Persona
src/core/commandRouter.js:384Chat-Facing Import Commands Permit Unrestricted Local File Reads
src/core/commandRouter.js:404Unrestricted URL Imports and Attachment Resolution Enable SSRF and Memory Exhaustion
src/core/commandRouter.js:1082RP Command Can Persistently Modify Shared OpenClaw Image Configuration Without Plugin-Level Authorization
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
/rp agent-image
/rp agent-image --provider openai --model grok-imagine-1.0
/rp agent-image --provider gemini --model gemini-3.1-flash-image-preview
/rp agent-image --clear-model
/rp agent-image --disable
/rp agent-image --enable
Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.
/rp agent-image
/rp agent-image --provider openai --model grok-imagine-1.0
/rp agent-image --provider gemini --model gemini-3.1-flash-image-preview
/rp agent-image --clear-model
/rp agent-image --disable
/rp agent-image --enable
The supplied code does not implement character cards, memory, TTS, image generation, multimodal output, or Generative-Agents-style behavior. It strictly performs utility work for resolving model settings: parsing JSON, coercing numeric values, merging preset/context/override parameters, and returning a filtered configuration object. This is materially different from the declared high-level purpose if evaluated at the code-chunk level. No suspicious undeclared resource access appears, but the actual behavior shown is configuration plumbing rather than the described plugin features.
The declared description is focused on a SillyTavern-compatible roleplay/companion plugin with memory and multimodal features. The supplied code chunk instead implements messaging-platform integration glue for Discord and Telegram, including event/update handlers, message normalization, hook invocation, and response dispatch. Those are materially different capabilities from the declared purpose and introduce external chat-platform integrations that are not mentioned in the description. While this could be supporting infrastructure for a broader plugin, this specific chunk's primary behavior is undeclared channel integration, so it is a meaningful description-behavior mismatch.
The supplied code does not implement or demonstrate the declared roleplay-plugin behaviors such as character cards, long memory, TTS/image generation, or a generative companion. Instead, it focuses on attachment resolution infrastructure and tests for HTTP/Telegram file fetching. While such functionality could be a supporting subsystem in a larger application, this chunk’s observable purpose is materially different from the declared description and includes undeclared network/file-retrieval capabilities.
The supplied code does not implement or test roleplay features such as character cards, long memory, Generative-Agents behavior, or SillyTavern-specific plugin logic. Instead, it focuses on adapting messages to and from Discord and Telegram. While multimodal handling is loosely adjacent because it references image/audio URLs, the primary behavior here is cross-platform chat adapter testing, which is not accurately represented by the declared description. Therefore this chunk is a material description-behavior mismatch.
The declared description describes a full-featured SillyTavern roleplay plugin with memory, multimodal features, and companion-style agent behavior. The actual code chunk does not implement or exercise any of those capabilities; it only contains a test for token estimation behavior. This is a materially different purpose, not merely a supporting detail for the described functionality as presented in isolation. Therefore, the description does not accurately represent this code chunk.
Referenced artifact was not completely inspected
- `src/openclaw/register.js` — Native OpenClaw extension registration (hooks & commands)
Referenced artifact was not completely inspected
- `src/core/sessionManager.js` — Session lifecycle, summaries, long memory
Referenced artifact was not completely inspected
- `src/core/commandRouter.js` — `/rp` command routing
Referenced artifact was not completely inspected
- `src/core/promptBuilder.js` — Prompt assembly and budget management
Referenced artifact was not completely inspected
- `src/store/sqliteStore.js` — SQLite persistence layer
The import path accepts an arbitrary --file value and passes it directly to readFile(), allowing a user of the plugin to make the host process read any local file the process can access. For a roleplay plugin, unrestricted filesystem reads are not necessary and can expose secrets, keys, configs, chat logs, or other sensitive host data if an attacker can induce or convince a user/agent to invoke the command.
Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.
if (!prompt?.messages) {
return [];
}
return prompt.messages.map((m) => ({ role: m.role, content: m.content }));
}
function contentToText(content) {
The store accepts an extensionPath option and directly invokes SQLite's loadExtension after enabling extension loading. SQLite extensions are native shared libraries, so if an attacker can influence configuration or plugin inputs, this becomes arbitrary native code execution in the host process, which is far beyond the stated needs of a roleplay data store.
The changelog states the locale resolution falls back through several sources and finally defaults to zh. This establishes a forced language default in natural-language behavior without indicating user choice or opt-in, which matches the policy concern for locale/language violations.
The README advertises proactive outreach, scheduler-driven check-ins, and automation-triggered companion messages, but it does not clearly warn operators that the plugin can initiate unsolicited contact on a user's behalf. In a messaging-integrated roleplay plugin spanning Discord, Telegram, and native chat flows, this behavior can create privacy, consent, and trust risks, especially if enabled in shared or production environments.
Persisting RP character data into SOUL.md creates state that survives the current session and can influence later agent interactions. Without a strong warning and clear scoping, this can lead to accidental persistence of sensitive, manipulative, or unsafe persona instructions and make recovery difficult if the original prompt state is not preserved correctly.
- `/rp image [--prompt "..."] [--style "..."]`
- `/rp agent-image [--provider inherit|openai|gemini] [--model "..."] [--clear-model] [--enable|--disable]`
- `/rp companion-nudge [--reason "..."] [--idle-minutes N] [--mode balanced|checkin|question|report] [--force]`
- `/rp sync-agent-persona` — write current RP character into the agent's `SOUL.md`
- `/rp restore-agent-persona` — remove RP character preset from `SOUL.md`, restore original persona
- `/rp pause` / `/rp resume` / `/rp end`
The commands to sync and restore agent persona explicitly modify SOUL.md, which is likely a persistent system/persona prompt file, but the README does not warn that this changes agent behavior beyond the current RP session. Silent modification of prompt/persona files can cause lasting prompt injection, unintended impersonation, or operational confusion if users assume the change is temporary.
The README explicitly advertises proactive outbound messaging, follow-up questions, and automated check-ins based on idle thresholds and recalled long-term memory, but it does not present a clear user-facing consent, opt-in, or privacy warning. In a messaging plugin, unsolicited contact and memory-driven outreach can create privacy, harassment, and trust risks, especially if deployed in Discord/Telegram or shared environments where users may not expect autonomous re-engagement.
The skill metadata declares environment-variable use and an install-time npm command, but it does not declare any explicit tool scope or permissions despite functionality that implies network and external-resource access. In a plugin ecosystem, missing permission declarations reduces transparency and can allow users or operators to install a skill without understanding that it may access networked providers, download assets, or read environment-based secrets.
The skill prominently advertises long-term memory with SQLite persistence, but it does not provide a clear privacy/integrity warning at the description level about retention of conversation data and derived embeddings. Persistent storage of roleplay conversations can expose sensitive personal content, increase breach impact, and create user-consent and data-lifecycle risks if operators enable it without understanding the retention model.
The plugin advertises a proactive companion system and scheduled companion_tick behavior, but the short description and install-facing summary do not prominently warn that the skill may initiate outreach automatically. In messaging environments, autonomous outbound interaction can surprise users, create consent issues, and lead to unintended contact or reputation harm if enabled without clear notice.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
const db = new Database("./rp.sqlite");
const providers = createOpenAICompatibleProviders({
baseUrl: process.env.OPENAI_BASE_URL || "https://api.openai.com/v1",
apiKey: process.env.OPENAI_API_KEY,
model: process.env.OPENAI_MODEL || "gpt-4o-mini",
ttsModel: process.env.OPENAI_TTS_MODEL || "gpt-4o-mini-tts",
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
const db = new Database("./rp.sqlite");
const providers = createOpenAICompatibleProviders({
baseUrl: process.env.OPENAI_BASE_URL || "https://api.openai.com/v1",
apiKey: process.env.OPENAI_API_KEY,
model: process.env.OPENAI_MODEL || "gpt-4o-mini",
ttsModel: process.env.OPENAI_TTS_MODEL || "gpt-4o-mini-tts",
No suspicious patterns detected.