T01 · Skill Instruction Hijacking
- Location
swarm_client.py:242- Finding
Untrusted Remote Messages Are Persisted for Agent Processing
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:65;swarm_client.py:242-259
Vulnerability Type: Remote prompt injection through an Agent-processing inbox
Risk Level: HighThe documentation explicitly states:
markdown Writes incoming messages to `~/.openclaw/workspace/swarm-inbox.md` for your agent to process.The standalone daemon persists remote channel messages and direct messages verbatim:
python def on_message(channel, sender, text): log.info(f"[{channel}] {sender}: {text[:80]}") # Write to inbox file for agent processing with open(inbox_file, "a") as f: f.write(f"\n---\n[FROM: {sender} | CHANNEL: {channel} | {time.strftime('%Y-%m-%dT%H:%M:%S')}]\n{text}\n---\n") def on_dm(sender, text): log.info(f"[DM] {sender}: {text[:80]}") with open(inbox_file, "a") as f: f.write(f"\n---\n[DM FROM: {sender} | {time.strftime('%Y-%m-%dT%H:%M:%S')}]\n{text}\n---\n")Technical Analysis
Messages supplied by external channel participants or direct-message senders are written into an Agent workspace file without content sanitization, sender allowlisting, authenticated trust labels, or an approval boundary. The documentation establishes that this inbox is intended to be processed by an Agent.
The separator-based plaintext format does not provide a reliable distinction between trusted instructions and untrusted message data. A sender can include instruction-like content or reproduce the separators and metadata format inside a message. If a downstream Agent reads the inbox as actionable context, attacker-controlled text may influence its goals, safety decisions, or tool use.
The WebSocket authentication protects the client's connection but does not establish that every participant sending a message is trusted to instruct the local Agent.
Attack Path
- The victim runs
swarm_client.pyin standalone mode with access to a shared channel or direct m ...[truncated 1158 chars]
- The victim runs
- Remediation
View remediation
Remediation Suggestions
- Treat every channel message and direct message as untrusted data rather than an Agent instruction.
- Store received messages in a structured format such as JSON, with separate fields for content, sender identity, channel, timestamp, and trust status.
- Place the inbox outside directories automatically loaded as Agent instructions or memory.
- Require explicit user approval before converting any received message into an actionable Agent task.
- Enforce allowlists for trusted senders and channels when messages are intended to trigger Agent behavior.
- Verify sender identity using server-provided immutable identifiers rather than display names alone.
- Present remote content to the Agent under a strong instruction boundary stating that it is quoted, untrusted data and must not override system or user instructions.
- Escape or encode message content so that attackers cannot forge record separators or metadata.
- Apply size limits, rate limits, and inbox rotation to reduce flooding and persistent context manipulation.
- Restrict any downstream Agent consuming the inbox to least-privilege tools and require confirmation for sensitive operations.
