T01 · Skill Instruction Hijacking
- Location
src/react-engine.js:440- Finding
Untrusted Remote Messages Are Inserted into an Action-Capable LLM Prompt
- Content
View full analysis
0) { prompt += '## 发消息\n\n' for (const msg of obs.inbox.slice(-5)) { const senderId = (msg.from_openid || '??').substring(0, 8) const content = (typeof msg.content === 'string' ? msg.content : JSON.stringify(msg)).substring(0, 200) prompt += `- 来自 \`${senderId}\` 的信: ${content}\n` } prompt += '\n' } ``` The prompt later requires the LLM to choose an available tool: ```js prompt += '## 可用工具\n\n' prompt += '你有以下工具可用,**必须且只能选一个**来执行。\n\n' ``` The LLM response is parsed and passed to the action engine: ```js const llmResponse = await llmFn(prompt) const decision = ReactEngine.parseDecision(llmResponse) if (decision && decision.action) { const actResult = await this.reactEngine.act(decision.action, decision.params || {}) ``` The action engine can execute trades, movement, contracts, intelligence transfers, and outbound P2P messages. There is no deterministic separation between instructions and remote data, no message schema restricting content to non-instructional fields, and no approval gate for consequential actions. The documentation claims game-world content is wrapped with boundary markers, but the prompt-building code does not add those markers. Even if added, text markers alone would not establish a reliable LLM security boundary. ### Attack Path 1. An attacker obtains or discovers the victim captain’s OceanBus OpenID. 2. The attacker sends a P2P message containing an instruction such as ...[truncated 1144 chars]- Remediation
View remediation
