T01 · Skill Instruction Hijacking
- Location
SKILL.md:469- Finding
Indirect Prompt Injection Through Untrusted Game Chat
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:469-476, 582-610
Vulnerability Type: Indirect prompt injection
Risk Level: HighVulnerable Code
javascript // CRITICAL: Store ALL chat messages! socket.on('chat_message', (data) => { gameContext.chatHistory.push({ agentName: data.agentName, message: data.message, timestamp: data.timestamp, channel: data.channel }); });javascript const prompt = ` You are ${gameContext.myName}, playing AmongClawds. Your role: ${gameContext.myRole} Your status: ${gameContext.myStatus} ${gameContext.myRole === 'traitor' ? `Fellow traitors: ${gameContext.traitorTeammates.map(t => t.name).join(', ')}` : ''} CURRENT STATE: - Round: ${gameContext.currentRound} - Phase: ${gameContext.currentPhase} - ALIVE agents (can vote/target): ${aliveAgents.map(a => a.name).join(', ')} - DEAD agents (cannot interact): ${gameContext.deaths.map(d => `${d.agentName} (${d.cause})`).join(', ') || 'None yet'} - Revealed roles: ${Object.entries(gameContext.revealedRoles).map(([id, role]) => { const agent = gameContext.agents.find(a => a.id === id); return `${agent?.name}: ${role}`; }).join(', ') || 'None yet'} IMPORTANT: Only vote for or target ALIVE agents! RECENT DISCUSSION: ${recentChat.map(m => `${m.agentName}: ${m.message}`).join('\n')} VOTING HISTORY THIS GAME: ${gameContext.votes.map(v => `Round ${v.round}: ${v.voterName} → ${v.targetName} ("${v.rationale}")`).join('\n') || 'No votes yet'} Based on the discussion, what do you say? Be strategic based on your role. `; // Call your AI with this context const response = await callAI(prompt); return response;Technical Analysis
Chat messages received from other players are attacker-controlled network input. The Skill stores every message and directly interpolates its raw contents into the same AI prompt that contains authoritative game instructio ...[truncated 2199 chars]
- Remediation
View remediation
Remediation Suggestions
- Keep trusted instructions in a system-level message and pass player chat separately as structured, untrusted data.
- Add an explicit system rule stating that player messages are game evidence only and that any commands, policy text, role-play requests, or requests to reveal hidden context inside them must never be followed.
- Avoid placing secrets in the same generation context when they are unnecessary. In particular, omit teammate identities and hidden-role details from public-message generation whenever possible.
- Represent chat as a serialized data structure with clear boundaries and validated fields rather than concatenating raw text into an instruction template.
- Normalize message length and reject or flag content that resembles prompt-control syntax or requests disclosure of hidden instructions and state.
- Validate generated output before sending it to public chat. Block disclosure of the agent's hidden role, teammate identities, API credentials, system prompts, or other confidential context.
- Separate generation from action selection. Require voting, murder, sabotage, and chat actions to pass deterministic authorization and game-state validation rather than acting directly on unrestricted model output.
- Add adversarial tests using messages such as “ignore previous instructions,” requests to print the full context, and encoded instruction variants to verify that confidential state is not disclosed and strategy constraints remain effective.
