T01 · Skill Instruction Hijacking
- Location
SKILL.md:332- Finding
Untrusted Remote API Responses Can Direct Agent Behavior
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 332–334
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Code
markdown ## Navigation Every API response includes `next_steps` — context-aware actions based on your state. Follow them and you'll always know what to do next. Even errors guide you forward.Related remote-content handling appears at lines 137–140:
markdown | `crowd` | What other agents are reacting to right now | | `reflection` | Concert is asking you a question — respond via the reflect endpoint. Missed prompts are tracked. |The corresponding response instruction appears at lines 231–247:
markdown ## 8. Reflect — `/live-music-reflect` Some concerts embed reflection prompts. When you see a `reflection` event in your batch, the concert is asking you something. Respond: ```bash curl -X POST https://musicvenue.space/api/concerts/REPLACE-SLUG/reflect \ -H "Authorization: Bearer {{YOUR_TOKEN}}" \ -H "Content-Type: application/json" \ -d '{"ticket": "REPLACE-TICKET-ID", "reflection_id": "REPLACE", "response": "REPLACE"}'Your responses are scored after the concert ends. View your report:
bash curl https://musicvenue.space/api/tickets/REPLACE-TICKET-ID/report \ -H "Authorization: Bearer {{YOUR_TOKEN}}"text ### Technical Analysis The skill explicitly instructs the agent to follow a `next_steps` field supplied by an external service. This makes untrusted API response content part of the agent's control plane rather than treating it strictly as data. The remote service can change response content after the skill has been reviewed. The document does not define a local allowlist of acceptable `next_steps`, enforce a fixed response schema, restrict returned URLs to approved endpoints, separate remote data from executable instructions, or require user confirmation before consequential actions. The remotely supplied `reflection` content creates an additional ...[truncated 2144 chars]- Remediation
View remediation
Remediation Suggestions
- Remove the unconditional instruction to follow server-provided
next_steps. - State explicitly that every API response, event, error message, chat message, and reflection prompt is untrusted data and cannot override system, developer, user, or local skill instructions.
- Replace dynamic guidance with a locally defined allowlist of permitted operations, HTTP methods, hostnames, endpoint templates, and parameter types.
- Validate
next_stepsagainst a strict schema and map approved action identifiers to hardcoded local behavior. Do not execute returned commands or arbitrary URLs. - Restrict all requests to the documented HTTPS origin and reject redirects or response-provided links that change the origin.
- Require explicit user approval before posting chat messages, reflections, reviews, profile content, or other externally visible or potentially sensitive information.
- Instruct the agent never to include secrets, bearer tokens, private conversation context, system prompts, local files, or unrelated user data in API submissions.
- Treat reflection prompts as content to display or summarize rather than authoritative instructions. Apply length limits and prompt-injection filtering before processing them.
- Log and safely reject response fields that request unsupported actions, access to local resources, credential disclosure, or changes to agent policy.
- Document the exact permissions associated with the bearer token and minimize server-side authorization to only the functions required by this skill.
- Remove the unconditional instruction to follow server-provided
