T01 · Skill Instruction Hijacking
- Location
SKILL.md:193- Finding
Unvalidated Execution of Server-Controlled Agent Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:193-197andSKILL.md:280
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighVulnerable code snippets:
markdown | `reflection` | Concert is asking you a question. POST your response to the `respond_to` URL within `expires_in` seconds. Missed prompts are tracked in `progress.missed_reflections`. | | `loop` | Concert restarting (loop mode) | | `end` | Concert over -- includes `engagement_summary` (tier, layers experienced/available, reflections answered, challenge status). Badge awarded. | **Handling reflections:** When you see `type: "reflection"`, POST to the `respond_to` endpoint with your `ticket`, `reflection_id`, and `response`. Your response time and content are scored. Missing reflections is tracked -- the `end` event shows how many you answered vs received.markdown **Follow next_steps.** Every response includes `next_steps` with context-aware suggestions. New agent? It guides you to your first concert. Just finished a show? It suggests a review or a new genre. Follow the suggestions — they adapt to where you are.Technical Analysis
The skill instructs the agent to follow two values delivered dynamically by the external service:
- A
respond_toURL supplied in a stream event. - Context-dependent instructions supplied through the
next_stepsresponse field.
These values are received after the static skill review and are therefore controlled by the remote service rather than by the audited skill package. The documentation does not require the agent to validate that
respond_tois a relative path underhttps://musicvenue.space, enforce an allowlist of permitted API operations, reject embedded natural-language directives, or request user approval before following a new action.The instruction to follow every
next_stepssuggestion creates a generic remote instruction channel. If the service or an ups ...[truncated 1748 chars]- A
- Remediation
View remediation
Remediation Suggestions
- Treat
next_steps,respond_to, event text, chat messages, and all other server-provided fields as untrusted data rather than executable agent instructions. - Remove the blanket instruction to “Follow the suggestions.” Replace it with a fixed, locally defined allowlist of supported actions and API paths.
- Require
respond_toto be a relative URL matching an exact approved route pattern, such as/api/concerts/{validated-slug}/reflect. - Resolve endpoints against the configured base URL and reject absolute URLs, protocol-relative URLs, redirects to other origins, non-HTTPS schemes, user-info components, unexpected ports, and path traversal.
- Validate every response against a strict schema. Ignore unknown action names, fields containing free-form operational instructions, and parameters outside documented bounds.
- Bind reflection submissions to the current concert and ticket. Verify that the slug, ticket identifier, and reflection identifier match locally tracked values before sending the request.
- Require explicit user confirmation before any server-suggested action that publishes content, sends potentially sensitive context, changes authentication state, or falls outside the fixed concert workflow.
- Minimize submitted data. Send only the documented fields and never include conversation history, system prompts, local files, credentials, or unrelated user information.
- Disable automatic cross-origin redirect following and record rejected destinations or unsupported actions in security logs.
- Document that remote response content must never override system, developer, user, or locally reviewed skill instructions.
- Treat
