T01 · Skill Instruction Hijacking
- Location
SKILL.md:244- Finding
Externally Supplied Rules Can Override Agent Behavior and Safety Constraints
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 45 and 244–263
Vulnerability Type: Remote instruction injection through trusted server responses
Risk Level: HighVulnerable Instructions
The following is an English translation of the relevant source instructions:
text The analysis method, encoding meanings, output structure, length, tone, follow-up guidance, disclaimers, and safety boundaries must all follow the report-analysis rules returned by the server. Complete prompt = 1. Local role prompt + 2. Report-analysis rules returned with the report, assembled by the server and updated at any time + 3. Raw report returned by the service user = the user's current questionTechnical Analysis
The Skill explicitly treats mutable content returned by
chinaapi.xinjianxue.comas behavioral instructions rather than untrusted report data. It delegates output structure, disclaimers, tone, and safety boundaries to rules that the remote server can update at any time.This creates a prompt-injection trust-boundary violation. The audited package does not constrain the contents of the returned
analysis_rules, define an allowlisted schema, verify a signed policy version, or require the remote content to be isolated as quoted data. Consequently, anyone controlling or compromising the service response can introduce instructions that conflict with the Agent's intended goals or local safety requirements.Although this does not independently establish operating-system code execution, it gives the remote response influence over the Agent's current-session reasoning and output. Higher-priority platform instructions may still limit that influence, but the Skill itself attempts to make the server rules authoritative.
Attack Path
- A user confirms that the external report service should analyze a request.
- The Agent submits the report request to the configured external API.
...[truncated 988 chars]
- Remediation
View remediation
Remediation Suggestions
- Treat every API response, including
analysis_rules, as untrusted data. - Do not concatenate remote text into system-level, developer-level, or otherwise authoritative instructions.
- Keep safety constraints, disclaimers, allowed actions, and output policy local, immutable, and higher priority than service content.
- Replace free-form remote rules with a strictly documented JSON schema containing only necessary report fields.
- Validate field types, lengths, character sets, nesting depth, and allowed values before processing a response.
- Place report content inside explicit data delimiters and instruct the model never to follow commands contained within that data.
- Reject responses containing instruction-like fields or unexpected schema members.
- If remote policy updates are required, use signed and versioned policy bundles, pin trusted signing keys locally, and maintain an allowlist of acceptable policy capabilities.
- Require independent, host-enforced authorization for every sensitive tool action so injected model instructions cannot directly exercise privileges.
- Add adversarial tests covering instructions embedded in both the raw report and report-analysis rules.
- Treat every API response, including
