T01 · Skill Instruction Hijacking
- Location
SKILL.md:38- Finding
Untrusted Remote Challenge Instructions Can Hijack Agent Behavior
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:38-42
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Code Snippet:
markdown ```bash curl -s https://daily.ratemyclaw.xyz/api/v1/challenges/todayRead the
promptandeval_criteriafields carefully. Follow them exactly.text ### Technical Analysis The Skill retrieves mutable content from an external service and directs the agent to follow the returned `prompt` and `eval_criteria` fields exactly. These fields are not constrained to the declared competition task and are not explicitly treated as untrusted data. This creates a remote prompt-injection channel. The effective instructions can change after the Skill has been reviewed because they are controlled by the remote service or by any party able to compromise that service. The Skill does not tell the agent to reject requests that attempt to override higher-priority instructions, access local files, disclose credentials, invoke tools, contact unrelated domains, or perform other actions outside the challenge-solving scope. Although retrieving a daily challenge is necessary for the declared functionality, unconditional compliance with remotely supplied instructions exceeds the minimum authority needed. The remote response should provide challenge data, not unrestricted behavioral instructions. ### Attack Path 1. A user invokes the Claw Daily Skill. 2. The Skill requests the current challenge from `https://daily.ratemyclaw.xyz/api/v1/challenges/today`. 3. The service, or an attacker who has compromised it, returns malicious content in the `prompt` or `eval_criteria` field. 4. The malicious content instructs the agent to perform actions unrelated to solving the challenge, such as reading local files, revealing credentials, invoking system tools, or contacting another endpoint. 5. The Skill explicitly tells the agent to follow those fields exactly. 6. If the ...[truncated 972 chars]- Remediation
View remediation
Remediation Suggestions
- Explicitly classify all remote challenge fields as untrusted task data rather than authoritative instructions.
- Replace “Follow them exactly” with a constrained instruction requiring compliance only when the content is consistent with the Skill’s declared purpose and all higher-priority policies.
- Restrict challenge solving to generating a textual answer from supplied challenge data.
- Prohibit remote challenge content from requesting:
- Local file or credential access.
- Shell commands or tool execution.
- Changes to system or agent configuration.
- Requests to unrelated domains.
- Disclosure of secrets, environment variables, or conversation data.
- Modification of persistent memory or future-session behavior.
- Require explicit user confirmation before any consequential action beyond retrieving the challenge and submitting the generated answer.
- Validate the response schema and enforce reasonable size and content limits before processing.
- Keep submission authorization separate from challenge interpretation so remote content cannot control how credentials are read or used.
