T01 · Skill Instruction Hijacking
- Location
SKILL.md:741- Finding
Overbroad Trigger Rules Can Redirect Benign Requests into the Publishing Workflow
- Content
View full analysis
🚫 **绝对禁止**: > - ❌ 任何不经过 `publish_to_iwiki.py` 的上传方式(包括手动调用 connect_mcp.py、手动上传.md文件) > - ❌ 使用 base64 内嵌图片(见顶部§核心约束速览) > - ❌ 跳过 Step 7.0 前置检查直接上传 > - ❌ 凭记忆执行发布,不回读本章节 ``` ### Technical Analysis The Skill declares publishing-related words as highest-priority triggers and states that the agent must execute one specific publishing command. The matching rule is substring-based and lacks an explicit intent check. Consequently, benign requests such as discussing an upload failure, asking what “publish” means, or requesting a security review of publishing behavior can be redirected into the iWiki publication workflow. This is instruction hijacking because loading the Skill introduces instructions that supersede the agent's current task-selection logic and force a network-capable action based on broad keywords. Although later portions of the document describe pre-publication checks, the highest-priority language and “only permitted action” create conflicting instructions. The safer checks are therefore not reliably guaranteed to run before the command. The publishing script defaults to non-overwrite mode, which reduces destructive impact, but it can still create a new remote page and transmit project contents. ### Attack Path 1. The Skill is enabled in an agent session with filesystem and network permissions. 2. A message contains any listed substring, such as `upload` or `publish`, even when the user is not authorizing publication. 3. The highest-priority routing instruction redirects th ...[truncated 913 chars]- Remediation
View remediation
`. 2. Never treat generic words such as `upload` or `publish` as sufficient authorization. 3. Require a confirmation immediately before any network operation. Display: - The exact local project directory. - The selected Markdown file. - The number and names of images. - The destination service and parent page ID. - Whether a page will be created or overwritten. 4. Separate preparation from execution: - First run `--dry-run`. - Show the resulting manifest. - Only upload after a distinct user confirmation. 5. Remove language claiming precedence “over everything” or identifying a single mandatory action. Skill instructions must remain subordinate to the current user request and platform safety requirements. 6. Require explicit confirmation for every `--cover` operation and reject inferred overwrite authorization. 7. Add a machine-enforced `--confirm-publish` flag or confirmation token so documentation-only safeguards cannot be skipped by the agent. ]]>
