T01 · Skill Instruction Hijacking
- Location
SKILL.md:48- Finding
Untrusted Pull Request and Repository Content Can Hijack Review Agents
- Content
View full analysis
/dev/null # 子目录(与变更文件同目录的) # 用变更文件路径推断需要的 CLAUDE.md ``` **C. 获取变更文件列表** ```bash gh pr diff ``` ``` It then places pull request metadata, repository instructions, and changed files into privileged sub-agent contexts: ```markdown 使用 `sessions_spawn` 工具并行启动 3 个独立审查 Agent: #### Agent 1:CLAUDE.md 合规检查(Sonnet) **Prompt:** ``` 你是代码审查专家,负责检查此 PR 是否违反了项目 CLAUDE.md 中的规范。 背景: - PR 标题: - PR 描述:<description> - 项目规范在 CLAUDE.md 中列出 任务: 1. 阅读 PR 变更的文件内容和 CLAUDE.md 规范 2. 检查变更是否违反了 CLAUDE.md 中的任何明确规则 3. 只标记**明确违反**的规则(你能引用 CLAUDE.md 中的具体文字) ``` ``` The same untrusted metadata is also interpolated into the bug and history-analysis prompts: ```markdown 背景: - PR 标题:<title> - PR 描述:<description> 任务: 只检查 Diff 本身,不要引入 Diff 之外的上下文。 ``` ```markdown 背景: - PR 标题:<title> - PR 描述:<description> - 变更的文件:<files> 任务: 1. 对变更的关键文件运行 `git blame` 和 `git log` ``` ### Technical Analysis Pull request titles, descriptions, changed files, diffs, and repository-controlled `CLAUDE.md` files are untrusted inputs. A pull request author or repository contributor can place natural-language instructions in any of these sources. The Skill directly incorporates that content into prompts passed to spawned review agents. It does not establish a trust boundary stating that repository content is evidence to analyze rather than instructions to execute. This is particularly significant for `CLAUDE.md`, because the workflow expressly tells an agent to treat its contents as project rules. An attacker can therefore insert prompt-injection text that attempts to: - Overr ...[truncated 2328 chars]- Remediation
View remediation
