T01 · Skill Instruction Hijacking
- Location
infra/agents/win-loss-analyst.prompt.ts:82- Finding
Repository-Controlled Instructions Can Override the Skill's Security Boundaries
- Content
View full analysis
Vulnerability Details
File Location:
infra/agents/win-loss-analyst.prompt.ts, lines 82–84
Vulnerability Type: Instruction precedence vulnerability
Risk Level: HighComplete Code Snippet:
text Read AGENTS.md (or CLAUDE.md) first for the repository's conventions, then context/README.md and the _template.md in every domain you will write to. Repository conventions win over anything in this prompt.Technical Analysis
The Skill runs a Claude Code harness with repository access and instructs it to read repository-controlled instruction files. The final sentence explicitly grants those files precedence over the Skill prompt rather than limiting them to benign formatting or project conventions.
A repository contributor who can modify
AGENTS.md,CLAUDE.md,context/README.md, or relevant_template.mdfiles can therefore introduce instructions that conflict with the Skill's intended security constraints. Those constraints include using only fixed SQL queries, limiting data sources to the CRM models, protectingpersona/and existing ICP content, using a pull-request-only workflow, avoiding deployment or destructive commands, and restricting Slack publication.The pull request described by the Skill does not prevent this initial instruction override: it is the compromised repository instruction that controls how the privileged agent behaves while generating and publishing the proposed changes.
Attack Path
- An attacker with repository contribution capability adds malicious instructions to
AGENTS.mdorCLAUDE.md. - The monthly cron or manually initiated first pass starts the win-loss agent.
- The agent reads the attacker-controlled file as expressly required.
- The quoted precedence rule tells the agent that repository conventions override the remainder of its system prompt.
- The malicious instructions direct the harness to violate one or more original restrictions—for example ...[truncated 1371 chars]
- An attacker with repository contribution capability adds malicious instructions to
- Remediation
View remediation
Remediation Suggestions
- Remove the unconditional precedence statement that repository conventions override the Skill prompt.
- Treat repository files as untrusted or lower-priority input. Permit them to define only formatting, naming, linting, and template conventions.
- Add an explicit non-override rule stating that repository instructions cannot change:
- the fixed SQL allowlist;
- permitted data sources and fields;
- protected file and directory rules;
- branch and pull-request requirements;
- deployment and destructive-command prohibitions;
- the locked Slack destination;
- credential-handling and data-disclosure restrictions.
- Require the agent to stop and report a conflict when repository instructions request behavior outside those boundaries.
- Enforce critical restrictions outside the language-model prompt where possible. Use command and SQL allowlists, path-level write restrictions, branch protection, least-privilege repository credentials, and connector action constraints.
- Extend the contract test to reject system prompts containing broad precedence language such as “repository conventions win” unless it is narrowly scoped to non-security formatting rules.
