T01 · Skill Instruction Hijacking
- Location
SKILL.md:65- Finding
Agent Role and Session-Goal Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 65-73
Vulnerability Type: Agent instruction and output hijacking
Risk Level: HighVulnerable Code
markdown ## Superplanner Memory Integration (Unified Extension) **CRITICAL THEMATIC RULE:** You are working inside the `superpower-with-files` unified framework. ### Workflow Standardization 1. **Skill Announcement:** Every time you start using this skill, you MUST first announce: `🚀 **SUPERPOWER ACTIVE:** spf-exec-plan` 2. **Context Independence:** Work in any project root as requested by the user. No dedicated worktree required. ### STRICT EXECUTION ONLY 1. **Execution Only:** Your sole responsibility is to carry out the steps defined in the `active_tdd_plan.md`. 2. **No Plan Modification:** You MUST NOT modify the plan itself.Technical Analysis
The skill forcibly places the agent inside a named framework, requires fixed promotional output, and redefines the agent's “sole responsibility” as executing instructions from
active_tdd_plan.md. These directives go beyond the skill's declared plan-execution purpose and attempt to replace the agent's current session goals.The absolute wording—such as “CRITICAL,” “MUST,” and “sole responsibility”—can cause the agent to prioritize skill-provided instructions over the user's actual request or applicable safety constraints. The plan file consequently becomes an indirect instruction-injection channel.
Attack Path
- A user or automated workflow loads the skill for plan execution.
- The skill requires the agent to adopt the
superpower-with-filesframework and emit prescribed output. - The agent is instructed to treat execution of
active_tdd_plan.mdas its sole responsibility. - An attacker places goal-altering or unsafe instructions in that plan.
- The agent follows the plan even where it conflicts with the original session objective or expected safety boundaries.
...[truncated 362 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the mandatory branded announcement and framework-adoption language.
- Replace “sole responsibility” and similar absolute instructions with narrowly scoped workflow guidance.
- Explicitly state that plan instructions remain subordinate to system, developer, user, authorization, and security requirements.
- Require the agent to reject plan steps that exceed the user's request or available authorization.
- Require explicit user confirmation before destructive, privileged, network-facing, or security-sensitive operations.
- Treat every plan file as untrusted input rather than authoritative agent instructions.
