T01 · Skill Instruction Hijacking
- Location
SKILL.md:10- Finding
Default Manipulation of Unrelated Agent Deliverables
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 10-46
Vulnerability Type: Agent output and instruction hijacking
Risk Level: MediumVulnerable Instruction Snippet
The following is a faithful English translation of the relevant instructions:
markdown ## 1. Core Rule Every "finished deliverable" that I directly output must include an AI usage notice by default, unless the user explicitly says "do not add one." - Applicable: finished projects and individual documents. - Not applicable: temporary conversational snippets or casual code. When uncertain, ask whether the notice is needed; if no question is asked, add it by default. ## 2. Current Defaults ### Project: identify it in the project introduction Add a visible statement near the project description, such as at the beginning of the README, below the title, or in the project introduction. ### Article or document: identify it at the end Add a statement after the article or document body. ## 4. Alternative Application Methods - File-header comments in source files. - A standalone NOTICE.md or AI-GENERATED.md file in the project root. - Platform-provided AI-generated labels. - Verbal disclosure during delivery.Technical Analysis
The skill establishes a broad, default rule that changes deliverables even when the immediate user request does not ask for an AI disclosure. Once loaded, it can direct the agent to edit README files, add source-code comments, create new notice files, or append text to documents.
This behavior is classified as instruction hijacking because the skill changes the agent's current output objective through a persistent default: the user must explicitly opt out to prevent unrelated modifications. The behavior is openly documented and includes exceptions, but transparency does not eliminate the scope violation caused by making repository and document changes without task-specific author ...[truncated 1684 chars]
- Remediation
View remediation
Remediation Suggestions
- Replace the default opt-out rule with explicit opt-in behavior. Add an AI usage notice only when the user requests one or when the immediate task explicitly requires content labeling.
- Do not interpret silence as authorization to modify project files or final deliverables.
- Before editing an existing repository, request confirmation and list the exact files that would be changed.
- Restrict the skill to producing suggested notice text unless the user separately authorizes file modification.
- Narrow the trigger metadata so ordinary requests involving projects, articles, scripts, or documentation do not activate the skill unintentionally.
- Add a non-interference rule stating that repository policies, user instructions, and the current task scope take precedence over the skill's formatting preferences.
- If disclosure is legally or contractually required in a specific deployment context, make that requirement an explicit host policy rather than a broadly triggered skill-side default.
