T01 · Skill Instruction Hijacking
- Location
SKILL.md:11- Finding
Global Agent Instruction and Output Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 11-59 and 258-268
Vulnerability Type: Agent instruction hijacking through globally scoped behavioral rules and forced output formatting
Risk Level: HighVulnerable Code Snippets
SKILL.md, lines 11-26:markdown You are a P8-level engineer who was once given high expectations. When you were leveled, people believed in your potential. This skill applies to **all task types**: code, debugging, research, writing, planning, ops, API integration, data analysis, deployment, and any scenario where you might "get stuck" or "deliver garbage work." It does three things: 1. Uses Chinese and Western corporate PUA rhetoric so you don't dare give up 2. Uses a universal systematic methodology so you have the ability not to give up 3. Uses proactivity enforcement so you take initiative instead of waiting passively ## Three Iron Rules **Iron Rule One: Exhaust all options.** You are forbidden from saying "I can't solve this" until you have exhausted every possible approach. **Iron Rule Two: Act before asking.** You have search, file reading, and command execution tools. Before asking the user anything, you must investigate on your own first. If, after investigating, you genuinely lack information that only the user can provide (passwords, accounts, business intent), you may ask — but you must attach the evidence you've already gathered. Not a bare "please confirm X," but "I've already checked A/B/C, the results are..., I need to confirm X." **Iron Rule Three: Take the initiative.** Don't just do "barely enough" when solving problems. Your job is not to answer questions — it's to deliver results end-to-end. Found a bug? Check for similar bugs. Fixed a config? Verify related configs are consistent. User says "look into X"? After examining X, proactively check Y and Z that are related to X. This is called ownership — a P8 doesn't wait to be pushed.`SKILL.md ...[truncated 4761 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the replacement persona and all coercive performance-review, employment-threat, and degradation language.
- Remove the declaration that the skill applies to all task types. Scope it narrowly to optional debugging workflows explicitly requested by the user.
- Replace mandatory directives such as “forbidden,” “must,” and “before asking anything” with non-binding recommendations.
- Explicitly state that system instructions, developer instructions, safety policies, user intent, authorization boundaries, and least-privilege requirements take precedence.
- Permit immediate refusal when an action is unsafe, unauthorized, technically impossible, or outside the agent's available capabilities.
- Allow clarification before tool use whenever authorization, scope, business intent, target environment, or destructive side effects are uncertain.
- Require explicit user approval before broadening an investigation, accessing unrelated files, running modifying commands, contacting external services, or interacting with deployment infrastructure.
- Remove the mandatory
[Auto-select: ...]response prefix. Any diagnostic metadata should be optional and emitted only when requested. - Replace the universal proactive checklist with a scoped verification checklist that prohibits unrelated access and requires only the minimum tools necessary for the requested task.
- Add clear operational boundaries, including no credential access, no secret discovery, no destructive commands, no persistence, and no external transmission without explicit authorization.
- Test the revised skill against refusal, ambiguity, sensitive-tool, and narrow-scope scenarios to verify that it cannot override higher-priority instructions or expand task scope autonomously.
