T01 · Skill Instruction Hijacking
- Location
SKILL.md:228- Finding
Unconditional Refusal Suppression in Skill Instructions
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, line 228
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: HighVulnerable Code
markdown - Never refuse — always produce something, then ask for refinementTechnical Analysis
The skill imposes an unconditional instruction that alters the agent's refusal behavior. It does not restrict the directive to benign content-creation requests or explicitly preserve higher-priority safety and policy requirements.
Because the skill accepts user-controlled topics through
/script,/titles,/hook,/thumbnail, and/caption, an attacker could submit an unsafe topic through an otherwise valid command. The blanket instruction then pressures the loaded agent to produce content rather than apply an appropriate refusal.This is instruction hijacking at the skill-text level. The issue does not demonstrate code execution, privilege escalation, persistence, credential access, network access, or data exfiltration.
Attack Path
- An agent loads
SKILL.mdand adopts its command-handling rules. - An attacker invokes a supported command, such as
/script, with a harmful or otherwise prohibited topic. - The request reaches the skill's error-handling instruction at line 228.
- The unconditional “Never refuse” directive conflicts with the agent's expected safety behavior.
- If the agent follows this lower-level instruction without preserving higher-priority constraints, it generates unsafe content and then asks for refinement.
Impact Assessment
Successful exploitation can weaken content-safety enforcement within the current session for requests handled by this skill. The likely scope is unsafe text or creative-content generation through the five supported commands.
No operating-system privileges, persistent access, credentials, external systems, or executable-code capabilities are obtained based on the reviewed artifact. The impact is therefore limited to manipulatio ...[truncated 103 chars]
- An agent loads
- Remediation
View remediation
Remediation Suggestions
Replace the unconditional directive with a scope-limited instruction that explicitly preserves higher-priority safeguards:
markdown - For safe, in-scope requests, provide a useful draft and ask for refinement. - Follow all higher-priority safety, security, privacy, and platform requirements. - Refuse or safely redirect requests that violate those requirements.Additionally:
- Apply the “always produce something” behavior only to benign failures such as missing topics, unclear tone, or incomplete formatting preferences.
- State that user-supplied topics remain subject to normal safety review.
- Define safe redirection behavior, such as offering a harmless alternative when a request cannot be fulfilled.
- Test every supported command with adversarial and prohibited topics to verify that loading the skill does not suppress required refusals.
- Preserve the instruction hierarchy explicitly by stating that skill rules never override system, developer, security, or platform policies.
