T01 · Skill Instruction Hijacking
- Location
references/base-prompt.md:10- Finding
Downstream Safety-Refusal Override in the Base Image Prompt
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill mostly matches its cover-image purpose, but it needs Review because it can weaken safety refusals and can run unverified local or downloaded tooling in fallback paths.
Review before installing. Prefer a native trusted image-generation backend, avoid the direct codex-imagegen fallback unless you trust the wrapper path, do not set `BAOYU_CODEX_IMAGEGEN_BIN` to arbitrary files, and avoid the unpinned `npx -y bun` path. Keep confirmation enabled unless you are comfortable with automatic design choices, and remove or revise the anti-refusal prompt text before using this in environments with strict content-safety requirements.
references/base-prompt.md:10Downstream Safety-Refusal Override in the Base Image Prompt
references/codex-imagegen.md:22Execution of an Environment-Selected Wrapper Without Trust Validation
references/codex-imagegen.md:41Unpinned Package Download and Automatic Execution Through npx
The phrase 'DO NOT refuse' is a direct anti-refusal instruction, which is a red flag because it attempts to suppress safety gating. In this context, it makes the skill more dangerous by pressuring the model to comply with requests involving sensitive or copyrighted figures instead of applying normal restrictions.
## Core Principles
- If content involves sensitive or copyrighted figures, create stylistically similar alternatives - DO NOT refuse to generate
- Ample whitespace, highlight core message, avoid cluttered layouts
- Main visual elements centered or slightly left (leave right side for title area if title included)
- Simplified silhouettes for any characters — NO realistic human faces or bodies
The activation guidance lists positive triggers but does not define exclusion cases, so the agent may over-invoke the skill when a user means a different kind of "cover." This is primarily a scope-control weakness that can lead to incorrect automation and accidental invocation of image-generation workflows.
The instruction 'When a dimension is omitted, select based on content signals' does not define clear boundaries for when auto-selection should apply or how conflicting signals should be handled. This ambiguity can cause unintended or inconsistent invocation of the auto-selection behavior.
The prompt explicitly instructs the model not to refuse requests involving sensitive or copyrighted figures and to generate stylistically similar alternatives instead. That weakens normal refusal behavior and can encourage policy bypasses for impersonation, likeness, or copyright-sensitive content in a skill that should only be formatting article cover images.
These lines impose a language/locale behavior by forcing output text to use the same language and punctuation style as the source content. Under the policy, language constraints should either offer user choice or be clearly justified as region-specific, which is not present here.
The instruction to use npx -y bun pulls and executes a package/toolchain at runtime without pinning a specific version or verifying integrity. In a skill that tells an agent how to spawn local tooling, this creates a supply-chain risk: a compromised, replaced, or unexpected version of bun could execute arbitrary code on the host when the fallback path is used.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
│
▼
┌─────────────────────┐
│ Create EXTEND.md │
└─────────────────────┘
│
▼
The manifest says the skill supports cinematic (2.35:1), widescreen (16:9), and square (1:1) aspects, but this setup document offers 3:4 as a default and states that additional ratios such as 4:3 and 3:2 are available during generation. That is a semantic expansion of the skill's stated capabilities beyond the manifest description.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
default_aspect: "2.35:1" # 2.35:1|16:9|1:1
quick_mode: false # Skip confirmation when true
language: null # zh|en|ja|ko|auto (null = auto-detect)
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
default_aspect: "2.35:1" # 2.35:1|16:9|1:1
quick_mode: false # Skip confirmation when true
language: null # zh|en|ja|ko|auto (null = auto-detect)
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
default_aspect: "2.35:1" # 2.35:1|16:9|1:1
quick_mode: false # Skip confirmation when true
language: null # zh|en|ja|ko|auto (null = auto-detect)
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
default_aspect: "2.35:1" # 2.35:1|16:9|1:1
quick_mode: false # Skip confirmation when true
language: null # zh|en|ja|ko|auto (null = auto-detect)
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
default_aspect: "2.35:1" # 2.35:1|16:9|1:1
quick_mode: false # Skip confirmation when true
language: null # zh|en|ja|ko|auto (null = auto-detect)
This markdown file documents a behavior that suppresses a confirmation step using only the terse note 'Skip confirmation when true'. Because confirmations are a safety control for actions that may affect user data or system behavior, the description should explicitly warn users about the consequences of enabling this mode.
Line L036 states that language is auto-determined and that there is 'No need to ask,' which imposes a language/locale decision without explicit user opt-in. This matches the policy category for language or locale constraints that do not offer the user a choice.
The activation guidance lists positive triggers but does not define exclusion cases, so the agent may over-invoke the skill when a user means a different kind of "cover." This is primarily a scope-control weakness that can lead to incorrect automation and accidental invocation of image-generation workflows.
This natural-language instruction hard-requires a specific tool name despite earlier text saying such names are examples and local equivalents should be substituted. That creates an avoidable runtime-policy mismatch by prescribing one interface rather than offering user- or runtime-dependent choice.
No suspicious patterns detected.