T01 · Skill Instruction Hijacking
- Location
references/style-guide.md:57- Finding
Unconditional Non-Refusal Directive for Sensitive or Copyright-Related Content
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This knowledge-card image skill mostly matches its purpose, but it includes under-disclosed public sharing, persistent style changes, and a direct instruction not to refuse sensitive or copyright-related image requests.
Install only if you are comfortable with a Chinese-first image workflow that may upload outputs to a public CDN. Do not use it for confidential, personal, regulated, or rights-sensitive material unless the publisher removes the non-refusal rule, adds consent and access controls for uploads, and constrains user-added styles to reviewed visual-only data.
references/style-guide.md:57Unconditional Non-Refusal Directive for Sensitive or Copyright-Related Content
references/style-guide.md:64Persistent Storage of Untrusted User-Provided Style Instructions
assets/prompts/base_prompt.md:75Generated Content Is Automatically Published Through a Public CDN URL
The phrase instructing the model to 'not refuse generation' directly conflicts with content-safety controls. This creates policy pressure to comply with requests involving sensitive persons or copyrighted material, increasing the chance of unsafe generation and weakening higher-level safeguards through prompt steering.
Telling the model to draw a 'similar replacement' for restricted subjects is a semantic workaround: it preserves the user's disallowed intent while changing surface form. That makes the file more dangerous because it operationalizes evasion inside a reusable default style guide, potentially affecting many downstream generations.
The activation phrases are very broad and map to common user intents like making a card or poster, which can cause the skill to trigger in situations the user did not intend. This increases the chance of context hijacking or unexpected execution of the skill workflow, especially because the skill then instructs the agent to perform retrieval and image-generation steps automatically.
The skill is written to operate in Chinese and mandates fixed phrasing and workflow without offering a language choice, which can cause mismatches with the user's language and reduce informed consent around what the skill is doing. While this is not a direct code-execution issue, it can lead to user confusion, mistaken confirmations, or unintended processing when the user cannot clearly understand the interaction.
The template hardcodes the badge text as '{DATE} · 知识卡片', which imposes a Chinese-language output element regardless of user preference. The file does not indicate that this is optional, user-selected, or required for a documented region-specific purpose.
The prompt workflow explicitly includes uploading generated images to a public CDN and then sending the resulting URL, but there is no indication of user notice, consent, or any warning about public exposure. If prompts or rendered images contain sensitive business, personal, or internal information, this can cause unintended data disclosure through a publicly accessible link.
The template title and instruction text are written entirely in Chinese and direct the agent to organize and send content using this template before user confirmation. This imposes a specific language/locale on generated user-facing content without indicating that the user can choose another language.
The style guide explicitly instructs the agent to create a similar replacement for sensitive persons or copyrighted content instead of refusing. In a visual-style definition file, this is an unnecessary behavioral override that can help users evade safety and intellectual-property restrictions by reframing disallowed requests as lookalikes.
This file contains natural-language instructions and headings predominantly in Chinese, which can impose a language expectation on users or maintainers without an explicit opt-in. The policy requires flagging language or locale constraints when the skill does not offer a user choice or document a justified regional restriction.
No suspicious patterns detected.