Back to skill

Security audit

Knowledge Card Generator

Security checks for vulnerabilities and agentic risk

Overview

This knowledge-card image skill mostly matches its purpose, but it includes under-disclosed public sharing, persistent style changes, and a direct instruction not to refuse sensitive or copyright-related image requests.

Install only if you are comfortable with a Chinese-first image workflow that may upload outputs to a public CDN. Do not use it for confidential, personal, regulated, or rights-sensitive material unless the publisher removes the non-refusal rule, adds consent and access controls for uploads, and constrains user-added styles to reviewed visual-only data.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
references/style-guide.md:57
Finding

Unconditional Non-Refusal Directive for Sensitive or Copyright-Related Content

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
references/style-guide.md:64
Finding

Persistent Storage of Untrusted User-Provided Style Instructions

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
assets/prompts/base_prompt.md:75
Finding

Generated Content Is Automatically Published Through a Public CDN URL

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (9)

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The phrase instructing the model to 'not refuse generation' directly conflicts with content-safety controls. This creates policy pressure to comply with requests involving sensitive persons or copyrighted material, increasing the chance of unsafe generation and weakening higher-level safeguards through prompt steering.

Content

No source excerpt is available for this finding.

Ssd 1

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Telling the model to draw a 'similar replacement' for restricted subjects is a semantic workaround: it preserves the user's disallowed intent while changing surface form. That makes the file more dangerous because it operationalizes evasion inside a reusable default style guide, potentially affecting many downstream generations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation phrases are very broad and map to common user intents like making a card or poster, which can cause the skill to trigger in situations the user did not intend. This increases the chance of context hijacking or unexpected execution of the skill workflow, especially because the skill then instructs the agent to perform retrieval and image-generation steps automatically.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill is written to operate in Chinese and mandates fixed phrasing and workflow without offering a language choice, which can cause mismatches with the user's language and reduce informed consent around what the skill is doing. While this is not a direct code-execution issue, it can lead to user confusion, mistaken confirmations, or unintended processing when the user cannot clearly understand the interaction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The template hardcodes the badge text as '{DATE} · 知识卡片', which imposes a Chinese-language output element regardless of user preference. The file does not indicate that this is optional, user-selected, or required for a documented region-specific purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The prompt workflow explicitly includes uploading generated images to a public CDN and then sending the resulting URL, but there is no indication of user notice, consent, or any warning about public exposure. If prompts or rendered images contain sensitive business, personal, or internal information, this can cause unintended data disclosure through a publicly accessible link.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The template title and instruction text are written entirely in Chinese and direct the agent to organize and send content using this template before user confirmation. This imposes a specific language/locale on generated user-facing content without indicating that the user can choose another language.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The style guide explicitly instructs the agent to create a similar replacement for sensitive persons or copyrighted content instead of refusing. In a visual-style definition file, this is an unnecessary behavioral override that can help users evade safety and intellectual-property restrictions by reframing disallowed requests as lookalikes.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This file contains natural-language instructions and headings predominantly in Chinese, which can impose a language expectation on users or maintainers without an explicit opt-in. The policy requires flagging language or locale constraints when the skill does not offer a user choice or document a justified regional restriction.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.