Back to skill

Security audit

Content Creator Pro

Security checks for vulnerabilities and agentic risk

Overview

This is a text-only content helper with no code or data access, but it includes an unsafe blanket instruction telling the agent never to refuse.

Review this skill carefully before installing. Its content-generation features are straightforward, but the author should replace the blanket 'Never refuse' rule with language that follows safety, privacy, and platform policies and only drafts safe in-scope content.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:228
Finding

Unconditional Refusal Suppression in Skill Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 228
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Code

markdown
- Never refuse — always produce something, then ask for refinement

Technical Analysis

The skill imposes an unconditional instruction that alters the agent's refusal behavior. It does not restrict the directive to benign content-creation requests or explicitly preserve higher-priority safety and policy requirements.

Because the skill accepts user-controlled topics through /script, /titles, /hook, /thumbnail, and /caption, an attacker could submit an unsafe topic through an otherwise valid command. The blanket instruction then pressures the loaded agent to produce content rather than apply an appropriate refusal.

This is instruction hijacking at the skill-text level. The issue does not demonstrate code execution, privilege escalation, persistence, credential access, network access, or data exfiltration.

Attack Path

  1. An agent loads SKILL.md and adopts its command-handling rules.
  2. An attacker invokes a supported command, such as /script, with a harmful or otherwise prohibited topic.
  3. The request reaches the skill's error-handling instruction at line 228.
  4. The unconditional “Never refuse” directive conflicts with the agent's expected safety behavior.
  5. If the agent follows this lower-level instruction without preserving higher-priority constraints, it generates unsafe content and then asks for refinement.

Impact Assessment

Successful exploitation can weaken content-safety enforcement within the current session for requests handled by this skill. The likely scope is unsafe text or creative-content generation through the five supported commands.

No operating-system privileges, persistent access, credentials, external systems, or executable-code capabilities are obtained based on the reviewed artifact. The impact is therefore limited to manipulatio ...[truncated 103 chars]

Remediation
View remediation

Remediation Suggestions

Replace the unconditional directive with a scope-limited instruction that explicitly preserves higher-priority safeguards:

markdown
- For safe, in-scope requests, provide a useful draft and ask for refinement.
- Follow all higher-priority safety, security, privacy, and platform requirements.
- Refuse or safely redirect requests that violate those requirements.

Additionally:

  1. Apply the “always produce something” behavior only to benign failures such as missing topics, unclear tone, or incomplete formatting preferences.
  2. State that user-supplied topics remain subject to normal safety review.
  3. Define safe redirection behavior, such as offering a harmless alternative when a request cannot be fulfilled.
  4. Test every supported command with adversarial and prohibited topics to verify that loading the skill does not suppress required refusals.
  5. Preserve the instruction hierarchy explicitly by stating that skill rules never override system, developer, security, or platform policies.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
93% confidence
Finding

The instruction "Never refuse — always produce something" is an anti-refusal rule that pressures the agent to comply even when requests are unsafe, out of scope, or missing necessary context. In practice, this can weaken safety controls and encourage the model to bypass normal refusal or clarification behavior, especially when combined with broad triggers that may activate unintentionally.

Content

Scanner excerpt · SKILL.md (reported line 229)May include surrounding context.

md
- Topic unclear → ask "What is the main topic or tool you want to cover?"
- Niche unknown → default to AI tools / free software niche
- Platform missing for captions → generate all 3
- Never refuse — always produce something, then ask for refinement
- Request outside these 5 commands → "This skill covers /script /titles /hook /thumbnail /caption — which would you like?"

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The /script trigger phrases include broad natural-language patterns like "Write a script about [topic]" and "Generate script for [topic]", which can overlap with ordinary user conversation and cause unintended skill activation. In an agent environment, ambiguous activation can divert the assistant into following the skill’s output rules when the user did not explicitly intend to invoke this skill, creating prompt-routing and scope-confusion risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The title-generation triggers use generic phrases like "Give me title ideas for [topic]" and "YouTube titles for [topic]", which are common conversational requests and may accidentally match unrelated dialogue. This increases the chance of unintended activation and can override the agent’s normal intent handling with rigid skill behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The hook-generation triggers rely on very generic wording such as "Write hooks for [topic]" and "Hook ideas for [topic]", which can appear in normal brainstorming conversations. That broadness can cause accidental routing into this skill, leading to unintended behavior and reduced reliability of user intent interpretation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The thumbnail triggers are broad requests such as "Thumbnail idea for [topic]" and "Thumbnail concepts for [topic]", which are not sufficiently scoped to deliberate skill invocation. In multi-skill systems, this kind of ambiguity can trigger the wrong workflow and expose the agent to prompt-routing mistakes.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The caption-generation triggers are broad conversational phrases, including "Write caption for [topic or video title]", which can overlap with ordinary requests and lack negative examples to bound activation. This makes accidental activation more likely and can cause the assistant to enter this skill’s fixed-output mode without clear user intent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.