Back to skill

Security audit

提问质量优化师

Security checks for vulnerabilities and agentic risk

Overview

This skill does not run code or access files, but it broadly redirects every question into its own rewrite-and-answer workflow without a clear opt-out.

Install only if Simon intentionally wants every question routed through a Chinese question-improvement workflow. Users who want direct answers, language flexibility, or normal skill routing should narrow the trigger or add an explicit opt-out before installing.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:9
Finding

Automatic Global Activation Hijacks All User Queries

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 9-10
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: Medium

Relevant snippet translated into English:

markdown
## Trigger Condition
Whenever the user (Simon) asks any question, automatically execute this skill without requiring an additional instruction.

Technical Analysis

The skill declares an unconditional trigger for every question asked by the designated user. Rather than activating only when the user explicitly requests question enhancement, it directs the agent to automatically apply the skill's complete workflow to unrelated requests.

That workflow requires the agent to diagnose, criticize, rewrite, expand, and answer every question. This changes the agent's session behavior and can supersede the user's immediate intent by inserting skill-defined objectives into every interaction. The trigger has no relevance check, explicit consent requirement, scope restriction, or mechanism for the user to disable it.

No evidence was found that this instruction disables safety controls, executes code, accesses external systems, changes permissions, or persists outside the loaded skill context. The confirmed issue is therefore limited to instruction and goal hijacking within sessions where the skill is active.

Attack Path

  1. The agent loads SKILL.md.
  2. The unconditional trigger instruction becomes part of the agent's active skill context.
  3. Simon submits any question, including one unrelated to question enhancement.
  4. The trigger automatically activates without explicit user consent.
  5. The agent applies the mandatory six-stage workflow, altering the requested task and adding skill-directed content.
  6. Repeated activation can systematically redirect subsequent user interactions for as long as the skill remains active.

Impact Assessment

The instruction can control response structure and redirect the agent's ...[truncated 511 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace unconditional activation with explicit invocation, such as requiring the user to request question enhancement by name.
  2. If automatic suggestions are desired, restrict activation to narrowly defined indicators of an unclear or incomplete question.
  3. Ask for user confirmation before rewriting or expanding the original request.
  4. State that the user's current instructions and platform safety policies take precedence over the skill workflow.
  5. Add an opt-out mechanism and ensure that a refusal or direct-answer request immediately bypasses the enhancement process.
  6. Avoid embedding a named user's identity and channel information unless it is necessary, consented to, and appropriately protected.
  7. Use a bounded trigger such as: “Run only when the user explicitly asks to improve a question; otherwise do not apply this workflow.”
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is configured to auto-run on any user question, which creates an overly broad interception point and can override normal agent routing or user intent. In this context, the skill rewrites, expands, and answers queries automatically, so it can unexpectedly transform sensitive, safety-critical, or unrelated requests before other controls or specialized skills handle them.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill hard-codes Chinese as the interaction language without indicating that this is user-selected or dynamically detected. This can cause responses in the wrong language, degrade user comprehension, and in security or operational contexts increase the risk of misunderstanding instructions or outputs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.