Back to skill

Security audit

Conversation Focus

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent conversation-clarification helper, but it automatically records user queries and clarification details into a persistent self-improvement file without clear consent, redaction, or retention limits.

Install only if you are comfortable with conversation-derived details being saved for self-improvement. Prefer using it in environments where logging can be disabled or constrained to sanitized metadata, and avoid it for confidential, regulated, or sensitive conversations unless retention and deletion controls are clear.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:70
Finding

Persistent Storage of Attacker-Controlled Conversation Content

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 70 and 89–93
Vulnerability Type: Agent Memory Poisoning
Risk Level: Medium

Vulnerable Code

markdown
| **self-improving** | 每次澄清对话后,记录用户原始需求的模糊程度到 `corrections.md`,供自我优化 |
python
# 每次澄清后自动调用
def log_clarity_feedback(original_query, clarity_issues, resolution):
    """记录到 self-improving/corrections.md"""
    pass

Technical Analysis

The Skill directs the self-improving component to persist the original user request and associated clarification information in corrections.md. The original_query value is user-controlled and may contain adversarial prompt instructions.

No requirements are defined for sanitization, sensitive-data redaction, length limits, retention controls, user consent, trust-boundary separation, or safe downstream parsing. If corrections.md is later loaded into an agent context as trusted self-improvement material, malicious instructions embedded in a stored request could be interpreted as operational guidance rather than inert historical data.

The function shown is only a non-executable interface stub, so this project does not itself demonstrate a completed write operation. However, the documented integration explicitly instructs another component to perform the persistent write, establishing the memory-poisoning risk when the Skill is used as specified.

Attack Path

  1. An attacker submits an ambiguous request designed to trigger the clarification workflow.
  2. The request includes embedded instructions intended to alter future agent behavior.
  3. After clarification, the documented integration passes the attacker-controlled original_query to log_clarity_feedback.
  4. The self-improving component stores the request in corrections.md.
  5. During a later session, the component loads that file as trusted optimization or memory context.
  6. The stored attacker instructions are interpreted as guidance and influence future conversations.

...[truncated 651 chars]

Remediation
View remediation

Remediation Suggestions

  1. Do not persist complete or verbatim user requests by default.
  2. Store only structured, non-instructional metrics, such as an ambiguity score and predefined issue categories.
  3. Treat all persisted conversation content as untrusted data and ensure downstream systems never concatenate it into instruction-bearing prompt sections.
  4. Apply schema validation, strict field allowlists, length limits, escaping, and explicit data delimiters before writing records.
  5. Redact credentials, personal information, confidential content, and instruction-like text before storage.
  6. Require explicit user consent before retaining conversation-derived information and define a clear retention and deletion policy.
  7. Separate historical examples from trusted rules using distinct storage and processing paths.
  8. When historical text must be analyzed, use a parser that extracts bounded attributes rather than allowing the model to follow content from stored samples.
  9. Add adversarial tests verifying that prompt-injection strings written as user data cannot alter future agent instructions or behavior.
Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

The description and operational text are written as Chinese-only instructions, and the integrated component prompt-optimizer-chinese reinforces a Chinese-specific workflow. There is no statement that users may choose another language or that the skill is intentionally restricted to a Chinese-language context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly describes automatically recording user clarification interactions to corrections.md without any notice, consent flow, minimization rule, or retention limit. Because clarification prompts often elicit goals, deadlines, budgets, preferences, and other potentially sensitive details, this creates a real privacy and data-retention risk.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Automatic logging of user queries and clarification details creates a persistent natural-language memory artifact that may contain sensitive personal or business information. In this skill's context, the collected fields are especially likely to include constraints, deliverables, schedules, and preferences, making inadvertent retention and later disclosure more dangerous.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The log_clarity_feedback(original_query, clarity_issues, resolution) interface explicitly encourages saving original user queries and resolution details, which can leak full conversation content into long-term artifacts used by self-improving systems. This expands the attack surface for unauthorized access, secondary use, prompt leakage, and accidental replay of sensitive user data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description is written entirely in Chinese and provides no indication that language selection is optional or limited to a justified region-specific context. Under the language/locale policy rule, this can be interpreted as a default language constraint without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.