Back to skill

Security audit

Xiaohongshu Founder Growth Writer

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent copywriting skill for Xiaohongshu posts with no code execution, persistence, credential use, or hidden data access.

Installers should expect this skill to default to Chinese Xiaohongshu-style founder marketing copy. The main improvement would be clearer wording that normal safety boundaries still apply, but the reviewed artifacts do not show hidden behavior or high-impact access.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
### If the user gives very little input

Do not refuse.

Infer a reasonable account type, goal, and tone from the context, then proceed.
Confidence
88% confidence
Finding
The instruction 'Do not refuse' weakens the model's safety posture by biasing it against declining unsafe, policy-violating, or under-specified requests. While this skill is focused on copywriting and does not request privileged actions or tool use, anti-refusal phrasing can still encourage the assistant to comply when it should instead set boundaries or ask clarifying questions.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The text repeatedly instructs conversion or rewriting into Chinese/Xiaohongshu-specific content, such as 'Turn this English brand paragraph into a Chinese Xiaohongshu post.' This imposes a specific language/locale context in the skill examples without stating that users can choose another language or opt into that locale-specific behavior.

Natural-Language Policy Violations

Low
Confidence
94% confidence
Finding
The skill specifies 'Chinese, unless the user asks otherwise,' which sets a default language choice without explicit user opt-in. This is a natural-language locale policy concern because it assumes a language preference rather than first offering a choice.

Static analysis

No suspicious patterns detected.