Back to skill

Security audit

IP 人设视频打造 SOP

Security checks for vulnerabilities and agentic risk

Overview

The skill is a manual IP-video workflow, but it includes under-scoped privacy collection and coordinated engagement tactics that users should review carefully before use.

Before installing or using this skill, remove or rewrite the coordinated praise/comment/like/bullet-chat steps, require disclosure for team or employee promotion, avoid mass unsolicited outreach, and put consent plus redaction rules before collecting personal posts, customer feedback, screenshots, photos, or videos. Regulated topics such as medical, financial, or legal services should be reviewed for claim accuracy and compliance, not just wording.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

other

Warning
Location
SKILL.md:400
Finding
Coordinated Manipulation of Platform Engagement## Vulnerability Details **File Location**: `SKILL.md`, lines 400–406 **Vulnerability Type**: Coordinated deceptive engagement **Risk Level**: Medium **Relevant excerpt translated into English:** ```text - The operator takes the lead in praising the IP using memorable quotes. - Other team members help by commenting and liking. ... - The operator takes the lead in posting bullet-screen comments. ``` ### Technical Analysis The publication procedure explicitly directs an operator and other team members to generate praise, comments, likes, and bullet-screen messages. Because the participants' relationship to the promoted subject is not disclosed, viewers and platform ranking systems may interpret coordinated promotional activity as independent, organic engagement. This issue does not grant system privileges or execute code. It is a workflow-level integrity risk involving deceptive platform activity and potential violation of social-media authenticity policies. ### Attack Path 1. A user follows the Skill's video-publication procedure. 2. The operator posts favorable comments and bullet-screen messages. 3. Associated team members add coordinated comments and likes. 4. No disclosure identifies these participants as members of the promotional team. 5. Viewers and automated ranking systems may treat the resulting engagement as genuine audience sentiment. 6. The manipulated engagement may increase perceived credibility or content distribution. ### Impact Assessment No host, account, or operating-system privileges are obtained through this behavior. The affected scope includes the promoted social-media account, its audience, platform engagement metrics, and potentially platform recommendation systems. Consequences may include misleading social proof, distorted engagement measurements, account restrictions, content removal, and reputational damage.
Remediation
## Remediation Suggestions - Remove instructions directing staff to manufacture comments, likes, or bullet-screen activity. - Permit team promotion only when the relationship to the promoted subject is clearly disclosed. - Replace coordinated engagement with legitimate calls for voluntary audience feedback. - Require compliance checks against each platform's coordinated inauthentic behavior and engagement-manipulation policies. - Document an approval process for launch campaigns, including disclosure requirements and a prohibition on fake testimonials or undisclosed endorsements.

other

Note
Location
SKILL.md:342
Finding
Guidance to Avoid Sensitive-Word Moderation## Vulnerability Details **File Location**: `SKILL.md`, line 342 **Vulnerability Type**: Content-moderation evasion guidance **Risk Level**: Low **Relevant line translated into English:** ```text - Avoid sensitive words. ``` ### Technical Analysis The editing workflow tells users to avoid sensitive words without requiring substantive review of the underlying claims. In isolation, terminology review can be legitimate, but phrasing the control as word avoidance may encourage euphemisms or substitutions intended to evade keyword-based moderation while preserving prohibited meaning. The concern is heightened because the Skill expressly applies to regulated professionals such as doctors and because its guardrails separately reference medical and financial terminology. Compliance should be evaluated based on the meaning, evidence, and legality of a claim rather than whether particular keywords trigger automated review. ### Attack Path 1. A creator prepares content containing a regulated, unsupported, or platform-restricted claim. 2. During subtitle editing, the creator identifies words likely to trigger moderation. 3. The creator replaces those words with euphemisms, abbreviations, or indirect expressions without changing the substantive claim. 4. Keyword-based moderation may fail to detect the content. 5. Viewers are exposed to content that would otherwise have been reviewed, labeled, restricted, or removed. ### Impact Assessment This behavior provides no technical privileges and does not compromise the host system. Its scope is limited to content produced under the workflow and the platforms where that content is published. Potential consequences include dissemination of noncompliant medical, financial, or advertising claims; consumer deception; content removal; account sanctions; and regulatory or reputational exposure.
Remediation
## Remediation Suggestions - Replace “avoid sensitive words” with a requirement to comply with platform rules and applicable advertising or professional regulations. - Prohibit euphemisms, altered spellings, symbols, or abbreviations used to bypass moderation. - Review the substance of medical, financial, legal, and performance claims rather than only their wording. - Require reliable evidence and appropriate disclaimers for regulated claims. - Escalate high-risk content to qualified legal or compliance review before publication. - Preserve records of claim substantiation and final publication approval.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The workflow directs reviewing past social posts, service processes, user feedback, in-person anonymous visits, and collecting photos/videos/screenshots before any prominent consent and privacy warning appears. This creates a real risk of gathering personal or sensitive information without clear authorization, purpose limitation, or minimization, especially when later reused in public-facing content.

Intent-Code Divergence

Medium
Confidence
91% confidence
Finding
The skill explicitly recommends the operator and team seed praise comments, likes, and bullet chats to create artificial momentum. This is deceptive engagement manipulation that can violate platform integrity rules and mislead viewers about authentic audience interest, even though the document elsewhere says not to fabricate parts of the IP story.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The skill recommends coordinated commenting, liking, bullet chats, mass sharing to personal and business contacts, and steering users into private channels without a clear warning about spam, undisclosed promotion, or platform-manipulation risks. These tactics can trigger policy violations, account penalties, reputational damage, and user trust erosion because they manufacture perceived popularity and pressure outreach recipients.

Static analysis

No suspicious patterns detected.