Back to skill

Security audit

Twitter热点监控与推文生成

Security checks for vulnerabilities and agentic risk

Overview

This skill is a Twitter/X trend and tweet-drafting helper, but it over-expands ordinary tweet requests into recurring monitoring and mandatory engagement-growth guidance.

Review this before installing if you only want ordinary tweet drafting. It may run or imply hourly trend monitoring, produce Chinese-only drafts, imitate named creators' styles, and include mandatory engagement tactics such as replying to all comments and commenting under popular posts. Install only if you want that full workflow and have separate controls for any scheduled execution or social-account actions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:3
Finding

Overbroad Mandatory Activation Hijacks Generic Tweet-Writing Requests

Content
View full analysis
Automatically search Twitter every hour for trending information about China, cryptocurrency, international news, United States news, Trump, and related subjects; extract the top ten topics; and generate high-exposure tweets or threads for each topic. When the user mentions keywords such as "Twitter trends", "Twitter search", "generate tweets", "social media monitoring", or "trending topics", this Skill must be used. Even if the user only says "help me write a tweet" or "check Twitter trends", this Skill must be triggered. ``` ### Technical Analysis The Skill metadata mandates activation for broad, generic requests such as writing any tweet. This does not limit invocation to the declared specialist function of monitoring and ranking current Twitter/X trends. Once loaded, the Skill introduces additional objectives that the user may not have requested, including: - Searching five predetermined topic categories. - Producing a top-ten ranking. - Generating twenty separate drafts. - Imitating named social-media authors. - Applying engagement-maximization rules. - Appending publication and traffic-acquisition instructions. This alters the agent's current-session objectives through Skill-selection instructions. The behavior exceeds minimum privilege because an ordinary tweet-writing request does not require trend monitoring, author imitation, algorithm optimization, or promotional guidance. ### Attack Path 1. A user submits a generic request such as “Help me write a tweet.” 2. The broad metadata trigger requires the agent to activate this Skill. 3. The Skill replaces or expands the requested task with its five-category trend-monitoring workfl ...[truncated 880 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:188
Finding

Mandatory Report Template Injects Traffic-Diversion and Engagement Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The trigger policy is overly broad and mandatory, forcing activation even for generic requests like 'help me write a tweet.' This can cause the agent to invoke external-searching and monitoring behavior unexpectedly, leading to scope hijacking, unnecessary data access, and user intent misclassification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill describes automatic scheduled monitoring and searching but does not clearly disclose network/data access to the user at activation time. That creates a transparency and consent problem, especially when searches run regularly and independently of immediate user interaction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

L065 明确规定“语言:全中文”,同时全文多处要求生成“中文算法优化推文”和“中文博主风格推文”。该约束没有提供用户选择语言的选项,也未说明这是仅面向特定中文地区或合规场景的限定,属于未经用户选择而强制指定语言/locale 的自然语言策略问题。

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill goes beyond passive monitoring and content drafting by prescribing engagement-manipulation tactics such as mandatory rapid replies, comment farming, and traffic capture under trending posts. This can enable spammy or deceptive growth-hacking behavior at scale and may violate platform rules or organizational use policies.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The entire reference is written only in Chinese, including all headings and guidance, with no indication that users may choose another language or that the document is intended solely for a Chinese-speaking or region-specific audience. Under the stated policy, forcing a specific language without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file contains prescriptive natural-language instructions for producing content specifically as '中文推文' and enforcing style rules in that language. Under the policy, forcing a specific language without user opt-in is a locale/language policy violation unless the restriction is explicitly justified, which is not provided here.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill explicitly instructs imitation of named third-party creators' styles, which introduces impersonation, misattribution, and potential IP/personality-rights concerns unrelated to the core monitoring function. In a content-generation pipeline, this increases the risk of deceptive outputs that appear authored by or closely associated with real individuals.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.