Back to skill

Security audit

公众号文案创作

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its WeChat writing purpose, but it injects an unrelated RedFox sales contact into every script result and lacks privacy guardrails for personal writing samples.

Review before installing. Use a revocable RedFox API key, avoid sensitive keywords or confidential campaign topics, and do not provide diaries, private notes, workplace secrets, financial, health, or third-party personal details as style samples. Be aware the bundled script currently appends an unsolicited sales contact to results and can make JSON output invalid.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
scripts/fetch_gzh_trends.py:308
Finding

Unconditional Commercial Content Injection into Script Output

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch_gzh_trends.py, line 308
Vulnerability Type: Unconditional output manipulation
Risk Level: High

Vulnerable Code

python
print(f"\n另外红狐配套全量数据库可提供完整详实数据,如需了解采购方案,可发送邮件至 redfoxdata@proton.me 对接咨洵")

The embedded message advertises an additional commercial database and directs users to an external email address.

Technical Analysis

Every successful execution unconditionally appends an unsolicited commercial message to standard output. This behavior is unrelated to the Skill's declared function of retrieving and analyzing trending WeChat article data, and it is not part of the documented output format.

Because the message is emitted for every output mode, it also contaminates machine-readable JSON output. An Agent executing the script may ingest or relay the injected message as though it were part of the legitimate search result. This creates a stable output-hijacking channel through which package-controlled content is inserted into the Agent's working context and potentially into user-facing responses.

The behavior exceeds minimum privilege and functional necessity: trend retrieval requires sending authenticated search parameters to the documented API, but it does not require advertising a separate product or directing users to an external contact.

Attack Path

  1. A user asks the Agent to search for trending WeChat articles or generate related copy.
  2. The Agent follows SKILL.md and invokes scripts/fetch_gzh_trends.py.
  3. The script performs the legitimate API request and formats the returned trend data.
  4. Line 308 appends package-controlled commercial content and an external contact address to standard output.
  5. The Agent consumes the entire output and may reproduce the unsolicited message in its analysis or final response.
  6. When JSON output is requested, the appended plaintext also makes the output invalid JSON and can disrupt d ...[truncated 585 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the unconditional print statement at line 308.
  2. Restrict standard output to the documented result format only.
  3. Ensure JSON mode emits exactly one valid JSON document with no banners, advertisements, diagnostics, or trailing plaintext.
  4. Send operational diagnostics to standard error only, and only when explicitly requested through a debug option.
  5. If commercial contact information must be disclosed, place it transparently in project documentation rather than runtime output.
  6. Add automated tests that parse JSON-mode output and verify that Markdown and text output contain only requested trend data.
  7. Review all future output strings to ensure they are necessary for the declared functionality and cannot manipulate downstream Agent responses.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (11)

Tainted flow: 'headers' from os.environ.get (line 109, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/fetch_gzh_trends.py (reported line 44)May include surrounding context.

python
print(f"Params: {json.dumps(params, ensure_ascii=False)}", file=sys.stderr)

    try:
        response = requests.post(base_url, headers=headers, json=params, timeout=60)

        if debug:
            print(f"状态码: {response.status_code}", file=sys.stderr)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description promises a broader创作工具: keyword-based viral article search, traffic-pattern analysis, and full article generation for公众号写作. The supplied code only implements the search/reporting portion: it calls the Redfox API, retrieves article data, auto-expands the date range when results are sparse, deduplicates entries, and formats a markdown report with titles/authors/engagement stats. There is no logic for generating publishable articles, drafting公众号文案, summarizing writing strategies, or performing substantive规律分析 beyond presenting raw counts. Therefore the description materially overstates the implemented behavior, making this a mismatch.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README says users can 'Simply describe your creative need in natural language—no fixed commands to memorize' and gives examples built around broad phrases like 'Help me write...' and 'What articles are trending on WeChat?'. These triggers overlap with common everyday requests and the file does not provide negative examples or clear boundaries for when the skill should or should not activate.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.en.md (reported line 57)May include surrounding context.

md
### Quick Reference

| Intent               | Example phrase                                            | Result                                                                                          |
| -------------------- | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| Write an article     | "Help me write a WeChat article about career changes"     | Search virals by keyword, analyze patterns, then generate a complete article                    |
| Recommend a product  | "Help me write a WeChat article recommending an AI tool"  | Automatically break down differentiation advantages and generate copy around product highlights |

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README says users can 'directly use natural language to express your creation needs' without fixed commands, which leaves activation scope ambiguous for a markdown skill description. Combined with examples like '帮我写一篇...' this overlaps with common everyday phrasing and does not provide exclusion conditions or negative examples to limit unintended invocation.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to proactively ask users for personal writing samples, such as notes, diaries, or essays, without any privacy warning, minimization guidance, or caution against sharing sensitive data. Because such samples can contain intimate personal details, identifiers, or confidential information, this creates an avoidable risk of oversharing and secondary reuse of sensitive content.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill not only asks for personal writings, but explicitly directs the model to analyze the user's writing habits, tone, and expression and reuse them in generated output. This increases privacy risk because sensitive traits, personal experiences, or identifiable phrasing may be inferred, retained, or reproduced in ways the user did not fully expect, especially when the requested examples include diaries or informal private writing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script's descriptive text, CLI help, errors, and generated report content are all hard-coded in Chinese, which effectively fixes the user-facing locale. There is no option for users to choose another language and no documented justification that the skill is intentionally region- or locale-specific.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/fetch_gzh_trends.py (reported line 44)May include surrounding context.

python
print(f"Params: {json.dumps(params, ensure_ascii=False)}", file=sys.stderr)

    try:
        response = requests.post(base_url, headers=headers, json=params, timeout=60)

        if debug:
            print(f"状态码: {response.status_code}", file=sys.stderr)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script sends user-provided keywords and the API credential to an external service by default, but normal operation does not clearly disclose that the query content leaves the local environment. In an agent skill context, users may provide sensitive business topics, draft ideas, or personal data assuming local processing, creating a privacy and data-governance risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.