Back to skill

Security audit

小红书违禁词检测

Security checks for vulnerabilities and agentic risk

Overview

This Xiaohongshu checker is mostly purpose-aligned, but it needs Review because it can read shell startup files for credentials, fetch arbitrary URLs from the host network, send checked content to RedFox, and create local output files.

Review before installing. Use this only with content you are comfortable sending to RedFox, avoid checking private or internal URLs, prefer explicitly setting REDFOX_API_KEY instead of relying on shell startup files, and remember that flagged results create a local downloadable text file containing the optimized copy.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/check_sensitive_words.py:31
Finding

Unnecessary Credential Discovery Through Shell Startup Files

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/check_sensitive_words.py:177
Finding

Server-Side Request Forgery Through Unrestricted Web Content Retrieval

Content
View full analysis
100: text = extracted break except Exception: continue if not text: text = page.inner_text('body') browser.close() return text.strip() except Exception: pass response = _http_request(url, method="GET", timeout=30) if response["status_code"] >= 400: raise Exception(f"Web request failed: HTTP {response['status_code']}") soup = BeautifulSoup(response["body"], 'html.parser') text = soup.get_text(separator='\n', strip=True) return text.strip() ``` ### Technical Analysis The `--url` argument is controlled by the user and is passed directly to Playwright's `pa ...[truncated 3238 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The documented behavior does not cleanly match the implemented/declared capabilities: it promises replacement suggestions and rewritten safe copy, but also includes undeclared third-party API transmission, webpage fetching, and text-extraction-only behavior. Such mismatches are dangerous because users may provide sensitive text under false assumptions about what the skill does locally versus what is sent to external services.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Scanning shell configuration files for credentials is an over-privileged behavior for a content-checking skill and can expose far more than the intended API key, including tokens, aliases, and other sensitive settings. Because the skill handles arbitrary user inputs, adding filesystem secret discovery makes the overall design materially more dangerous than its stated purpose suggests.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README invites invocation via very broad natural-language phrases such as "Xiaohongshu prohibited words," "note sensitive words," and related vague intents without clearly constraining what inputs, sources, or actions the skill should accept. In a skill that can process pasted text, uploaded files, images, and arbitrary URLs, this ambiguity can cause over-triggering, unintended handling of sensitive content, or invocation in contexts the user did not explicitly authorize, increasing the chance of data exposure or misuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README states that the skill will fetch and analyze text from user-provided webpage links, but it does not clearly warn that this requires network access and may transmit or retrieve third-party content. In practice, this can expose private URLs, internal resources, or sensitive browsing targets to external services without sufficiently informed user consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill advertises very broad, free-form trigger phrases such as general natural-language requests about prohibited words, sensitive words, or compliance. In an agent environment, this can cause unintended activation on loosely related user input, leading the skill to process content or files when the user did not clearly intend to invoke it.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill documents capabilities to read environment variables, access local files, and make network requests, but it does not declare any explicit tool scope or permissions boundary. This weakens reviewability and informed consent: a host or user cannot easily tell that uploaded content, shell config files, and external URLs may be accessed during execution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill sends user-provided text, file contents, or extracted webpage content to a third-party API, but the user-facing description does not present this as a clear up-front warning before use. This creates a privacy and data-handling risk because users may submit unpublished marketing copy, internal documents, or sensitive business content without realizing it will leave the local environment.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest and usage sections state that users can upload images and have text extracted for detection, but the listed dependencies and documented script architecture only cover DOC/DOCX/TXT parsing, webpage extraction, and API-based sensitive-word checking. No OCR library, image-processing dependency, or image extraction function is declared, so this capability appears unjustified relative to the described implementation.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill claims user data is not stored locally, yet it also says it generates and sends a downloadable optimized text file. This contradiction can mislead users about persistence and retention of sensitive content, which is especially important because the tool handles potentially confidential drafts and uploads.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The workflow broadens credential access by instructing the skill to read shell startup files to discover API keys, rather than limiting access to explicitly provided environment variables. Reading user shell configuration files is unrelated to the visible purpose of banned-word checking and creates a path to access secrets and unrelated local configuration data.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The workflow mandates writing the processed user content to a local file and sending it back as a downloadable artifact, which expands the skill from transient analysis into persistent file handling. This increases data exposure risk because sensitive user text is retained on disk and redistributed through an additional channel not necessary for prohibited-word detection.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Persisting processed user content to disk and then transmitting it back as a file creates a secondary data exposure channel beyond the immediate response. Even if the content originates from the user, local persistence can leave residual sensitive data on the host and increases the chance of accidental disclosure or cross-session access.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill searches users' shell startup files to extract an API key if the environment variable is not set. Reading unrelated local credential sources is unnecessary for prohibited-word checking and expands the skill's access to secrets beyond user intent, creating a clear secret-harvesting and privacy boundary violation. In this skill's context, that behavior is more dangerous because the same script also sends content to an external service, so it combines local secret access with outbound network capability.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script sends user-provided content, including extracted file or webpage text, to a third-party API for scanning without an explicit warning, consent prompt, or local-only option. This can expose sensitive document contents, proprietary text, or personal data to an external service unexpectedly; in this skill's context that is especially relevant because users may submit private drafts or internal files for checking.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill offers one-click export of an optimized text file but does not warn that it will generate a downloadable file containing processed user content. This can create unintended persistence of potentially sensitive text on disk or in platform storage, especially if the user assumes results are only shown transiently in chat.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The entire skill documentation, trigger phrases, examples, and expected interaction flow are presented only in Chinese, with no indication that users may choose another language or that Chinese is required for a documented compliance reason. Under the language/locale policy, a forced single-language interaction should either be optional or explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The script's user-visible messages and CLI descriptions are entirely in Chinese, including dependency errors and operational prompts. This imposes a fixed language on users without opt-in or any documented locale constraint, which matches the language/locale policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

Argument descriptions and runtime error messages are presented only in Chinese, which can force a specific locale on all users. There is no mechanism to choose another language or documentation justifying a Chinese-only constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest describes a skill for detecting Xiaohongshu prohibited words, highlighting them, and providing replacement suggestions or revised copy. The --extract-only path returns extracted text and length without performing any prohibited-word check, which adds a standalone content extraction capability not reflected in the manifest description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.