Back to skill

Security audit

多平台违禁词检测

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real prohibited-word checker, but it needs review because it handles secrets unsafely and fetches/submits user content with too little containment.

Review before installing. Use only non-confidential content unless you trust the RedFox service and any configured gateway, avoid checking private or internal URLs, do not let the agent print your API key, and prefer a session-scoped secret or dedicated secret store over shell-profile persistence.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:182
Finding

Mandatory Promotional Output Hijacks Agent Responses

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/extract_text.py:266
Finding

Server-Side Request Forgery Through Unrestricted Web Extraction

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:72
Finding

API Key Persisted and Displayed in Plaintext

Content
View full analysis
`;若用户不会配置,Agent应主动帮用户设置: - **macOS/Linux**:将 `export REDFOX_API_KEY=<值>` 追加到 `~/.zshrc`(zsh)或 `~/.bashrc`(bash),然后 `source` 对应文件使其全局生效 - **Windows**:使用 `[Environment]::SetEnvironmentVariable("REDFOX_API_KEY", "<值>", "User")` 设置用户级永久环境变量(需重启终端生效) - 配置完成后应验证:`echo $REDFOX_API_KEY`(macOS/Linux)或 `echo %REDFOX_API_KEY%`(Windows),确保换一个skill也能读取到 ``` ### Technical Analysis The skill instructs the agent to persist the API key in shell startup files or a user-level Windows environment variable. It additionally recommends verifying configuration by printing the complete secret to the terminal. Shell startup files commonly contain plaintext and may be readable by backup software, support tooling, malware running under the same account, accidental diagnostic uploads, or other processes with user-level access. Printing the key can expose it through terminal history, session recording, screenshots, logs, copied transcripts, or agent tool output. The instruction is especially risky because it tells the agent to perform the persistence operation proactively when the user does not know how to configure the key. This expands the agent's activity from using a credential for the current task to modifying persistent user configuration. ### Attack Path 1. A user provides a valid `REDFOX_API_KEY` to configure the skill. 2. Following `SKILL.md`, the agent writes the key in plaintext to a shell startup file or persistent user environment storage. 3. The agent verifies the configuration by printing the complete value. 4. A local process, backup, terminal recorder, log collector, screenshot, or shared transcript captures the secret. 5. An unauthorized party obtains the key and submits requests to the associated API ...[truncated 629 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (24)

Tainted flow: 'api_url' from os.environ.get (line 85, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
94% confidence
Finding

The request destination is taken from the PROHIBITED_WORD_API_URL environment variable and arbitrary user content plus the X-API-KEY credential are sent to it. Although the code requires HTTPS, that does not prevent exfiltration to an attacker-controlled host if the environment is poisoned or misconfigured, making this a real SSRF/external-exfiltration risk in a skill that processes potentially sensitive text.

Content

Scanner excerpt · scripts/check_sensitive_words.py (reported line 115)May include surrounding context.

python
try:
        # 发起 HTTPS POST 请求
        response = requests.post(api_url, headers=headers, json=params, timeout=30)

        if response.status_code >= 400:
            raise Exception(f"HTTP请求失败: {response.status_code}, {response.text[:500]}")

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill advertises broader detection and media-processing capabilities than are actually implemented, while the real behavior includes web fetching and parsing. This can cause users to rely on incomplete safety checks and unknowingly grant the agent more access than expected, particularly for remote content retrieval.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises broader detection and media-processing capabilities than are actually implemented, while the real behavior includes web fetching and parsing. This can cause users to rely on incomplete safety checks and unknowingly grant the agent more access than expected, particularly for remote content retrieval.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README advertises uploading text files, images, and URLs for automatic processing but does not clearly warn users that submitted content may be sent to an external service or otherwise leave the local environment. This creates a real privacy and data-handling risk because users may submit sensitive copy, internal documents, or private page content without informed consent.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.en.md (reported line 53)May include surrounding context.

md
### Quick Reference

| Intent                | Example phrase                                                  | Result                                                                                   |
| --------------------- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| Check WeChat copy     | `Check this WeChat article for prohibited words: [copy]`        | Detects against WeChat rules, outputs flagged results and replacement suggestions        |
| Switch to Xiaohongshu | `Check this Xiaohongshu post for prohibited words: [copy]`      | Switches to Xiaohongshu word library for detection                                       |

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Documenting automatic URL body extraction without a privacy or integrity warning is risky because users may cause the tool to fetch and process internal, private, or sensitive pages without understanding the consequences. It also increases the chance of unsafe retrieval behavior or unintentional transmission of page contents to a third-party backend.

Content

No source excerpt is available for this finding.

Whitespace Padding

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.

Content

Scanner excerpt · README.en.md (reported line 76)May include surrounding context.

md
## Use Cases

| Scenario                       | Role                   | Example question                                   | Benefit                                                                                                    |
| ------------------------------ | ---------------------- | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| Pre-publish self-audit         | WeChat operator        | `Check this article before I send it out`          | Catch advertising law violations and false claims before mass send, reducing rejection or throttling risk  |
| Xiaohongshu post compliance    | Xiaohongshu creator    | `Xiaohongshu check: [post content]`                | Flag extreme claims and community-banned phrases, get replacement suggestions tuned to the seeding context |

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README states that pasted webpage links will be automatically fetched and analyzed, but it does not warn users that this causes external network access and may transmit sensitive URLs or page content for processing. In a privacy-sensitive agent setting, automatic fetching can expose internal, private, tokenized, or user-specific links without informed consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README encourages invocation through very broad natural-language phrases such as generic requests to 'check this copy' or 'detect this page,' which can overlap with ordinary user conversation. In an agent environment, this increases the chance of accidental or over-broad skill activation on content the user did not clearly intend to submit, potentially causing unintended processing of sensitive text, files, images, or links.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill declares capabilities that require environment access, file reading, and network access, but it does not explicitly scope or constrain those tools. This weakens the trust boundary and makes it harder for a host agent or reviewer to determine what actions are expected or permissible, especially because the skill also handles credentials and external requests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
71% confidence
Finding

The skill description and prescribed interaction are entirely in Chinese and assume Chinese-language commands and outputs, with no indication that users may choose another language. Because the file does not document a justified locale restriction or provide opt-in language selection, this can constitute a language/locale policy issue.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to persistently modify user shell profiles or environment settings in order to store an API key. Persistently changing user configuration is beyond the core purpose of word checking and can create unwanted side effects, credential exposure risks, and durable changes the user may not understand or be able to audit.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill directs persistent API key setup without a clear warning that shell profiles or user environment configuration will be modified. This undermines informed consent and can leave long-lived secrets stored in places that may be readable by other processes, backed up unintentionally, or difficult for the user to remove.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly recommends verifying the API key by echoing it back, which normalizes exposing secrets in plain output. This is particularly dangerous in agent environments because responses, terminal output, logs, chat transcripts, and screenshots may capture the credential, leading to straightforward secret compromise.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation tells the agent to run shell or PowerShell commands to set up and verify credentials, which expands the agent’s operational scope beyond content analysis. Command execution for secret handling is dangerous because it can alter user systems, leak secrets into logs or terminal history, and normalize unsafe operational behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest describes a capability grounded in an 'official prohibited-word library' and multi-format checking, which implies the skill itself performs or directly embodies that checking logic. This file instead sends content to a third-party backend and states that no local library is held, which materially differs from the manifest's representation of how the checking is performed.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script reads multiple shell startup files in the user's home directory to extract REDFOX_API_KEY, which exceeds the minimum data access needed for a word-checking tool. This broad local-file access can expose unrelated secrets or sensitive configuration content and is especially concerning in an agent skill context where users may not expect inspection of personal shell profiles.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The tool transmits the full input content to an external service for analysis. In the context of a content-checking skill, users may submit drafts, unpublished marketing copy, or regulated/sensitive text, so outbound transfer to a third party creates a genuine confidentiality and compliance risk even if it is the intended product behavior.

Content

Scanner excerpt · scripts/check_sensitive_words.py (reported line 115)May include surrounding context.

python
try:
        # 发起 HTTPS POST 请求
        response = requests.post(api_url, headers=headers, json=params, timeout=30)

        if response.status_code >= 400:
            raise Exception(f"HTTP请求失败: {response.status_code}, {response.text[:500]}")

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The stated skill scope includes multiple input forms such as files, images, and links. In this file, the CLI only takes a text string and passes it as the content field, with no code for file parsing, image OCR, or link retrieval/inspection, so the implemented behavior is narrower than the manifest claims.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The declared skill scope includes 图片 and multiple input forms, but this implementation only accepts a fixed set of text-oriented file extensions and a web URL mode. There is no image OCR, image parsing, or image-type handling here, so the behavior does not match the stated multimodal scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Multiple extraction paths require the presence of Chinese characters or preferentially extract Chinese text blocks, effectively constraining successful operation to Chinese-language content. The file does not offer a language choice or document this locale restriction as an intentional, justified limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The fallback strategy explicitly extracts continuous Chinese text blocks and ignores other languages, which is a natural-language locale constraint. Because this restriction is neither user-selectable nor clearly documented in this file, it violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest claims a moderation-oriented capability that outputs prohibited-word markings and replacement suggestions across multiple input forms. In this file, the implemented behavior is limited to reading text from supported local text files or fetching/parsing webpage content, with no prohibited-word library lookup, marking, auditing-standard comparison, image handling, or suggestion generation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code fetches arbitrary user-supplied URLs server-side without any visible warning, validation, or restriction. In a hosted skill environment, this can enable SSRF-style access to internal services or unintentionally send requests to attacker-controlled endpoints, exposing network metadata and creating privacy and security risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.