T01 · Skill Instruction Hijacking
- Location
SKILL.md:182- Finding
Mandatory Promotional Output Hijacks Agent Responses
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a real prohibited-word checker, but it needs review because it handles secrets unsafely and fetches/submits user content with too little containment.
Review before installing. Use only non-confidential content unless you trust the RedFox service and any configured gateway, avoid checking private or internal URLs, do not let the agent print your API key, and prefer a session-scoped secret or dedicated secret store over shell-profile persistence.
SKILL.md:182Mandatory Promotional Output Hijacks Agent Responses
scripts/extract_text.py:266Server-Side Request Forgery Through Unrestricted Web Extraction
SKILL.md:72API Key Persisted and Displayed in Plaintext
The request destination is taken from the PROHIBITED_WORD_API_URL environment variable and arbitrary user content plus the X-API-KEY credential are sent to it. Although the code requires HTTPS, that does not prevent exfiltration to an attacker-controlled host if the environment is poisoned or misconfigured, making this a real SSRF/external-exfiltration risk in a skill that processes potentially sensitive text.
try:
# 发起 HTTPS POST 请求
response = requests.post(api_url, headers=headers, json=params, timeout=30)
if response.status_code >= 400:
raise Exception(f"HTTP请求失败: {response.status_code}, {response.text[:500]}")
The skill advertises broader detection and media-processing capabilities than are actually implemented, while the real behavior includes web fetching and parsing. This can cause users to rely on incomplete safety checks and unknowingly grant the agent more access than expected, particularly for remote content retrieval.
The skill advertises broader detection and media-processing capabilities than are actually implemented, while the real behavior includes web fetching and parsing. This can cause users to rely on incomplete safety checks and unknowingly grant the agent more access than expected, particularly for remote content retrieval.
The README advertises uploading text files, images, and URLs for automatic processing but does not clearly warn users that submitted content may be sent to an external service or otherwise leave the local environment. This creates a real privacy and data-handling risk because users may submit sensitive copy, internal documents, or private page content without informed consent.
Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.
### Quick Reference
| Intent | Example phrase | Result |
| --------------------- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| Check WeChat copy | `Check this WeChat article for prohibited words: [copy]` | Detects against WeChat rules, outputs flagged results and replacement suggestions |
| Switch to Xiaohongshu | `Check this Xiaohongshu post for prohibited words: [copy]` | Switches to Xiaohongshu word library for detection |
Documenting automatic URL body extraction without a privacy or integrity warning is risky because users may cause the tool to fetch and process internal, private, or sensitive pages without understanding the consequences. It also increases the chance of unsafe retrieval behavior or unintentional transmission of page contents to a third-party backend.
Large whitespace padding was detected (a block of blank lines or a long run of spaces). This can push injected instructions below or to the right of the visible area so a human reviewer never sees them while the agent still reads them. Manual review of the hidden content is recommended.
## Use Cases
| Scenario | Role | Example question | Benefit |
| ------------------------------ | ---------------------- | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| Pre-publish self-audit | WeChat operator | `Check this article before I send it out` | Catch advertising law violations and false claims before mass send, reducing rejection or throttling risk |
| Xiaohongshu post compliance | Xiaohongshu creator | `Xiaohongshu check: [post content]` | Flag extreme claims and community-banned phrases, get replacement suggestions tuned to the seeding context |
The README states that pasted webpage links will be automatically fetched and analyzed, but it does not warn users that this causes external network access and may transmit sensitive URLs or page content for processing. In a privacy-sensitive agent setting, automatic fetching can expose internal, private, tokenized, or user-specific links without informed consent.
The README encourages invocation through very broad natural-language phrases such as generic requests to 'check this copy' or 'detect this page,' which can overlap with ordinary user conversation. In an agent environment, this increases the chance of accidental or over-broad skill activation on content the user did not clearly intend to submit, potentially causing unintended processing of sensitive text, files, images, or links.
The skill declares capabilities that require environment access, file reading, and network access, but it does not explicitly scope or constrain those tools. This weakens the trust boundary and makes it harder for a host agent or reviewer to determine what actions are expected or permissible, especially because the skill also handles credentials and external requests.
The skill description and prescribed interaction are entirely in Chinese and assume Chinese-language commands and outputs, with no indication that users may choose another language. Because the file does not document a justified locale restriction or provide opt-in language selection, this can constitute a language/locale policy issue.
The skill instructs the agent to persistently modify user shell profiles or environment settings in order to store an API key. Persistently changing user configuration is beyond the core purpose of word checking and can create unwanted side effects, credential exposure risks, and durable changes the user may not understand or be able to audit.
The skill directs persistent API key setup without a clear warning that shell profiles or user environment configuration will be modified. This undermines informed consent and can leave long-lived secrets stored in places that may be readable by other processes, backed up unintentionally, or difficult for the user to remove.
The skill explicitly recommends verifying the API key by echoing it back, which normalizes exposing secrets in plain output. This is particularly dangerous in agent environments because responses, terminal output, logs, chat transcripts, and screenshots may capture the credential, leading to straightforward secret compromise.
The documentation tells the agent to run shell or PowerShell commands to set up and verify credentials, which expands the agent’s operational scope beyond content analysis. Command execution for secret handling is dangerous because it can alter user systems, leak secrets into logs or terminal history, and normalize unsafe operational behavior.
The manifest describes a capability grounded in an 'official prohibited-word library' and multi-format checking, which implies the skill itself performs or directly embodies that checking logic. This file instead sends content to a third-party backend and states that no local library is held, which materially differs from the manifest's representation of how the checking is performed.
The script reads multiple shell startup files in the user's home directory to extract REDFOX_API_KEY, which exceeds the minimum data access needed for a word-checking tool. This broad local-file access can expose unrelated secrets or sensitive configuration content and is especially concerning in an agent skill context where users may not expect inspection of personal shell profiles.
The tool transmits the full input content to an external service for analysis. In the context of a content-checking skill, users may submit drafts, unpublished marketing copy, or regulated/sensitive text, so outbound transfer to a third party creates a genuine confidentiality and compliance risk even if it is the intended product behavior.
try:
# 发起 HTTPS POST 请求
response = requests.post(api_url, headers=headers, json=params, timeout=30)
if response.status_code >= 400:
raise Exception(f"HTTP请求失败: {response.status_code}, {response.text[:500]}")
The stated skill scope includes multiple input forms such as files, images, and links. In this file, the CLI only takes a text string and passes it as the content field, with no code for file parsing, image OCR, or link retrieval/inspection, so the implemented behavior is narrower than the manifest claims.
The declared skill scope includes 图片 and multiple input forms, but this implementation only accepts a fixed set of text-oriented file extensions and a web URL mode. There is no image OCR, image parsing, or image-type handling here, so the behavior does not match the stated multimodal scope.
Multiple extraction paths require the presence of Chinese characters or preferentially extract Chinese text blocks, effectively constraining successful operation to Chinese-language content. The file does not offer a language choice or document this locale restriction as an intentional, justified limitation.
The fallback strategy explicitly extracts continuous Chinese text blocks and ignores other languages, which is a natural-language locale constraint. Because this restriction is neither user-selectable nor clearly documented in this file, it violates the language/locale policy criteria.
The manifest claims a moderation-oriented capability that outputs prohibited-word markings and replacement suggestions across multiple input forms. In this file, the implemented behavior is limited to reading text from supported local text files or fetching/parsing webpage content, with no prohibited-word library lookup, marking, auditing-standard comparison, image handling, or suggestion generation.
The code fetches arbitrary user-supplied URLs server-side without any visible warning, validation, or restriction. In a hosted skill environment, this can enable SSRF-style access to internal services or unintentionally send requests to attacker-controlled endpoints, exposing network metadata and creating privacy and security risk.
No suspicious patterns detected.