Back to skill

Security audit

Scrapling Fetch Pro

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to be a real web scraper, but it needs Review because it promotes anti-bot scraping and accepts arbitrary URLs without clear scope or safety guidance.

Install only if you intend to use a broad web-scraping tool and can control where it runs. Avoid using it against sites you are not authorized to access, be cautious with stealth mode, do not run it from environments that can reach private services or cloud metadata, and treat scraped content as untrusted before giving it to an AI agent.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/scrapling_fetch.py:243
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/scrapling_fetch.py, lines 243-266
Vulnerability Type: Server-Side Request Forgery (SSRF)
Risk Level: High

Vulnerable Code

python
parser.add_argument("url", help="目标网页 URL")
parser.add_argument("max_chars", nargs="?", type=int, default=30000,
                    help="最大输出字符数(默认: 30000)")
parser.add_argument("--mode", choices=["basic", "stealth", "auto"], default="auto",
                    help="抓取模式: basic(快速)/ stealth(隐身)/ auto(自动检测)")
parser.add_argument("--json", action="store_true", help="JSON 格式输出")
parser.add_argument("--debug", action="store_true", help="显示调试信息")

args = parser.parse_args()

# 确定抓取模式
if args.mode == "auto":
    mode = "stealth" if needs_stealth_mode(args.url) else "basic"
    if args.debug:
        print(f"[DEBUG] 自动选择模式: {mode}", file=sys.stderr)
else:
    mode = args.mode

try:
    # 抓取页面
    if mode == "stealth":
        html, selector, is_wechat = fetch_stealth(args.url, args.debug)
    else:
        html, selector, is_wechat = fetch_basic(args.url, args.debug)

Technical Analysis

The command-line caller has complete control over args.url, which is passed to either fetch_stealth or fetch_basic without security validation. The implementation does not:

  • Restrict URLs to the expected http and https schemes.
  • Reject loopback, private, link-local, multicast, reserved, or unspecified IP addresses.
  • Resolve hostnames and validate all returned IP addresses.
  • Revalidate redirect destinations.
  • Enforce an approved-domain allowlist.
  • Prevent DNS rebinding between validation and connection.

Consequently, when the Skill runs in an environment with access to internal networks, local services, or cloud metadata endpoints, an attacker can use it as an SSRF primitive.

Attack Path

  1. An attacker supplies a URL targeting a non-public resource, such as a loopback service, priv ...[truncated 1483 chars]
Remediation
View remediation

Remediation Suggestions

  1. Parse URLs with a standards-compliant URL parser and permit only http and https.
  2. Reject URLs containing credentials or ambiguous hostname representations.
  3. Resolve the destination hostname before connecting and reject every address classified as loopback, private, link-local, multicast, reserved, or unspecified.
  4. Apply the same checks to IPv4, IPv6, IPv4-mapped IPv6 addresses, and alternative numeric address forms.
  5. Disable automatic redirects or validate the destination after every redirect before issuing the next request.
  6. Prefer an explicit allowlist of approved domains when the expected scraping scope is known.
  7. Mitigate DNS rebinding by connecting to a validated resolved address while preserving the intended hostname for TLS verification and the HTTP Host value.
  8. Enforce outbound network restrictions at the container, firewall, or proxy layer so the process cannot reach metadata, loopback, or private network services.
  9. Add automated tests covering loopback addresses, RFC 1918 networks, IPv6 local addresses, cloud metadata addresses, encoded IP representations, redirects, and DNS rebinding scenarios.

T01 · Skill Instruction Hijacking

Warning
Location
references/usage.md:123
Finding

Documented AI Workflow Directly Injects Untrusted Web Content into a Prompt

Content
View full analysis

Vulnerability Details

File Location: references/usage.md, lines 123-130
Vulnerability Type: Indirect prompt injection through untrusted webpage content
Risk Level: Medium

Vulnerable Code

bash
# 抓取文章内容并保存
content=$(python3 scripts/scrapling_fetch.py "https://example.com" 20000)

# 让 AI 总结内容
echo "请总结以下文章:\n\n$content" | ai-chat

Technical Analysis

The documented workflow retrieves arbitrary webpage content and interpolates it directly into a prompt sent to ai-chat. Webpage text is attacker-controlled data and may contain instructions intended for an AI model, such as demands to ignore the user's summarization request, disclose context, invoke tools, or perform unrelated actions.

The example does not establish a trust boundary between the user's instruction and the fetched document. It also provides no explicit instruction that directives contained in the webpage must be treated as quoted data rather than executable instructions. Markdown conversion and noise removal do not neutralize natural-language prompt injection.

The vulnerability becomes more consequential if ai-chat has access to tools, credentials, files, network operations, persistent memory, or other privileged capabilities.

Attack Path

  1. An attacker publishes or modifies a webpage so that its visible or extracted content includes malicious AI instructions.
  2. A user follows the documented workflow and fetches that webpage.
  3. The scraper converts the attacker-controlled content to Markdown and stores it in the content shell variable.
  4. The workflow concatenates that content directly with the summarization request.
  5. ai-chat receives the malicious webpage instructions in the same prompt context as the legitimate request.
  6. If the model prioritizes or follows the embedded instructions, it may produce manipulated output or misuse any tools available to it.

Impact Assessment

This issue does not indepe ...[truncated 603 chars]

Remediation
View remediation

Remediation Suggestions

  1. Clearly document that all fetched content is untrusted and may contain prompt-injection instructions.
  2. Use a structured prompt that explicitly states that the webpage is data and that instructions, requests, links, or commands inside it must never be followed.
  3. Place fetched content in a clearly delimited or structured field rather than directly concatenating it with the user's instruction.
  4. Use a dedicated summarization model or execution profile with all tools, file access, network access, persistent memory, and credential access disabled.
  5. Require explicit user confirmation before allowing any downstream action derived from fetched content.
  6. Validate model output before passing it into another command, tool invocation, or privileged workflow.
  7. Treat filtering as defense in depth only; keyword removal cannot reliably neutralize natural-language prompt injection.
  8. Update the usage example to demonstrate safe separation between trusted instructions and untrusted document content.
Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises and demonstrates network-enabled scraping behavior but does not declare any tool scope such as permissions or allowed-tools. That omission weakens governance and review controls, making it easier for an agent to invoke network access implicitly for broad web collection, including anti-bot bypass scenarios described in the skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill promotes scraping third-party sites, WeChat public-account articles, and anti-bot bypass features without any warning about authorization, privacy, terms-of-service, or legal/compliance constraints. In this context, the omission materially increases misuse risk because the skill is specifically positioned for broad extraction from protected or monitored sites.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation explicitly promotes a stealth mode to bypass Cloudflare and anti-bot protections, but provides no warning about authorization, terms-of-service, or legal/compliance boundaries. In a scraping-focused skill, this materially increases the likelihood of misuse by normalizing evasion of access controls and anti-abuse mechanisms.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

SQP-3 applies to all file types and covers language or locale policy violations in natural-language strings. Here, the module docstring, usage examples, and later CLI help strings present the tool only in Chinese, with no opt-in, alternative locale, or justification that the tool is region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description is entirely in Chinese, which can function as a language-specific policy choice in user-facing skill metadata. There is no indication that the skill is region-specific or that users can opt into this locale, so it may violate language/locale neutrality expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The markdown provides a ready-to-run batch script that saves scraped page content into an output directory, but it does not disclose that this will persist potentially sensitive or copyrighted material locally. For markdown guidance, safety-relevant behaviors that affect user data should be accompanied by a warning.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.