Back to skill

Security audit

Scraper

Security checks for vulnerabilities and agentic risk

Overview

This scraper is mostly straightforward, but its URL-fetching scripts do not enforce the promised public-only boundary and could be pointed at local, private-network, or local-file resources.

Review before installing. Use this only with URLs you personally trust and expect to be public. Avoid giving it URLs from untrusted pages or users until the implementation restricts schemes and destinations, blocks local/private/link-local targets and unsafe redirects, and enforces response-size limits.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/fetch_page.py:7
Finding

Unrestricted URL Fetching Enables SSRF and Local Resource Access

Content
View full analysis

Vulnerability Details

File Location: scripts/fetch_page.py:7-18, scripts/extract_text.py:21-32, and scripts/save_output.py:35-47
Vulnerability Type: Server-Side Request Forgery (SSRF) and unsafe URL scheme handling
Risk Level: High

Complete Vulnerable Code Snippets

scripts/fetch_page.py:7-18

python
parser.add_argument("--url", required=True, help="Public URL to fetch")
args = parser.parse_args()

req = urllib.request.Request(
    args.url,
    headers={
        "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)"
    }
)

with urllib.request.urlopen(req, timeout=20) as resp:
    html = resp.read().decode("utf-8", errors="replace")

scripts/extract_text.py:21-32

python
parser.add_argument("--url", required=True, help="Public URL to fetch and clean")
args = parser.parse_args()

req = urllib.request.Request(
    args.url,
    headers={
        "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)"
    }
)

with urllib.request.urlopen(req, timeout=20) as resp:
    html = resp.read().decode("utf-8", errors="replace")

scripts/save_output.py:35-47

python
parser.add_argument("--url", required=True, help="Public URL")
parser.add_argument("--title", required=True, help="Local title")
args = parser.parse_args()

req = urllib.request.Request(
    args.url,
    headers={
        "User-Agent": "Mozilla/5.0 (compatible; ScraperSkill/1.0)"
    }
)

with urllib.request.urlopen(req, timeout=20) as resp:
    html = resp.read().decode("utf-8", errors="replace")

Technical Analysis

All three scripts pass a user-controlled URL directly to urllib.request.Request and urllib.request.urlopen without validating its scheme, hostname, resolved IP address, destination port, or redirect targets.

Consequently, an attacker may direct requests toward loopback, private, link-local, or otherwise reserved network addresses. Depending on the handlers available to urllib.request, unsupported non-web scheme ...[truncated 2514 chars]

Remediation
View remediation

Remediation Suggestions

  1. Parse every URL with urllib.parse.urlsplit and permit only explicitly supported schemes, preferably https and, only if required, http.
  2. Reject URLs containing embedded credentials, malformed hostnames, fragments where inappropriate, or destination ports outside an explicit allowlist.
  3. Resolve the hostname before connecting and inspect every returned IPv4 and IPv6 address with the ipaddress module.
  4. Reject loopback, private, link-local, multicast, reserved, unspecified, and other non-global addresses.
  5. Explicitly block known metadata destinations, including link-local cloud metadata addresses.
  6. Disable automatic redirects or implement a custom redirect handler that reparses, resolves, and validates every redirect destination.
  7. Account for DNS rebinding by ensuring that the validated address is the address actually used for the connection, preferably through network-level egress controls.
  8. Apply an outbound firewall or proxy policy that prevents the scraper process from reaching local, private, management, and metadata networks.
  9. Stream response bodies in bounded chunks and enforce a strict maximum response size instead of using unbounded resp.read().
  10. Validate response content types before processing and provide controlled error handling without printing sensitive response details.
  11. Enforce robots, authorization, and public-resource restrictions in code where those guarantees are claimed, rather than relying solely on documentation.
  12. Apply the same centralized URL-validation function to all three affected scripts to prevent inconsistent protections.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill describes network access and local file read/write behavior, but it does not declare an explicit tool scope such as permissions or allowed-tools. That creates a governance gap: an agent runtime may permit broader capabilities than intended, making it easier for the skill to fetch arbitrary URLs or write local data without clear enforcement boundaries.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This code performs an HTTP request and retrieves remote content, but there is no confirmation prompt, user-facing log/print notice before the request, or comment/docstring disclosing the network activity. For code files, network calls that transmit user or system data should include some form of visible disclosure unless the warning is provided elsewhere.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code retrieves content from a user-supplied URL and writes the processed result to a local file, which are safety-relevant side effects under the warning criteria for code files. Although the script prints a success message after completion, there is no prior disclosure, confirmation, or comment/docstring warning that it will make a network request and save data to disk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.