Back to skill

Security audit

Page Fetch

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does webpage extraction as advertised, but it can fetch arbitrary URLs and handle raw session cookies without enough guardrails.

Install only if you are comfortable with the agent making live web requests from its runtime. Do not use it on untrusted URLs in environments that can reach private services, localhost admin panels, or cloud metadata endpoints, and avoid passing real account cookies unless you trust the exact target and understand that command-line secrets can leak through process or tool logs.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/fetch_page.py:39
Finding

Arbitrary URL Fetching Enables Server-Side Request Forgery

Content
View full analysis
Tuple[str, requests.Response, List[str]]: notes = [] headers = {"User-Agent": DEFAULT_UA, "Accept-Language": "en-US,en;q=0.9"} try: response = requests.get(url, headers=headers, timeout=timeout) response.raise_for_status() return response.text, response, notes except requests.RequestException as exc: notes.append(f"requests failed: {exc}") raise ``` The browser fallback also navigates directly to the unvalidated URL: ```javascript (async () => { const argv = process.argv.slice(-2); const url = argv[0]; const waitMs = Number(argv[1] || 2500); const browser = await chromium.launch({ headless: true }); const page = await browser.newPage({ userAgent: "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36" }); await page.goto(url, { waitUntil: "domcontentloaded", timeout: 30000 }); ``` The unified runner forwards user-controlled URLs into both affected paths: ```python if is_wechat_url(args.url): steps.append("wechat-first") code, payload = run_json(FETCH_WECHAT, args.url, timeout=max(args.timeout, 30), max_chars=args.max_chars, cookie=args.cookie) if not is_good_enough(payload): steps.append("wechat-fallback-general") code, payload = run_json(FETCH_PAGE, args.url, timeout=args.timeout, max_chars=args.max_chars) if not is_good_enough(payload): steps.append("wechat-fallback-browser") code, payload = run_json(RENDER_PAGE, args.url, timeout=0, max_chars=args.max_chars, wait_ms=args.wait_ms) else: steps.append("general-first") code, payload = run_json(FETCH_PAGE, ar ...[truncated 2662 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/fetch_wechat_article.py:23
Finding

Weak WeChat Hostname Validation Can Disclose Caller-Supplied Cookies

Content
View full analysis
bool: parsed = urlparse(url) return parsed.netloc.endswith("mp.weixin.qq.com") ``` ```python def fetch_html(url: str, timeout: int, cookie: Optional[str] = None): headers = dict(MOBILE_HEADERS) if cookie: headers["Cookie"] = cookie response = requests.get(url, headers=headers, timeout=timeout, allow_redirects=True) response.raise_for_status() return response.text, response ``` The cookie is accepted from the command line and passed to the request: ```python parser.add_argument("--cookie", default=None) ``` ```python html, response = fetch_html(args.url, args.timeout, args.cookie) ``` The unified runner uses the same unsafe suffix check and forwards the cookie: ```python def is_wechat_url(url: str) -> bool: return urlparse(url).netloc.endswith("mp.weixin.qq.com") ``` ```python if is_wechat_url(args.url): steps.append("wechat-first") code, payload = run_json(FETCH_WECHAT, args.url, timeout=max(args.timeout, 30), max_chars=args.max_chars, cookie=args.cookie) ``` ### Technical Analysis The code checks `netloc.endswith("mp.weixin.qq.com")` without requiring a DNS label boundary. A hostname such as `evilmp.weixin.qq.com` satisfies the string suffix test even though it is not `mp.weixin.qq.com` and is not a subdomain separated by a dot. Using `netloc` instead of the normalized `hostname` property also mixes host, port, and possible user-information syntax into security-sensitive validation. After this weak check succeeds, the script places the caller-provided cookie directly into th ...[truncated 1657 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a generic, stable workflow for extracting readable content from webpages broadly, with a sequence of raw HTML fetch, embedded-data inspection, and browser escalation when lightweight methods fail. The supplied code does only a subset of that idea and is specialized for a single site family: WeChat public articles. It rejects non-WeChat URLs, parses WeChat-specific metadata variables and DOM containers, and includes handling for WeChat-specific access restrictions. While this is still related to webpage content extraction, the actual scope, target resources, and workflow are materially narrower and different from the declared general-purpose capability.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/render_page.py (reported line 74)May include surrounding context.

python
def node_env_with_global_modules() -> dict:
    env = os.environ.copy()
    root = npm_global_root()
    if root:
        existing = env.get("NODE_PATH", "")

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill invokes network access, shell execution, and optional file writes, but it does not declare any explicit tool scope or permission boundaries. That creates a real least-privilege and review gap: a caller or platform may treat the skill as harmless documentation while it actually enables outbound fetches, browser execution, and persistence behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script accepts a user-supplied cookie and forwards it directly in an HTTP request to the target WeChat domain. While this is likely intended to access pages gated by login or anti-bot controls, it creates a credential-handling risk because sensitive session material may be sent, exposed in logs, shell history, or reused without clear warning or scoping safeguards.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/page_fetch.py (reported line 29)May include surrounding context.

python
if cookie:
        cmd.extend(["--cookie", cookie])

    run = subprocess.run(cmd, capture_output=True, text=True)
    stdout = (run.stdout or "").strip()
    stderr = (run.stderr or "").strip()
    if not stdout:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The runner accepts a raw cookie value and forwards it to child fetch scripts for WeChat access, but provides no safeguards around secret handling, redaction, or storage in process arguments. Command-line arguments can be exposed via process listings, logs, crash reports, or orchestration telemetry, potentially leaking session credentials.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
93% confidence
Finding

The browser-render fallback is invoked with timeout=0, which disables the timeout parameter entirely in run_json and can allow the child rendering process to run indefinitely. A malicious or pathological page could hang rendering, consume CPU/memory, or tie up worker capacity, creating a denial-of-service condition in an agent environment that fetches untrusted URLs.

Content

Scanner excerpt · scripts/page_fetch.py (reported line 98)May include surrounding context.

python
code, payload = run_json(FETCH_PAGE, args.url, timeout=args.timeout, max_chars=args.max_chars)
            if not is_good_enough(payload):
                steps.append("wechat-fallback-browser")
                code, payload = run_json(RENDER_PAGE, args.url, timeout=0, max_chars=args.max_chars, wait_ms=args.wait_ms)
    else:
        steps.append("general-first")
        code, payload = run_json(FETCH_PAGE, args.url, timeout=args.timeout, max_chars=args.max_chars)

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
93% confidence
Finding

This second browser fallback path also disables timeout enforcement by passing timeout=0, allowing arbitrary pages to hold the renderer open without a deadline. In this skill context, fetching attacker-controlled URLs is expected behavior, which makes unbounded execution materially more dangerous because it can be triggered remotely through normal use.

Content

Scanner excerpt · scripts/page_fetch.py (reported line 104)May include surrounding context.

python
code, payload = run_json(FETCH_PAGE, args.url, timeout=args.timeout, max_chars=args.max_chars)
        if not is_good_enough(payload):
            steps.append("general-fallback-browser")
            code, payload = run_json(RENDER_PAGE, args.url, timeout=0, max_chars=args.max_chars, wait_ms=args.wait_ms)

    notes = payload.get("notes")
    if isinstance(notes, list):

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/render_page.py (reported line 66)May include surrounding context.

python
npm = shutil.which("npm")
    if not npm:
        return None
    run = subprocess.run([npm, "root", "-g"], capture_output=True, text=True)
    if run.returncode != 0:
        return None
    root = run.stdout.strip()

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/render_page.py (reported line 86)May include surrounding context.

python
node = find_node()
    if not node:
        return False
    check = subprocess.run(
        [node, "-e", "require.resolve('playwright'); console.log('ok')"],
        capture_output=True,
        text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This skill intentionally performs browser-based retrieval of arbitrary URLs with no apparent guardrails or disclosure around the security implications. In an agent setting, that makes SSRF and unintended access to internal-only web resources more dangerous, especially because a full browser may execute JavaScript, follow redirects, and reach destinations lightweight fetchers might not.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
82% confidence
Finding

The script launches a real browser to fetch an arbitrary user-supplied URL, which can be abused for server-side request forgery against internal services or cloud metadata endpoints if the skill runs in a network-reachable environment. Although the subprocess call itself is not shell-injectable, the browser execution meaningfully expands attack surface by executing remote page content and following redirects without URL restrictions.

Content

Scanner excerpt · scripts/render_page.py (reported line 128)May include surrounding context.

python
return 2

    cmd = [node, "-e", RENDER_SNIPPET, args.url, str(args.wait_ms)]
    run = subprocess.run(cmd, capture_output=True, text=True, env=node_env_with_global_modules())
    if run.returncode != 0:
        payload = {
            "url": args.url,

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The request headers hard-code "Accept-Language": "en-US,en;q=0.9", which imposes a specific language/locale on all fetches. This is a natural-language policy concern because the script does not offer user opt-in or explain why English is required.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The request headers force Language/zh_CN in the User-Agent and Accept-Language: zh-CN,..., imposing a specific locale rather than letting the user choose. This is a natural-language locale policy concern because the skill does not offer any opt-in or configuration for language/region behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

This code can persist the fetched payload to a local file when --save-json is used, but there is no confirmation prompt or runtime disclosure describing that potentially sensitive fetched content will be written to disk. The argument help mentions persistence, but the write operation itself lacks a stronger warning despite handling webpage content that may include private or restricted data.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.