Back to skill

Security audit

Deep Scraper

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed browser-based scraping skill, but it accepts arbitrary URLs without meaningful network or destination limits, which could expose private/internal services reachable from the container.

Review before installing. Only run this skill against URLs you are authorized to scrape, and avoid using it in environments where the container can reach private services, cloud metadata endpoints, intranet sites, or sensitive local web apps. Prefer adding strict URL allowlists, private-network blocking, redirect validation, pinned dependencies, and a complete pinned Docker build before use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
assets/main_handler.js:10
Finding

Unrestricted URL Navigation Enables Server-Side Request Forgery

Content
View full analysis
document.body.innerText); ``` ```js crawler.run([targetUrl]); ``` `assets/youtube_handler.js:3-4, 33, 90`: ```js const targetUrl = process.argv[2]; const videoId = targetUrl.split('v=')[1]?.split('&')[0]; ``` ```js await page.goto(targetUrl, { waitUntil: 'networkidle' }); ``` ```js crawler.run([targetUrl]); ``` ### Technical Analysis Both handlers accept a URL directly from a command-line argument and pass it to Crawlee and Playwright without validating its scheme, hostname, resolved IP address, or redirect destination. In `main_handler.js`, a URL is treated as generic unless its text contains `youtube.com`. Generic pages are opened and their title and visible body text are returned to stdout. Consequently, an attacker who can control the command-line URL can direct the browser toward HTTP services accessible from the container, including loopback interfaces, private network addresses, link-local services, and potentially cloud instance metadata endpoints. Validating only the original textual hostname would not be sufficient. A robust defense must account for DNS rebinding, IPv4 and IPv6 representations, redirects to prohibited destinations, and hostnames that resolve to private or reserved addresses. The browser is additionally launched with `--no-sandbox` and `--disable-setuid-sandbox` in `assets/main_handler.js:23` ...[truncated 2224 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
Findings (10)

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · assets/main_handler.js (reported line 31)May include surrounding context.

js
async requestHandler({ page, log }) {
        log.info(`Deep-Scraper starting in ${mode} mode for: ${targetUrl}`);
        
        // Clear context to ensure a fresh session (avoid cache leakage)
        const context = page.context();
        await context.clearCookies();

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly performs live scraping through a Dockerized browser stack against third-party sites, but it does not clearly warn users that supplied target URLs and fetched content will be transmitted over the network and processed inside a containerized automation environment. This can lead to unintended collection, transmission, and handling of sensitive URLs or data, especially when operators assume the skill is purely local or passive.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The script launches a full browser, intercepts all requests, and forcibly extracts transcript-related traffic from arbitrary user-supplied YouTube URLs. That broad scraping behavior increases data collection beyond the minimum needed and could be repurposed to capture additional request metadata or content without clear constraints, making it a real security/privacy concern even if the apparent goal is transcript retrieval.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill retrieves transcript XML and, on failure, falls back to collecting visible page description text, but provides no user-facing disclosure or consent mechanism about this network-driven content extraction. In an agent context, silent retrieval of page content from a supplied URL can violate user expectations and make unintended data collection harder to detect or audit.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

Low
Category
Tool Misuse
Confidence
15% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 16)May include surrounding context.

Standard Interface (CLI)

bash
docker run -t --rm -v $(pwd)/skills/deep-scraper/assets:/usr/src/app/assets clawd-crawlee node assets/main_handler.js [TARGET_URL]

Output Specification (JSON)

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Multiple log and console messages are written in Chinese, which imposes a specific language on users without opt-in or explanation. This is a natural-language policy concern because the file does not indicate that the skill is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
91% confidence
Finding

The dependency uses a caret range, which allows newer minor/patch releases to be installed without explicit review. This creates supply-chain risk and reduces build reproducibility, especially for a scraper that relies on browser automation and external downloads.

Content

Scanner excerpt · package.json (reported line 19)May include surrounding context.

json
"author": "Joseph",
  "license": "MIT",
  "dependencies": {
    "crawlee": "^3.0.0",
    "playwright": "^1.40.0"
  },
  "openclaw": {

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The Playwright dependency is not pinned to an exact version, so installations may resolve to different releases over time. For software that drives browsers inside containers, this increases exposure to upstream package compromise or introduction of insecure behavior through automatic version drift.

Content

Scanner excerpt · package.json (reported line 20)May include surrounding context.

json
"license": "MIT",
  "dependencies": {
    "crawlee": "^3.0.0",
    "playwright": "^1.40.0"
  },
  "openclaw": {
    "requires": {

Unverifiable Dependency: playwright has 1 known advisory(ies) (CVE-2025-59288 (Playwright downloads and installs browsers without verifying the authenticity of)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

The manifest references Playwright without pinning an exact version, while Playwright has a known advisory involving downloading and installing browsers without authenticity verification. Because the resolved version is not fixed, it is impossible to determine from this file whether deployments will pull an affected release, which is more concerning in a containerized scraper that automates browser setup.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.