Back to skill

Security audit

Web Crawler

Security checks for vulnerabilities and agentic risk

Overview

This crawler skill has no hidden executable payload, but it openly teaches CAPTCHA, anti-bot, fingerprint, and proxy evasion techniques that require careful review before use.

Review this skill before installing or using it. Use it only for sites you own or have permission to crawl, remove or ignore the CAPTCHA/anti-detection/proxy-bypass guidance, prefer official APIs, honor robots.txt and terms, and install only necessary pinned dependencies in an isolated environment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:625
Finding

Unpinned Third-Party Dependencies Installed from Unverified Sources

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 625–648
Vulnerability Type: Supply-chain exposure through unpinned dependencies
Risk Level: Medium

Vulnerable Code

bash
# Basic
pip install requests beautifulsoup4 lxml pandas openpyxl

# Dynamic
pip install playwright
playwright install chromium

# Anti-detection
pip install DrissionPage undetected-chromedriver fake-useragent

# Framework
pip install scrapy scrapy-redis

# Async
pip install aiohttp httpx

# JS execution
pip install PyExecJS
# Node.js required for PyExecJS

# Image processing
pip install Pillow

# OCR
pip install ddddocr  # Chinese captcha OCR

Technical Analysis

The Skill directs users to install numerous packages from the default Python package index without pinned versions, cryptographic hashes, a lockfile, or package-source restrictions. It also installs a browser binary through Playwright without specifying or validating the expected artifact version.

Python package installation can execute package-controlled build and installation hooks. If a referenced package or transitive dependency is compromised, taken over, replaced, or unexpectedly changed in a future release, installation may execute attacker-controlled code with the privileges of the user running pip.

The risk is increased by the breadth of optional dependencies. Packages providing browser automation, JavaScript execution, CAPTCHA processing, distributed crawling, and anti-detection features are recommended collectively even though many are unnecessary for a typical crawler. This exceeds minimum dependency privileges because users may install high-capability components unrelated to their specific task.

No evidence establishes that any named package is currently malicious. The vulnerability is the unsafe and non-reproducible dependency installation practice.

Attack Path

  1. An attacker compromises a named package, one of its ...[truncated 1575 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace free-form installation commands with a reviewed dependency manifest containing exact versions, such as requirements.txt or pyproject.toml.

  2. Generate a hash-locked requirements file and require hash verification, for example:

    bash
    python -m pip install --require-hashes -r requirements.lock
    
  3. Separate dependencies into minimal optional groups, such as basic, browser, distributed, and ocr, and instruct users to install only the group needed for the requested crawler.

  4. Pin and review all transitive dependencies using a reproducible locking tool such as pip-tools, Poetry, or uv.

  5. Restrict installation to an approved package index or internal mirror and verify package ownership, provenance, release history, and signatures where available.

  6. Install packages in an isolated virtual environment or disposable container using a non-root account and minimal filesystem and network permissions.

  7. Add automated dependency scanning and update review rather than automatically accepting new releases.

  8. Pin the Playwright package and browser revision together, and download browser artifacts only from an approved source with integrity validation.

  9. Avoid including high-capability anti-detection, JavaScript-execution, distributed-crawling, or OCR packages unless the user’s requested functionality specifically requires them.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill explicitly provides CAPTCHA-solving and challenge-bypass guidance, including third-party solving services and techniques for defeating reCAPTCHA, hCaptcha, Turnstile, and slider challenges. In the context of a web-crawling skill, this goes beyond ordinary automation and directly enables circumvention of access controls and anti-abuse protections, increasing the likelihood of unauthorized scraping at scale.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The documentation recommends stealth and evasion tactics such as fingerprint spoofing, disabling webdriver indicators, simulating human behavior, decoy requests, and residential proxy use. These are anti-detection measures intended to avoid bot defenses rather than simply retrieve public content, making the skill materially more dangerous by facilitating abuse and evasion of platform protections.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest trigger list contains generic terms such as "scraper", "crawler", "scrape", and Chinese equivalents like "爬取" and "抓取数据" without any narrowing context or exclusion rules. In a markdown/manifest context, this makes activation scope ambiguous and increases the chance the skill is invoked for casual or unrelated mentions rather than explicit requests for this specific skill.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 104)May include surrounding context.

md
import random

# Config
BASE_URL = "https://api.example.com/data"
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36",
    "Accept": "application/json",

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The Playwright example hard-codes locale="zh-CN", which imposes a specific language/locale behavior regardless of the user's preference. This is a natural-language policy issue because the file does not present locale choice or explain that the setting is only for region-specific targets.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs users to extract and execute target-site JavaScript via ExecJS/Node to reproduce obfuscated signing and encryption logic. In this context, that is practical reverse-engineering of client-side protections used to gate access to APIs, which can help users bypass intended request validation and scrape protected endpoints.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The anti-detection header example sets Accept-Language to prefer zh-CN,zh ahead of English, which enforces a language/locale preference in generated traffic. Because the skill does not offer the user a language choice or limit this behavior to China-specific targets, it violates the locale-choice policy.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.