Back to skill

Security audit

Crawler

Security checks for vulnerabilities and agentic risk

Overview

This is a documentation-only web crawling reference; it includes some anti-blocking advice to use cautiously, but I found no hidden execution, data collection, or persistence.

Install this as a reference skill only if you are comfortable with scraping guidance that includes anti-blocking techniques. Use it for lawful, permissioned crawling, prefer official APIs, follow site terms and robots.txt, and avoid using the proxy, fingerprint, CAPTCHA, or cookie-rotation advice to bypass a site's access controls.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Behavior Manipulation

Medium
Category
Prompt Injection
Confidence
70% confidence
Finding

Subtle instructions detected that may alter agent decision-making or introduce hidden biases.

Content

Scanner excerpt · scripts/script.sh (reported line 200)May include surrounding context.

sh
- Always check ToS before scraping

## Ethical Scraping Guidelines
  1. Check for API first (always prefer official API over scraping)
  2. Read and follow robots.txt
  3. Rate limit your requests (minimum 1s between requests per domain)
  4. Identify your crawler (descriptive User-Agent with contact info)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The content explicitly recommends anti-detection measures such as rotating proxies, rotating headers, and varying TLS fingerprints to avoid blocking. In a web-scraping reference skill, this crosses from neutral education into operational guidance for evading site defenses, which can facilitate abusive or unauthorized scraping.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/script.sh (reported line 302)May include surrounding context.

sh
//div[contains(@class,"x")] Class contains
    //table/tr[position()>1]  Skip header row

## curl with Cookie Jar
  curl -c cookies.txt -b cookies.txt -L "https://example.com/login"   -d "user=x&pass=y"
  curl -b cookies.txt "https://example.com/data"

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The FAQ gives step-by-step advice on how to avoid getting blocked, including rotating user agents, using residential proxies, matching browser fingerprints, and handling CAPTCHAs. That materially enables evasive scraping behavior and bypass of site-imposed defenses, making the skill more dangerous despite its reference-style framing.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.