T09 · Insecure Skill Coding Practices
- Location
scripts/scrapling_scrape.py:43- Finding
Unrestricted User-Controlled URL Fetching Enables Server-Side Request Forgery
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This appears to be a legitimate scraping helper, but it needs review because it enables broad network scraping, anti-bot bypass, proxy workflows, and local file effects without tight scoping.
Install only in an isolated, unprivileged environment with outbound network controls. Use it only for targets you are authorized to scrape, avoid stealth/proxy/login modes unless explicitly permitted, and add allowlists or private-network blocking before letting an agent run its scripts or MCP tools automatically.
scripts/scrapling_scrape.py:43Unrestricted User-Controlled URL Fetching Enables Server-Side Request Forgery
SKILL.md:13Unpinned Third-Party Package and Browser Artifact Installation
The skill clearly guides network-capable actions through MCP and scraping commands, but it declares no explicit tool scope such as permissions or allowed-tools. In an agent environment, that ambiguity can enable overbroad invocation or unintended network use, especially because the skill includes anti-bot, stealth, proxy rotation, and login-session guidance that expands operational risk.
The description uses broad language like 'Use this skill' for web scraping strategy and execution guidance without tight trigger conditions or safety boundaries. That can cause agents to apply the skill too generally, including in contexts involving protected sites, anti-bot evasion, session handling, or authentication flows, increasing the chance of policy violations or misuse.
The documentation explicitly exposes fetch_stealthy functionality with Cloudflare-bypass behavior, which is materially more sensitive than ordinary scraping guidance. If this capability is not clearly disclosed in the skill's declared behavior, users and agents may invoke evasive collection features without appropriate review, making misuse and policy circumvention more likely.
The skill documents stealth fetching and Cloudflare-solving as normal usage without any warning about authorization, legal constraints, terms-of-service violations, or operational impact on target sites. Presenting anti-bot evasion as a routine recipe lowers the barrier to potentially unauthorized access and can facilitate abusive scraping against protected services.
Although described as a guidance-oriented skill, this file provides concrete MCP commands and parameters for running live scraping, extraction, and crawling operations, including writing crawl output to disk. That mismatch can cause operators or downstream systems to grant the skill more trust or fewer safeguards than warranted, increasing the risk of unintended execution and data collection.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
"""Scrape via API endpoints instead of HTML."""
name = "api_products"
api_base = "https://api.example.com/v1"
def start_requests(self):
for page in range(1, 100):
The manifest describes this skill as providing MCP-native guidance, recipes, and best practices for using a Scrapling server, which justifies scraping and crawling examples. However, the recipes also instruct saving downloaded media to local files (L203-L207) and deleting checkpoint directories with shutil.rmtree (L221-L222), which are local filesystem modification capabilities not clearly justified by a guidance-focused scraping skill description.
The markdown includes a reset example that deletes the crawl checkpoint directory with shutil.rmtree("./crawl_checkpoint"). There is no accompanying warning that this permanently removes local crawl state or may discard resumable progress, which is the kind of destructive operation the markdown warning rule covers.
The media downloader recipe writes response bodies to ./downloads via open(path, 'wb'), but the surrounding markdown does not warn users that running the example will create files on disk. Because this affects local data and storage, a short disclosure is expected in markdown descriptions.
This code makes outbound requests to user-supplied URLs via FetcherSession, DynamicSession, and StealthySession. While the script purpose implies scraping, there is no explicit user-facing warning, log, or comment disclosing that network access will occur or that remote sites will be contacted.
This code performs network requests to each URL provided by the user via the selected fetcher. Although it prints status after the request, there is no explicit warning or disclosure in the script's help text or docstring that running it will make outbound HTTP requests to remote sites.
No suspicious patterns detected.