Back to plugin

Security audit

ScraperAPI Skills

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed ScraperAPI integration bundle, but it should be used deliberately because it sends web queries, URLs, page content, and API credentials to third-party services.

Install only if you intend to use ScraperAPI as a third-party scraping provider. Do not send private URLs, credentials in query strings, internal systems, regulated personal data, or sensitive business content unless you have authorization and a clear data-processing basis. Review lead-enrichment, crawler, DataPipeline, webhook, and n8n scheduling workflows carefully because they can collect personal contact details, run recurring jobs, send results to webhooks, and consume ScraperAPI credits.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (44)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The trigger phrases include broad requests like 'get web data for my agent', 'scrape a website in my app', and 'add web scraping to my project', which can activate this skill in contexts that are not explicitly asking for ScraperAPI. Over-broad invocation increases the chance that user-supplied URLs, queries, and content are routed to a third-party service unexpectedly, creating privacy and data-handling risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

This example sends a user-selected URL and the API key to ScraperAPI, and the skill metadata already states that user-supplied queries, URLs, and content are transmitted externally. In a skill that may be triggered broadly, this creates a real data egress risk because sensitive internal URLs, tokens in URLs, or proprietary content could be forwarded to a third party without sufficiently explicit consent or sanitization.

Content

Scanner excerpt · SKILL.md (reported line 152)May include surrounding context.

md
import os, requests

r = requests.get(
    "https://api.scraperapi.com/",
    params={"api_key": os.environ["SCRAPERAPI_API_KEY"], "url": "https://httpbin.org/ip"}
)
print(r.status_code, r.text[:200])

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The REST examples instruct direct transmission of target URLs and authentication material to an external scraping service. Because the skill is designed to operationalize arbitrary web scraping requests, this can expose sensitive user inputs or cause unintended third-party processing when users provide confidential targets or data-bearing URLs.

Content

Scanner excerpt · SKILL.md (reported line 214)May include surrounding context.

md
curl "https://api.scraperapi.com/?api_key=$SCRAPERAPI_API_KEY&url=https://example.com"

# With JS rendering
curl "https://api.scraperapi.com/?api_key=$SCRAPERAPI_API_KEY&url=https://example.com&render=true"

# Structured data — Google SERP
curl "https://api.scraperapi.com/structured/google/search?api_key=$SCRAPERAPI_API_KEY&query=web+scraping"

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This example shows externally sending rendered-page requests and structured search queries to ScraperAPI, which may include user intent, search terms, and browsing targets. In combination with the broad triggers and onboarding role, the skill can normalize third-party transmission of potentially sensitive prompts or investigative queries that users may not realize are leaving the local environment.

Content

Scanner excerpt · SKILL.md (reported line 217)May include surrounding context.

md
curl "https://api.scraperapi.com/?api_key=$SCRAPERAPI_API_KEY&url=https://example.com&render=true"

# Structured data — Google SERP
curl "https://api.scraperapi.com/structured/google/search?api_key=$SCRAPERAPI_API_KEY&query=web+scraping"

# Structured data — Amazon product
curl "https://api.scraperapi.com/structured/amazon/product?api_key=$SCRAPERAPI_API_KEY&asin=B09V3KXJPB"

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The structured Amazon and async job examples continue the pattern of sending user-controlled inputs and API credentials to external endpoints for processing. While expected for this product, it is still a true security concern because the skill facilitates broad data egress and could be used to transmit sensitive URLs, job payloads, or business data to a third party at scale.

Content

Scanner excerpt · SKILL.md (reported line 220)May include surrounding context.

md
curl "https://api.scraperapi.com/structured/google/search?api_key=$SCRAPERAPI_API_KEY&query=web+scraping"

# Structured data — Amazon product
curl "https://api.scraperapi.com/structured/amazon/product?api_key=$SCRAPERAPI_API_KEY&asin=B09V3KXJPB"

# Submit an async job
curl -X POST "https://async.scraperapi.com/jobs" \

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This duplicate finding points to the same outbound POST to ScraperAPI carrying credentials and target request data. Even though the transmission is expected for a scraping integration, it is still a genuine security-relevant behavior because it exports user-controlled data and authentication material to an external service.

Content

Scanner excerpt · SKILL.md (reported line 57)May include surrounding context.

md
API_KEY = os.environ["SCRAPERAPI_API_KEY"]

# Submit
r = requests.post(
    "https://async.scraperapi.com/jobs",
    json={
        "apiKey": API_KEY,

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This duplicate finding points to the same outbound POST to ScraperAPI carrying credentials and target request data. Even though the transmission is expected for a scraping integration, it is still a genuine security-relevant behavior because it exports user-controlled data and authentication material to an external service.

Content

Scanner excerpt · SKILL.md (reported line 57)May include surrounding context.

md
API_KEY = os.environ["SCRAPERAPI_API_KEY"]

# Submit
r = requests.post(
    "https://async.scraperapi.com/jobs",
    json={
        "apiKey": API_KEY,

YARA rule 'agent_skill_credential_exfiltration_webhook': AI agent skill credential harvesting followed by webhook or external exfiltration [agent_skills]

Critical
Category
YARA Match
Confidence
85% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.md (reported line 54)May include surrounding context.

python
import os, requests, time

API_KEY = os.environ["SCRAPERAPI_API_KEY"]

# Submit
r = requests.post(

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger scope is broad enough to activate on generic research or contact-finding requests, which increases the chance the skill runs in situations where users did not explicitly intend third-party scraping or personal-data enrichment. Because the skill then performs external searches and profile synthesis automatically, over-triggering can cause unnecessary collection and disclosure of personal information.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly instructs the model to gather and present personal contact details, including emails, phone numbers, locations, and social profiles, into a consolidated dossier. This creates a doxxing-style enrichment workflow that can facilitate targeted phishing, harassment, or unauthorized profiling even when data is scraped from public sources.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Using an email address as a seed to derive a person's identity, employer, and related company details materially increases privacy risk because a single identifier can be expanded into a fuller profile without the subject's consent. This supports deanonymization and targeted social-engineering workflows, especially when combined with search and scraping across multiple sources.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The synthesis step turns scattered public and scraped data into a structured contact card, which lowers the effort needed to operationalize personal information for phishing, impersonation, or surveillance. Aggregation meaningfully raises the sensitivity of the data beyond the risk of any single source because it produces a ready-made intelligence profile.

Content

No source excerpt is available for this finding.

Unrestricted Tool Access

Medium
Category
Excessive Agency
Confidence
83% confidence
Finding

The skill authorizes broad external search and arbitrary URL fetching driven by discovered links, with no meaningful allowlist or destination restrictions. That enables exfiltration of user-supplied queries and URLs to third parties and increases the risk of collecting data from unexpected or sensitive sources during automated browsing.

Content

Scanner excerpt · SKILL.md (reported line 44)May include surrounding context.

md
Run searches by *what you're looking for*, not by which site to target. Google will surface whatever sources exist — company website, Crunchbase, Wikipedia, news, directories, LinkedIn, G2, etc. Collect all promising URLs from `organic_results[].link` and carry them into Phase 3.

**Search tool:** Call `mcp__ScraperAPI__google_search` with `query`, `num: 10`, and `countryCode: "us"`. Read snippets carefully — they often contain the data you need without an extra fetch.

### 2a. Person name as seed

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The trigger text is broad enough to match common informational requests like general market questions, pricing questions, or trend questions, which can cause the agent to invoke this skill when the user did not explicitly intend external web research. Because the skill transmits user-supplied queries, URLs, and content to ScraperAPI, over-triggering increases unnecessary data exfiltration and can lead to privacy and compliance issues.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger scope is unusually broad, covering generic activities like web scraping, research, shopping, SEO, market research, and product lookup. That can cause the skill to activate for many ordinary user requests and route user queries, URLs, and page content to a third-party scraping service, increasing the chance of unnecessary data exposure, unexpected tool use, and bypass of safer default web tools.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
75% confidence
Finding

The Maps guidance explicitly recommends always supplying precise latitude/longitude for local queries but provides no caution about collecting, transmitting, or retaining sensitive geolocation data. In this skill context, user-supplied queries and coordinates are sent to a third-party service, which increases privacy risk and can expose exact user or target locations unnecessarily.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guide explicitly promotes scraping that 'handles proxy rotation, CAPTCHAs, and anti-bot measures automatically' without any caution about authorization, terms-of-service, privacy, or legal constraints. In this skill context, that omission materially increases misuse risk because the tool is positioned for broad scraping tasks and transmits user-supplied URLs and queries to a third party, making it easier for users to bypass protections without informed consent or compliance checks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The documentation recommends geo-targeting and premium proxy use but does not warn that requests may traverse third-party proxy infrastructure or different jurisdictions, which can create legal, compliance, and data-handling exposure. Given the skill metadata explicitly states that user-supplied queries, URLs, and content are transmitted to ScraperAPI, this omission is more dangerous because users may unknowingly send regulated or sensitive data through external networks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The webhook workflow takes user-supplied company input and sends it directly to ScraperAPI/Google Search, which is a third-party service, without any built-in notice, consent check, or data-minimization guidance in the example itself. While the skill metadata notes that user-supplied queries are transmitted externally, someone copying this workflow may expose user input to an external processor without realizing the privacy and compliance implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The crawler configuration requires a crawlerCallbackUrl where ScraperAPI will POST crawl results, which means scraped page contents and metadata are transmitted to an externally reachable endpoint. The reference documents the mechanics but does not clearly warn about this data flow, increasing the risk that users unknowingly send sensitive, proprietary, or regulated data off-platform or to an exposed webhook.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

The skill explicitly instructs use of external services and includes a direct API fetch from api.n8n.io, while the skill metadata also states that user-supplied queries, URLs, and content are transmitted to ScraperAPI. In this context, generated workflows may send user-controlled or sensitive data to third-party infrastructure without strong minimization, consent, or trust-boundary warnings at the point of use.

Content

Scanner excerpt · SKILL.md (reported line 514)May include surrounding context.

md
- Node listing: `https://n8n.io/integrations/scraperapi/`
- Price-drop digest template (page): `https://n8n.io/workflows/15609-send-daily-price-drop-digest-emails-for-amazon-walmart-and-google-via-scraperapi/`
- Price-drop digest template (API, returns full workflow JSON): `https://api.n8n.io/api/templates/workflows/15609`
- npm package: `https://www.npmjs.com/package/n8n-nodes-scraperapi-official`
- ScraperAPI docs: `https://docs.scraperapi.com/`
- Dashboard: `https://dashboard.scraperapi.com/`

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
80% confidence
Finding

Code accesses environment variables that may contain secrets (API keys, tokens). This is a common pattern for credential theft.

Content

Scanner excerpt · SKILL.md (reported line 251)May include surrounding context.

md
import os, requests
from scraperapi_sdk import ScraperAPIClient

client = ScraperAPIClient(os.environ["SCRAPERAPI_API_KEY"])

def scrape(url, params=None):
    try:

Tainted flow: 'SCRAPERAPI_API_KEY' from os.environ.get (line 26, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/research_agent.py (reported line 66)May include surrounding context.

python
"""Search Google via ScraperAPI structured endpoint."""
    print(f"  Searching: {query!r}")
    try:
        resp = requests.get(
            "https://api.scraperapi.com/structured/google/search",
            params={
                "api_key": SCRAPERAPI_API_KEY,

Tainted flow: 'SCRAPERAPI_API_KEY' from os.environ.get (line 26, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/research_agent.py (reported line 114)May include surrounding context.

python
if _should_skip(url):
        return None
    try:
        resp = requests.get(
            "https://api.scraperapi.com/",
            params={
                "api_key": SCRAPERAPI_API_KEY,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill clearly requires environment variables, writes output files, and performs network requests, yet it does not declare explicit permissions for those capabilities. This creates a transparency and policy-enforcement gap: users or orchestrators may invoke the skill without understanding that sensitive queries and scraped content will be transmitted externally and persisted locally.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.