Back to skill

Security audit

Scrapeless LLM Chat Scraper Skill

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward Scrapeless API wrapper for querying supported AI chat services, with privacy and supply-chain cautions but no artifact-backed malicious behavior.

Install only if you are comfortable sending prompts and scraper results through Scrapeless. Do not include secrets, regulated data, private customer information, or confidential prompts; treat X_API_TOKEN as a secret and keep .env out of source control. Prefer pinned dependencies or a lockfile, and consider removing prompt-content logging before use in shared terminals, CI, or centralized logs.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/llm_chat_scraper.py:107
Finding

Sensitive Prompt Content Exposed Through INFO-Level Logging

Content
View full analysis

Vulnerability Details

File Location: scripts/llm_chat_scraper.py, line 107
Vulnerability Type: Sensitive information exposure through application logs
Risk Level: Medium

Vulnerable Code

python
logger.info(f"Creating task for {actor} with prompt: {prompt[:50]}...")

Technical Analysis

The Skill records the first 50 characters of every user prompt at INFO level. Prompt content can contain API credentials, personal information, proprietary material, internal system details, or other confidential data.

Logging prompt text is not required to create or retrieve a Scrapeless API task. Because INFO logging is enabled globally, the disclosure occurs during normal operation rather than only in an explicitly enabled debugging mode. The resulting text may be retained in terminal transcripts, CI/CD output, centralized logging systems, Agent execution records, or process supervisors.

The exposure is limited to the first 50 characters of the prompt, but secrets and identifying information commonly occur at the beginning of a prompt. This issue does not directly grant system privileges or permit code execution.

Attack Path

  1. A user or calling Agent supplies a prompt containing confidential information within its first 50 characters.
  2. The Skill passes that prompt to create_task().
  3. Before the network request is made, the INFO-level logging statement writes the prompt fragment to the configured logging destination.
  4. A user, service, or attacker with access to retained execution logs reads the disclosed prompt fragment.
  5. If the fragment contains a usable credential or sensitive business information, it can be misused outside the Skill.

Exploitation therefore requires the ability to cause sensitive content to be submitted and access to the resulting logs.

Impact Assessment

The primary impact is loss of confidentiality. Depending on prompt contents and log distribution, exposure may ...[truncated 262 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove prompt content from routine logs:
    python
    logger.info("Creating task for actor: %s", actor)
    
  • Log only non-sensitive operational metadata, such as the actor name and server-generated task identifier.
  • If request correlation is necessary, use an opaque locally generated correlation ID rather than prompt text.
  • Do not use an ordinary unkeyed prompt hash as a secrecy mechanism because predictable prompts may be recoverable through guessing.
  • Make any diagnostic content logging explicitly opt-in, disabled by default, and protected by redaction and retention controls.
  • Review existing log storage and retention policies for previously captured prompt fragments.

T08 · Insecure Dependencies

Note
Location
requirements.txt:1
Finding

Unbounded Dependency Versions Create Non-Reproducible Supply-Chain Exposure

Content
View full analysis

Vulnerability Details

File Location: requirements.txt, lines 1-2
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Low

Vulnerable Code

text
requests>=2.31.0
python-dotenv>=1.0.0

Technical Analysis

Both dependencies use lower-bound-only version constraints. Consequently, a future installation may resolve to any later release without that release having been reviewed with this Skill. This makes installations non-reproducible and permits dependency behavior to change after the Skill itself has been audited.

The package names and configured source shown in the project do not demonstrate dependency confusion, typosquatting, or a currently malicious package. The risk is prospective: a compromised, malicious, or incompatible future release accepted by these broad constraints could affect installation or runtime security.

Attack Path

  1. A user follows the documented installation process and runs pip install -r requirements.txt.
  2. The package resolver queries the configured Python package index and selects versions satisfying the lower bounds.
  3. A later dependency version that was not reviewed with the Skill is selected.
  4. Package installation logic or imported runtime code executes with the permissions of the installing or running user.
  5. If that selected release has been compromised, it could access data and resources available to that user, including environment variables such as X_API_TOKEN.

This path is conditional on compromise or malicious modification of an accepted future dependency release; no such compromise was established during the static audit.

Impact Assessment

A compromised dependency could execute code with the installer or runtime user's privileges. Potentially exposed resources include project files, environment variables, the Scrapeless API token, network access, and any other resources available to that account. The present finding does ...[truncated 107 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin each dependency to a reviewed exact version.
  • Generate and commit a lock file containing transitive dependency versions.
  • Require package hashes during installation, for example by using a hash-locked requirements file with pip --require-hashes.
  • Install packages only from explicitly trusted indexes.
  • Automate vulnerability and integrity scanning for direct and transitive dependencies.
  • Update pinned versions through a controlled process that includes code review, compatibility testing, and security review.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (16)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 95)May include surrounding context.

  1. Create a .env file in the root directory based on the .env.example file:
bash
cp .env.example .env
  1. Add your Scrapeless API token to the .env file:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/llm_chat_scraper.py (reported line 25)May include surrounding context.

python
)
logger = logging.getLogger(__name__)

# Load environment variables from .env file
load_dotenv()

try:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/llm_chat_scraper.py (reported line 54)May include surrounding context.

python
)
logger = logging.getLogger(__name__)

# Load environment variables from .env file
load_dotenv()

try:

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · scripts/llm_chat_scraper.py (reported line 237)May include surrounding context.

python
self.poll_for_result(
                task_id=task_id,
                interval=poll_interval,
                max_retries=max_retries
            )
            
        except Exception as e:
            return {
                'success': False,
                'error': type(e).__name__,
                'message': str(e)
            }


def main():
    parser = argparse.ArgumentParser(
        description='LLM Chat Scraper Tool',
        formatter_class=argparse.RawDescriptionHelpFormatter,
        epilog="""
Examples:
  python3 llm_chat_scraper.py chatgpt --query "AI trends"
  python3 llm_chat_scraper.py gemini --query "Best restaurants" --country US
  python3 llm_chat_scraper.py perplexity --query "Latest news" --web-search
  python3 llm_chat_scraper.py copilot --query "Explain ML" --mode reasoning
  python3 llm_chat_scraper.py aimode --query "Programming tips"
  python3 llm_chat_scraper.py grok --query "Quantum physics" --mode MODEL_MODE_EXPERT
        """
    )
    
    subparse

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README promotes scraping AI chat conversations, including proxy-based collection and structured extraction of model outputs, but provides no warning about privacy, consent, platform terms, or transmission of potentially sensitive prompts/responses to a third-party service. In this context, users may deploy the skill against accounts or chats containing confidential data without understanding the legal and data-handling risks, which increases the chance of unsafe use.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares environment, network, and shell capabilities but does not constrain them with an explicit permission or allowed-tools scope. In practice, this gives the skill broader operational authority than users may expect, increasing the risk of unintended command execution, unrestricted outbound requests, or misuse of the API token by downstream code.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs users to submit arbitrary prompts to a third-party scraping API but does not clearly warn that those prompts, and potentially returned conversation data, are transmitted off-platform. This can cause accidental disclosure of sensitive personal, organizational, or proprietary information because users may assume the interaction is local or governed by the host agent's normal privacy expectations.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

This code intentionally transmits prompt data to an external API endpoint. In a security review of an agent skill, that is a meaningful data-exposure surface because prompts may contain sensitive information and the tool's purpose is to forward them to a remote service for processing.

Content

Scanner excerpt · scripts/llm_chat_scraper.py (reported line 120)May include surrounding context.

python
mode=mode
        )
        
        response = requests.post(
            f"{self.api_base_url}{self.request_endpoint}",
            headers=self.headers,
            json=payload,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The tool sends user-supplied prompts to a third-party scraping service without an explicit consent boundary or strong user-facing disclosure at the point of transmission. In an agent-skill context, users may enter secrets, internal data, or regulated content into prompts, making this an avoidable data exfiltration/privacy risk.

Content

No source excerpt is available for this finding.

Tainted flow: 'task_id' from requests.get (line 129, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/llm_chat_scraper.py (reported line 142)May include surrounding context.

python
raise RuntimeError(f"Unexpected status code {response.status_code}: {response.text}")
    
    def get_task_result(self, task_id: str) -> Optional[Dict[str, Any]]:
        response = requests.get(
            f"{self.api_base_url}{self.result_endpoint}/{task_id}",
            headers=self.headers,
            timeout=30

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The setup instructions tell users to place an API token in a local .env file but do not warn that the token is sensitive, should not be committed to source control, and should be stored using standard secret-handling practices. While this is common developer workflow, omission of basic credential hygiene guidance can lead to accidental exposure through repos, logs, or shared environments.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency specifier uses a lower bound only (requests>=2.31.0), which allows future versions to be installed without review and prevents reproducible builds. This creates supply-chain risk and makes it impossible to verify whether a deployed version includes known security fixes or newly introduced vulnerabilities.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31.0
python-dotenv>=1.0.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
89% confidence
Finding

requests has known advisories, and because the manifest does not pin a specific version, there is no way to verify whether the installed release is affected. In a skill that scrapes AI chat platforms and likely performs authenticated HTTP requests, using an affected requests version could expose credentials, session data, or transport security guarantees.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency specifier uses a lower bound only (python-dotenv>=1.0.0), so installations may resolve to different unreviewed versions over time. This weakens build reproducibility and increases supply-chain exposure because security posture depends on whatever version happens to be installed.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.31.0
python-dotenv>=1.0.0

Unverifiable Dependency: python-dotenv has 2 known advisory(ies) (CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
83% confidence
Finding

python-dotenv has known advisories, and the unpinned manifest prevents confirming whether a safe version will be installed. If the skill reads or writes environment configuration in unsafe contexts, an affected release could contribute to file overwrite or environment-manipulation issues.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

Defaulting the country/locale to US without explicit user choice can cause privacy, compliance, or behavioral mismatches, especially if locale affects routing, content generation, or legal jurisdiction. While not a severe exploit by itself, it can surprise users and weaken informed consent around how requests are processed.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.