Back to skill

Security audit

openclaw-serper

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed web-search tool, but it automatically fetches unvalidated result pages in a way that could reach internal network resources from the user's environment.

Review before installing in environments with access to internal web services, cloud metadata endpoints, or private networks. Use a restricted network sandbox or require URL validation before allowing automatic full-page extraction. Avoid entering secrets or sensitive internal queries, and prefer a pinned dependency installation in an isolated virtual environment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/search.py:89
Finding

Unvalidated Search Result URLs Permit Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/search.py:89-96 and scripts/search.py:190-193
Vulnerability Type: Server-Side Request Forgery through unvalidated remote URL fetching
Risk Level: High

Vulnerable Code:

python
def _extract_content(url: str) -> Optional[str]:
    """Fetch a URL and extract clean readable text using trafilatura."""
    try:
        downloaded = trafilatura.fetch_url(url, config=_traf_config)
        if not downloaded:
            return None
        return trafilatura.extract(downloaded, include_links=False, include_images=False,
                                   include_tables=True, deduplicate=True) or None
    except Exception:
        return None
python
futures = {}
pool = ThreadPoolExecutor(max_workers=max(1, len(results)))
for i, r in enumerate(results):
    if r.get("url"):
        futures[i] = pool.submit(_extract_content, r["url"])

Technical Analysis

URLs supplied by the Serper search response are passed directly to trafilatura.fetch_url() without validation. The implementation does not:

  • Restrict URLs to HTTP and HTTPS.
  • Resolve hostnames and reject loopback, private, link-local, reserved, or multicast addresses.
  • Block cloud instance metadata endpoints.
  • Revalidate destinations after HTTP redirects.
  • Enforce an explicit response-body size limit.

Search results constitute remotely influenced input. An attacker may operate an indexed page or compromise a page that appears in search results. That page can redirect the content-fetching client to an internal service. If the underlying fetching library follows redirects, the Skill may issue requests that the user could not issue directly from outside the execution environment.

The API credential is not attached to these page-fetching requests; it is only sent to the hard-coded Serper API endpoint. Consequently, this issue does not directly disclose the Serper cred ...[truncated 1591 chars]

Remediation
View remediation

Remediation Suggestions

  1. Accept only http and https URLs and reject URLs containing credentials or malformed hostnames.
  2. Resolve the destination hostname before connecting and reject every address in loopback, private, link-local, reserved, multicast, and unspecified ranges for both IPv4 and IPv6.
  3. Disable automatic redirects or validate the scheme, hostname, DNS resolution, and IP address after every redirect.
  4. Explicitly block common metadata destinations, including 169.254.169.254, even where link-local filtering is already present.
  5. Protect against DNS rebinding by ensuring that validation and connection use the same resolved public address.
  6. Set strict connection, read, total-operation, and response-size limits.
  7. Run page extraction in a sandbox with outbound network policy restricted to public Internet destinations.
  8. Log rejected destinations without logging credentials or sensitive response data.
  9. Add tests covering direct private addresses, encoded IP representations, IPv6 loopback, DNS names resolving to private addresses, and redirect chains to prohibited destinations.

T08 · Insecure Dependencies

Warning
Location
README.md:23
Finding

Unpinned Third-Party Dependency Creates Supply-Chain Exposure

Content
View full analysis

Vulnerability Details

File Location: README.md:23-38
Vulnerability Type: Unpinned dependency installation from a mutable package source
Risk Level: Medium

Vulnerable Code:

bash
# Install for your user
pip install --user trafilatura

# Or if you use pip3 explicitly
pip3 install --user trafilatura

The Skill metadata also recommends an unconstrained installation:

yaml
compatibility: Requires Python 3, trafilatura (pip install trafilatura), and network access.

Technical Analysis

The installation instructions retrieve the latest available trafilatura release and its transitive dependencies without version constraints or integrity hashes. The resolved code can therefore change after this Skill has been audited.

Although no malicious dependency is present in the audited repository itself, the instructions establish a mutable external supply-chain trust boundary. A compromised package account, malicious future release, dependency compromise, or incompatible update could cause users to install code that was not reviewed with the Skill.

Python packages may execute code during installation, and imported package code executes with the privileges of the Python process. The use of --user avoids system-wide installation but still modifies the user’s Python environment and permits package code to run with that user’s permissions.

Attack Path

  1. A maintainer account or a relevant package in the dependency graph is compromised, or a malicious release is published.
  2. A user follows the documented pip install --user trafilatura instruction.
  3. Pip resolves the current mutable release and transitive dependency set.
  4. The package executes installation-time logic or is subsequently imported by scripts/search.py.
  5. Malicious code executes with the permissions of the installing or invoking user.

This path requires compromise or malicious modification of the upstream package ...[truncated 692 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin trafilatura to a specific version that has been reviewed and tested.

  2. Pin all transitive dependencies using a lock file generated by a tool such as pip-tools.

  3. Publish cryptographic hashes and install with pip --require-hashes.

  4. Provide a checked-in requirements file, for example:

    text
    trafilatura==<reviewed-version> --hash=sha256:<verified-hash>
    
  5. Recommend installation inside a dedicated virtual environment rather than the shared user Python environment.

  6. Use automated dependency vulnerability and release monitoring.

  7. Review and deliberately update pinned versions instead of accepting arbitrary future releases.

  8. Keep the dependency source restricted to the official package index or an organization-controlled mirror.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (13)

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'network' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 54)May include surrounding context.

md
# =============================================================================
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 181)May include surrounding context.

md
# =============================================================================
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 182)May include surrounding context.

md
# =============================================================================
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/search.py (reported line 34)May include surrounding context.

python
# =============================================================================
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/search.py (reported line 76)May include surrounding context.

python
# =============================================================================
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/search.py (reported line 37)May include surrounding context.

python
# Auto-load .env from skill directory
# =============================================================================
def _load_env_file():
    env_path = Path(__file__).parent.parent / ".env"
    if env_path.exists():
        with open(env_path) as f:
            for line in f:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README clearly states that the skill sends search queries to the Serper API and then fetches third-party result pages for full-text extraction, but it does not prominently warn users about the privacy and data-handling implications of doing so. In a web-search skill, this is contextually expected behavior, but the lack of explicit disclosure can still mislead users into exposing sensitive queries or causing unanticipated outbound requests to external sites.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill prominently advertises that it returns 'clean readable text' and 'full article text,' but it does not clearly warn users that invoking it causes retrieval and extraction of full third-party page content rather than limited search snippets. That can create a meaningful transparency and privacy issue because user queries may trigger broad content acquisition from external sites, with larger-than-expected data transfer and potential ingestion of copyrighted, sensitive, or policy-restricted material.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill hard-codes an English/global default and mandates locale selection rules based on detected language or region without presenting this as a user choice. While not an exploit primitive by itself, this can cause unintended routing of queries to country/language-specific search contexts, affecting privacy expectations, fairness, and result integrity for multilingual or location-sensitive users.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The spec explicitly recommends that skill descriptions include 'specific keywords that help agents identify relevant tasks' but does not require boundaries, exclusions, or narrow activation criteria. In a skill ecosystem, this can incentivize overly broad descriptions that cause inappropriate skill activation, which may expose agents to unnecessary networked or high-privilege behavior outside the user's true intent.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest describes a web search and page-content extraction tool, but this code also implements local filesystem access to discover and load configuration secrets from a .env file. While using an API key is understandable for Serper access, automatic local secret loading is an extra capability beyond the stated user-facing purpose.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.