Back to skill

Security audit

OpenClaw Multi-LLM Adapter

Security checks for vulnerabilities and agentic risk

Overview

This skill is a plausible multi-LLM adapter, but it can route prompts and API keys to configurable endpoints while its documentation overstates capabilities and under-discloses data exposure.

Review before installing. Use only trusted configuration files, avoid custom base URLs unless you control the endpoint, do not send secrets or private data in prompts, and prefer pinned dependencies or a lock file. Treat the advertised Gemini/LiteLLM/load-balancing/cost features as unsupported unless the publisher adds matching code and documentation.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
lib/client.py:85
Finding

Unrestricted OpenAI Base URL Can Expose API Credentials and Conversation Data

Content
View full analysis

Vulnerability Details

File Location: lib/client.py, lines 85-89 and 253-271
Vulnerability Type: Unvalidated custom API endpoint and sensitive-data transmission
Risk Level: Medium

Vulnerable Code

python
self.client = OpenAI(
    api_key=config.api_key,
    base_url=config.base_url
)
python
def load_config(self, path: str) -> None:
    """Load configuration from file"""
    config = json.loads(Path(path).read_text())
    
    for name, provider_config in config.get("providers", {}).items():
        # Expand environment variables
        api_key = provider_config.get("api_key", "")
        if api_key.startswith("${") and api_key.endswith("}"):
            env_var = api_key[2:-1]
            api_key = os.environ.get(env_var)
        
        self.add_provider(ProviderConfig(
            name=name,
            api_key=api_key,
            model=provider_config.get("model", "gpt-4"),
            base_url=provider_config.get("base_url"),
            priority=provider_config.get("priority", 1),
            timeout=provider_config.get("timeout", 30)
        ))

Technical Analysis

The configuration loader accepts an arbitrary base_url and passes it to the OpenAI SDK together with the configured API key. The code does not enforce HTTPS, restrict destination hostnames, reject URLs containing embedded credentials, or require explicit approval for nonstandard endpoints.

Sending prompts and API credentials over the network is necessary for the declared LLM-adapter functionality. However, permitting an unrestricted endpoint exceeds the minimum privilege required for normal OpenAI access. When a custom endpoint is selected, the SDK can send its authentication header, system prompts, user messages, tool schemas, and other request metadata to that destination.

Exploitation requires the attacker to influence the configuration file or convince the user to load an untr ...[truncated 1329 chars]

Remediation
View remediation

Remediation Suggestions

  • Use the official OpenAI endpoint by default and allow custom endpoints only through an explicit, documented opt-in.
  • Require HTTPS for all non-loopback destinations.
  • Validate URLs using a structured URL parser and permit only expected schemes.
  • Reject URLs containing embedded usernames, passwords, fragments, or malformed host components.
  • Maintain an allowlist of trusted provider hostnames where operationally possible.
  • Do not automatically reuse an official provider credential with a custom endpoint. Require a separate credential explicitly designated for that endpoint.
  • Warn users that configuration files are security-sensitive and must not be loaded from untrusted sources.
  • Consider resolving and validating destinations to prevent unintended access to loopback, link-local, metadata-service, and private-network addresses where custom remote endpoints are not required.
  • Add automated tests covering HTTP URLs, untrusted hosts, embedded credentials, and environment-variable credential expansion.

T09 · Insecure Skill Coding Practices

Warning
Location
lib/client.py:174
Finding

Unrestricted Ollama Endpoint Allows Prompt Disclosure and Unbounded Network Requests

Content
View full analysis

Vulnerability Details

File Location: lib/client.py, lines 174-188 and 201-211
Vulnerability Type: Unvalidated outbound request destination, cleartext sensitive-data transmission, and missing timeout
Risk Level: Medium

Vulnerable Code

python
def __init__(self, config: ProviderConfig):
    super().__init__(config)
    self.base_url = config.base_url or "http://localhost:11434"

def chat(self, messages: List[Message], tools: Optional[List[Dict]] = None,
         **kwargs) -> LLMResponse:
    import requests
    
    response = requests.post(
        f"{self.base_url}/api/chat",
        json={
            "model": self.config.model,
            "messages": [m.to_dict() for m in messages],
            "stream": False
        }
    )
python
def chat_stream(self, messages: List[Message], tools: Optional[List[Dict]] = None,
                **kwargs) -> Iterator[str]:
    import requests
    
    response = requests.post(
        f"{self.base_url}/api/chat",
        json={
            "model": self.config.model,
            "messages": [m.to_dict() for m in messages],
            "stream": True
        },
        stream=True
    )

Technical Analysis

OllamaProvider accepts any configured base URL and transmits the complete message collection to its /api/chat route. There is no URL scheme validation, destination restriction, TLS requirement for remote hosts, or explicit user confirmation before sending data outside the local machine.

Although the default endpoint is the expected loopback Ollama service, a supplied configuration can change it to a remote attacker-controlled service or another network-accessible target. If plain HTTP is used remotely, prompts can also be intercepted or modified in transit.

The ProviderConfig.timeout field is not passed to either requests.post() call. Consequently, a remote service can leave the connection or s ...[truncated 1355 chars]

Remediation
View remediation

Remediation Suggestions

  • Restrict Ollama to loopback destinations by default, including validated IPv4 and IPv6 loopback addresses.
  • Require explicit configuration and user approval before connecting to a remote Ollama server.
  • Require HTTPS for remote endpoints and validate certificates using the default trusted certificate store.
  • Parse and validate the configured URL rather than constructing requests through unchecked string concatenation.
  • Reject unsupported schemes, embedded credentials, malformed hosts, and unintended link-local or metadata-service destinations.
  • Pass timeout=self.config.timeout to both requests.post() calls. Prefer separate connect and read timeouts where streaming requirements differ.
  • Call response.raise_for_status() before processing a streaming response.
  • Apply response-size and streaming-duration limits to reduce denial-of-service exposure.
  • Clearly disclose to users when prompts will leave the local machine.

T08 · Insecure Dependencies

Note
Location
requirements.txt:4
Finding

Open-Ended Dependency Constraints Prevent Reproducible and Reviewed Installations

Content
View full analysis

Vulnerability Details

File Location: requirements.txt, lines 4-10
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Low

Vulnerable Code

text
# OpenAI
openai>=1.0.0

# Anthropic
anthropic>=0.18.0

# HTTP requests (for Ollama)
requests>=2.28.0

Technical Analysis

The requirements specify only minimum versions and have no upper bounds, lock file, or package hashes. An installation can therefore resolve to any future release satisfying these constraints. Different installations may execute materially different dependency code without corresponding changes to the audited project.

The listed package names and package source declarations do not show direct evidence of typosquatting, dependency confusion, or a known malicious package. The risk arises from the inability to reproduce the reviewed dependency set and from automatically admitting future, unreviewed versions.

Attack Path

  1. A future dependency version satisfying one of the open-ended constraints is compromised, malicious, or introduces a security regression.
  2. A user installs the project after that version becomes available.
  3. The package resolver selects the affected release because it satisfies the minimum-version constraint.
  4. The project imports and executes the dependency in a process that may have access to LLM API credentials and sensitive prompts.
  5. Malicious or vulnerable dependency behavior can affect the confidentiality and integrity of the host process.

Impact Assessment

The potential scope is that of the Python process installing or running the Skill. Dependencies may access process environment variables, API credentials supplied to SDK constructors, conversation content, files available to the process, and network resources allowed by the host.

This finding does not establish that the currently named packages or any particular released versions are malicious. It identifies a suppl ...[truncated 70 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin dependencies to exact versions that have been reviewed and tested.
  • Generate and commit a lock file appropriate to the deployment workflow.
  • Use package hashes, such as with pip --require-hashes, to verify downloaded artifacts.
  • Obtain packages only from an explicitly configured trusted package index.
  • Scan pinned versions for known vulnerabilities during continuous integration.
  • Review dependency updates through a controlled process rather than accepting all future releases automatically.
  • Regularly refresh pins so security fixes are adopted without sacrificing reproducibility.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The core theme of the description and code align in that this is an LLM adapter with a unified interface and fallback behavior. However, the declared description materially overstates supported providers and features. The code only defines provider classes for OpenAI, Anthropic, and Ollama, and add_provider only instantiates those three. There is no Google Gemini provider, no LiteLLM dependency or abstraction for 100+ providers, and no load-balancing logic such as round-robin, weighted distribution, or health-based traffic sharing. The automatic selection implemented in chat_auto is simple priority-ordered failover, which supports the fallback claim but not load balancing. Because these are central advertised capabilities rather than minor implementation details, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The description presents this skill as a universal multi-provider adapter with broad provider coverage and automatic fallback/load balancing. The supplied code chunk is narrower: it is a CLI wrapper around an underlying LLMClient, with explicit environment setup only for OpenAI and Anthropic, limited provider choices for chat (openai/anthropic/ollama), and provider listing based on environment variables. Google Gemini and '100+ providers via LiteLLM' are not supported or referenced in the visible code. Likewise, fallback/load balancing are not implemented here, though an 'auto' call suggests some delegated selection behavior may exist elsewhere. Additionally, the code exposes comparison and provider-management CLI features that are not reflected in the declared purpose. Because the visible behavior only partially matches and materially overstates supported providers/capabilities, this is a mismatch.

Content

No source excerpt is available for this finding.

Context Leakage

High
Category
Data Exfiltration
Confidence
85% confidence
Finding

Code or instructions that leak agent conversation context to external services, potentially exposing sensitive user interactions.

Content

Scanner excerpt · lib/client.py (reported line 69)May include surrounding context.

python
@abstractmethod
    def chat(self, messages: List[Message], tools: Optional[List[Dict]] = None,
             **kwargs) -> LLMResponse:
        """Send chat request"""
        pass
    
    @abstractmethod

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill declares powerful tools (read, write, exec) and demonstrates use of environment variables and outbound calls to external LLM providers, but it does not define any explicit permission scope such as allowed tools, files, hosts, or network boundaries. In an agent setting, this weakens least-privilege controls and makes accidental overreach or abuse easier, especially when combined with tool execution and provider configuration workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is designed to send prompts, configuration, and possibly tool outputs to third-party LLM providers, yet it gives no clear privacy warning to users. This creates a real data-exposure risk because users may unknowingly submit sensitive prompts, credentials, internal content, or tool-derived data to external services with different retention and logging policies.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documented --execute-tools workflow encourages executing model-requested tools without a strong warning about system-affecting actions. In an agentic context, an LLM can be induced to call tools that read files, modify state, invoke commands, or access network resources, so undocumented execution safety materially increases the chance of prompt-injection-driven or unintended actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code sends the full message list to the OpenAI API, which can include user or system data, but there is no confirmation prompt, user-facing log/print, or nearby comment/docstring warning that data is transmitted off-box. The same pattern appears to be part of a general-purpose client adapter rather than a narrowly scoped 'deploy' or similarly self-evident network action.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The Anthropic provider sends assembled chat messages and optional tools to an external API, which may expose user or system data to a third party. The file does not provide any visible confirmation, user-facing logging, or warning text around this transmission.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
85% confidence
Finding

This performs outbound transmission of full chat messages to a configurable Ollama base URL with no validation that the endpoint is local, trusted, or encrypted. In this skill context, a 'universal adapter' increases exposure because callers may unknowingly route sensitive prompts, system messages, or tool outputs to arbitrary endpoints, including remote hosts over plain HTTP.

Content

Scanner excerpt · lib/client.py (reported line 218)May include surrounding context.

python
**kwargs) -> LLMResponse:
        import requests
        
        response = requests.post(
            f"{self.base_url}/api/chat",
            json={
                "model": self.config.model,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

These requests post the message history to an HTTP endpoint, which is a network transmission of user-provided content, yet the code includes no confirmation prompt, visible logging, or explanatory warning. Even though the default target is localhost, the configurable base URL means data may be sent beyond the local machine.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

The streaming variant also sends conversation data to a configurable endpoint without trust validation or transport-security requirements. Because streaming is often used for live user inputs, this can expose sensitive data continuously to unintended remote services if the base URL is malicious or misconfigured.

Content

Scanner excerpt · lib/client.py (reported line 241)May include surrounding context.

python
**kwargs) -> Iterator[str]:
        import requests
        
        response = requests.post(
            f"{self.base_url}/api/chat",
            json={
                "model": self.config.model,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The chat command sends user-supplied messages, optional system prompts, and tool definitions to external LLM providers without any explicit user-facing privacy warning or confirmation at send time. In a multi-provider adapter, users may not realize their prompts are leaving the local environment, which can lead to accidental disclosure of secrets, proprietary data, or personal information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The compare command duplicates the same user message across multiple providers, increasing the data exposure surface without making that expansion obvious to the user. This is especially risky in a universal adapter because one command can broadcast sensitive content to several third-party services with different retention, logging, and training policies.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The skill documentation is primarily written in Chinese and does not indicate that users may choose another language or that the locale restriction is intentional and justified. Under the stated policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The configuration loader resolves API keys from environment variables, which is access to sensitive credentials, but there is no comment, docstring, or user-facing notice explaining that the skill reads secrets from the environment. Users may not realize the skill depends on and consumes credential material from their environment.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency is specified with a minimum version only, which allows future installs to pull in different releases over time. This weakens build reproducibility and can inadvertently introduce vulnerable or breaking upstream versions into the skill.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
# Multi-LLM Adapter Dependencies

# OpenAI
openai>=1.0.0

# Anthropic
anthropic>=0.18.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The anthropic package is unpinned, so deployments may resolve to different versions depending on install time and repository state. In a security-sensitive adapter that brokers multiple LLM providers, this increases supply-chain uncertainty and makes it harder to verify exposure to known issues.

Content

Scanner excerpt · requirements.txt (reported line 7)May include surrounding context.

text
openai>=1.0.0

# Anthropic
anthropic>=0.18.0

# HTTP requests (for Ollama)
requests>=2.28.0

Unverifiable Dependency: anthropic has 4 known advisory(ies) (CVE-2026-34450 (Claude SDK for Python has Insecure Default File Permissions in Local Filesystem ); CVE-2026-34452 (Claude SDK for Python: Memory Tool Path Validation Race Condition Allows Sandbox); CVE-2026-34450 (The Claude SDK for Python provides access to the Claude API from Python applicat) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
89% confidence
Finding

The manifest does not pin anthropic, and the package has known advisories, so the installed version cannot be verified as safe. While the file alone does not prove an exploitable vulnerable version is in use, the uncertainty itself is a supply-chain risk that is more concerning in an adapter intended to connect to external AI services.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

Using requests>=2.28.0 permits installation of any later version, which reduces reproducibility and may allow an affected release to be installed if the ecosystem changes. Because requests is a core HTTP client, dependency drift can have security consequences across all outbound API calls.

Content

Scanner excerpt · requirements.txt (reported line 10)May include surrounding context.

text
anthropic>=0.18.0

# HTTP requests (for Ollama)
requests>=2.28.0

# Optional: Google Gemini
# google-generativeai>=0.3.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The requests dependency is unpinned despite known advisories affecting some releases, so consumers cannot determine from this manifest whether installation will select a safe version. Since requests handles outbound HTTP interactions, an unsafe resolved version could expose network communications or credentials to known issues.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.