Back to skill

Security audit

Gateway Validator

Security checks for vulnerabilities and agentic risk

Overview

This skill is a real OpenClaw gateway validator, but it can transmit provider API keys to configurable endpoints and can replace and restart production gateway configuration with weak controls.

Review carefully before installing. Use only with trusted gateway configs and trusted provider endpoints, avoid validating untrusted baseUrl changes with real API keys, and do not run the apply script unless you have reviewed the exact config file and are prepared for a production gateway restart.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/validate.py:77
Finding

Provider API Credentials Can Be Transmitted to Arbitrary or Plaintext Endpoints

Content
View full analysis

Vulnerability Details

File Location: scripts/validate.py:77-78, scripts/validate.py:111-121, scripts/validate.py:145-146, scripts/validate.py:218-220; scripts/test_provider.py:64-68, scripts/test_provider.py:95-108, scripts/test_provider.py:134-139, scripts/test_provider.py:218-220
Vulnerability Type: Unrestricted credential destination and server-side request forgery
Risk Level: High

Vulnerable Code

From scripts/validate.py:

python
def test_openai(api_key: str, base_url: Optional[str], model: Optional[str]) -> Tuple[bool, str]:
    """Test OpenAI API."""
    url = (base_url or 'https://api.openai.com/v1') + '/models'
    req = urllib.request.Request(url, headers={'Authorization': f'Bearer {api_key}'})
python
def test_anthropic(api_key: str, base_url: Optional[str], model: Optional[str]) -> Tuple[bool, str]:
    """Test Anthropic API."""
    url = (base_url or 'https://api.anthropic.com') + '/v1/messages'
    test_model = model or 'claude-3-5-sonnet-20241022'

    data = json.dumps({
        'model': test_model,
        'max_tokens': 1,
        'messages': [{'role': 'user', 'content': 'Hi'}]
    }).encode()

    req = urllib.request.Request(
        url, data=data,
        headers={
            'Content-Type': 'application/json',
            'x-api-key': api_key,
            'anthropic-version': '2023-06-01'
        }
    )
python
def test_provider(name: str, settings: Dict) -> Tuple[bool, str]:
    """Test a provider."""
    api_key = settings.get('apiKey', '')
    base_url = settings.get('baseUrl')
    model = settings.get('model')

From scripts/test_provider.py:

python
def test_openai_api(api_key: str, base_url: str = None, model: str = None) -> Tuple[bool, str]:
    """Test OpenAI API connectivity and model."""
    url = (base_url or 'https://api.openai.com/v1') + '/models'

    req = urllib.request.Reques
...[truncated 3129 chars]
Remediation
View remediation

Remediation Suggestions

  1. Normalize URLs before use and permit only https for credential-bearing provider requests.
  2. Maintain provider-specific hostname allowlists, such as the official OpenAI and Anthropic API hosts.
  3. Treat custom endpoints as a separate, explicit feature requiring informed user confirmation.
  4. Reject URLs containing user information, fragments, unsupported ports, loopback addresses, link-local addresses, and private network addresses unless local access is explicitly required.
  5. Resolve destination hostnames and validate all resulting IP addresses to reduce DNS rebinding risk.
  6. Disable automatic redirects or validate every redirect target. Never forward an API key when the scheme, hostname, or port changes.
  7. Build authentication headers only after the final destination has passed validation.
  8. Add tests covering plaintext URLs, attacker-controlled domains, cross-origin redirects, loopback targets, private addresses, and link-local metadata endpoints.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/validate.py:274
Finding

Validation Unnecessarily Transmits Credentials for Every Configured Provider

Content
View full analysis

Vulnerability Details

File Location: scripts/validate.py:274-295; scripts/test_provider.py:246-255
Vulnerability Type: Excessive secret access and outbound network privilege
Risk Level: Medium

Vulnerable Code

From scripts/validate.py:

python
# Test default first
if default and default in providers:
    print(f"      {default}...", end=" ")
    success, msg = test_provider(default, providers[default])
    if success:
        print(f"✅ {msg}")
    else:
        print(f"❌ {msg}")
        return False, f"Default provider '{default}' failed: {msg}"

# Test others
for name, settings in providers.items():
    if name == default:
        continue
    print(f"      {name}...", end=" ")
    success, msg = test_provider(name, settings)
    if success:
        print(f"✅ {msg}")
    else:
        print(f"❌ {msg}")
        return False, f"Provider '{name}' failed: {msg}"

From scripts/test_provider.py:

python
for provider_name, settings in providers.items():
    print(f"\n  Testing {provider_name}...", end=" ")

    success, message = test_provider(provider_name, settings)

    if success:
        print(f"✅ {message}")
    else:
        print(f"❌ {message}")
        all_passed = False

Technical Analysis

validate.py tests the default provider and then every remaining provider regardless of which configuration field changed. The standalone provider tester likewise authenticates to every configured provider without a provider-selection mechanism.

This behavior expands secret access and network activity beyond what is necessary to validate many changes. For example, changing a temperature value for one provider does not require transmitting credentials for all unrelated providers. It also amplifies the arbitrary-destination weakness because every configured credential may be sent when only one provider needs validation.

Attack Path

  1. T ...[truncated 1308 chars]
Remediation
View remediation

Remediation Suggestions

  1. Determine which provider and fields are affected by the requested changes.
  2. Test only the modified provider when credential, model, or endpoint validation is necessary.
  3. Avoid provider connectivity tests for changes that can be validated locally, such as numeric ranges or logging settings.
  4. Add an explicit provider-selection option to test_provider.py.
  5. Make full-provider testing a separate opt-in operation with a clear warning that it will use every configured credential.
  6. Defer loading each API key until its provider has been selected for testing.
  7. Combine selective testing with strict destination validation so an affected credential is sent only to its authorized endpoint.
  8. Add tests proving that unrelated provider functions are not invoked for single-provider or syntax-only changes.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/apply_change.py:66
Finding

Production Configuration Can Be Replaced Without Revalidation

Content
View full analysis

Vulnerability Details

File Location: scripts/apply_change.py:66-87
Vulnerability Type: Unverified production configuration replacement
Risk Level: Medium

Vulnerable Code

python
def main():
    """CLI entry point."""
    if len(sys.argv) < 2:
        print("Usage: apply_change.py <validated_config.yaml>")
        sys.exit(1)

    source_path = Path(sys.argv[1])
    if not source_path.exists():
        print(f"❌ Source config not found: {source_path}")
        sys.exit(1)

    target_path = find_config_file()
    if not target_path:
        print("❌ Could not find production config location")
        sys.exit(1)

    print(f"📝 Applying validated config...")

    # Backup current config
    backup_path = backup_config(target_path)
    print(f"   💾 Backup created: {backup_path.name}")

    # Apply new config
    if apply_config(source_path, target_path):
        print(f"   ✅ Config applied to {target_path}")
    else:
        print(f"   ❌ Failed to apply config")
        sys.exit(1)

The called copy function performs no additional validation:

python
def apply_config(source_path: Path, target_path: Path) -> bool:
    """Copy validated config to production location."""
    try:
        shutil.copy2(source_path, target_path)
        return True
    except Exception as e:
        print(f"❌ Failed to apply config: {e}")
        return False

Technical Analysis

The apply script treats any existing source path as a validated configuration. It does not parse the YAML, invoke schema or semantic validation, verify that validation previously succeeded, or bind the source content to an approved digest.

The source is also not required to be a regular file, and symlink behavior is not explicitly rejected. The destination is overwritten using shutil.copy2 rather than an atomic write-and-replace sequence. The subsequent gateway restart can a ...[truncated 1583 chars]

Remediation
View remediation

Remediation Suggestions

  1. Parse and fully validate the source configuration inside apply_change.py immediately before applying it.
  2. Reuse strict schema, provider, endpoint, and security-policy validation rather than trusting a filename or previous process.
  3. Bind approval to the exact file contents using a digest or a short-lived validation artifact containing the approved hash.
  4. Require the source to be a regular file and reject symlinks, devices, directories, and files with unsafe ownership or permissions.
  5. Open files defensively to reduce time-of-check/time-of-use substitution.
  6. Write the new configuration to a securely created temporary file in the target directory, set restrictive permissions, flush it, and atomically replace the destination.
  7. Preserve the existing owner, group, and restrictive mode only after validating that they are appropriate.
  8. Run a post-write validation before restarting the gateway.
  9. If restart or health verification fails, automatically restore the backup and report the rollback.
  10. Consider limiting application to an approved patch rather than permitting unrestricted replacement of the entire production configuration.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A skill that presents itself as live/test-environment validation but performs only static schema or file checks can falsely certify changes that will fail at runtime. In the context of a production gateway, this undermines the control's core safety purpose and can lead to service disruption, failed provider calls, or misconfigured authentication once changes are applied.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A skill that presents itself as live/test-environment validation but performs only static schema or file checks can falsely certify changes that will fail at runtime. In the context of a production gateway, this undermines the control's core safety purpose and can lead to service disruption, failed provider calls, or misconfigured authentication once changes are applied.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

A skill that presents itself as live/test-environment validation but performs only static schema or file checks can falsely certify changes that will fail at runtime. In the context of a production gateway, this undermines the control's core safety purpose and can lead to service disruption, failed provider calls, or misconfigured authentication once changes are applied.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A skill that presents itself as live/test-environment validation but performs only static schema or file checks can falsely certify changes that will fail at runtime. In the context of a production gateway, this undermines the control's core safety purpose and can lead to service disruption, failed provider calls, or misconfigured authentication once changes are applied.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This script does not merely validate configuration changes; it copies a supplied config into the user's production OpenClaw configuration location. In the context of a skill described as a validator, that mismatch is dangerous because using the skill can directly alter production settings, including providers, models, and API keys, without a clearly separated deployment step.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code restarts the production gateway after applying the config, turning a supposed validator into an active deployment mechanism. In this skill context, that increases risk because a user expecting safe preflight checks could instead trigger immediate service-impacting changes, downtime, or activation of malicious or broken configuration.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
95% confidence
Finding

Copying the full parent environment into a spawned process can expose sensitive credentials such as provider API keys to a child binary that is discovered from PATH and not strongly trusted. In this skill's context, gateway validation likely runs with highly sensitive secrets loaded, which makes indiscriminate environment inheritance especially dangerous.

Content

Scanner excerpt · scripts/test_gateway.py (reported line 78)May include surrounding context.

python
if not openclaw:
        return None, "openclaw binary not found in PATH", None
    
    env = os.environ.copy()
    env['OPENCLAW_CONFIG'] = str(config_path)
    env['OPENCLAW_LOG_LEVEL'] = 'error'

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises operational behavior that implies filesystem, network, process, and possibly environment access, yet it declares no explicit tool scope or permission boundaries. This creates an authorization ambiguity where a reviewer or execution framework cannot reliably constrain what the skill is allowed to do, increasing the risk of over-privileged execution against production configuration and secrets.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger scope activates on broad phrases like any config modification, which can cause the skill to run in situations beyond tightly scoped gateway validation. In a skill that may access config files, secrets, network endpoints, or restart services, broad auto-activation increases the chance of unintended privileged actions or misuse from ambiguous user requests.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module docstring explicitly states that it applies validated configuration to production, which contradicts the manifest's stated validation-focused purpose. This kind of intent mismatch is security-relevant because it can mislead operators and downstream automation into granting broader trust or execution than the skill's true behavior deserves.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/apply_change.py (reported line 51)May include surrounding context.

python
try:
        # Try to restart via openclaw command
        result = subprocess.run(
            ['openclaw', 'gateway', 'restart'],
            capture_output=True,
            text=True,

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module documentation claims it verifies completion processing, but the main flow never performs the advertised end-to-end request. In a validation skill for production gateway changes, this can create a false sense of safety and allow broken or misconfigured deployments to be approved.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/test_gateway.py (reported line 22)May include surrounding context.

python
def find_openclaw_binary() -> Optional[str]:
    """Find openclaw binary in PATH."""
    result = subprocess.run(['which', 'openclaw'], capture_output=True, text=True)
    if result.returncode == 0:
        return result.stdout.strip()
    return None

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
79% confidence
Finding

The script executes whatever openclaw binary is found via PATH and passes it a full copy of the current environment, which may include API keys and other secrets. In a compromised or untrusted environment, a malicious openclaw earlier in PATH could be executed and exfiltrate credentials or run arbitrary code under the user's privileges.

Content

Scanner excerpt · scripts/test_gateway.py (reported line 84)May include surrounding context.

python
# Start gateway
    try:
        proc = subprocess.Popen(
            [openclaw, 'gateway', 'start', '--port', '0'],
            stdout=subprocess.PIPE,
            stderr=subprocess.PIPE,

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The success message says the gateway 'is responding' even though the code only verifies the process remains alive for a short time. This is misleading in a production-change validator and may cause operators or automation to treat nonfunctional configurations as safe.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/test_provider.py (reported line 20)May include surrounding context.

python
# Provider API endpoints for validation
PROVIDER_ENDPOINTS = {
    'openai': {
        'models_url': 'https://api.openai.com/v1/models',
        'test_model': 'gpt-4o-mini',
    },
    'anthropic': {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/validate.py (reported line 85)May include surrounding context.

python
# Provider API endpoints for validation
PROVIDER_ENDPOINTS = {
    'openai': {
        'models_url': 'https://api.openai.com/v1/models',
        'test_model': 'gpt-4o-mini',
    },
    'anthropic': {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/test_provider.py (reported line 29)May include surrounding context.

python
'api_version': '2023-06-01',
    },
    'moonshot': {
        'models_url': 'https://api.moonshot.cn/v1/models',
        'test_model': 'kimi-k2.5',
    },
    'google': {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/validate.py (reported line 146)May include surrounding context.

python
'api_version': '2023-06-01',
    },
    'moonshot': {
        'models_url': 'https://api.moonshot.cn/v1/models',
        'test_model': 'kimi-k2.5',
    },
    'google': {

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This code sends the configured API key in an Authorization header to a URL that can be overridden via settings.baseUrl. If an attacker can influence the gateway config, they can redirect validation traffic to an arbitrary host and capture provider credentials, turning a legitimate connectivity check into SSRF-plus-secret-exfiltration.

Content

Scanner excerpt · scripts/test_provider.py (reported line 64)May include surrounding context.

python
def test_openai_api(api_key: str, base_url: str = None, model: str = None) -> Tuple[bool, str]:
    """Test OpenAI API connectivity and model."""
    url = (base_url or 'https://api.openai.com/v1') + '/models'
    
    req = urllib.request.Request(
        url,

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This function likewise sends the bearer token to a configurable base_url with no validation. In the context of a gateway configuration validator, this is more dangerous because the tool is expected to process untrusted configuration changes, so a malicious config can cause credential exfiltration to attacker-controlled infrastructure.

Content

Scanner excerpt · scripts/test_provider.py (reported line 137)May include surrounding context.

python
def test_moonshot_api(api_key: str, base_url: str = None, model: str = None) -> Tuple[bool, str]:
    """Test Moonshot API."""
    url = (base_url or 'https://api.moonshot.cn/v1') + '/models'
    
    req = urllib.request.Request(
        url,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The Google validation path places the API key directly in the request URL as a query parameter. Query-string secrets are more likely to be exposed via logs, proxies, browser/history equivalents, monitoring systems, exception traces, or intermediary infrastructure, making credential leakage more likely than header-based authentication.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill metadata promises validation by testing changes on an isolated gateway or by checking provider connectivity, but this script only performs static YAML/schema-style checks. In a production-change workflow, that mismatch can create a false sense of safety: invalid credentials, unreachable endpoints, or incompatible provider settings may be approved and later break production.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring explicitly says validation is done 'without running a gateway,' which conflicts with the advertised behavior of testing changes on an isolated gateway when possible. This inconsistency is dangerous in an operational security context because users may rely on the skill to catch connectivity or integration failures that it never attempts to detect.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.