Back to skill

Security audit

Llm Provider Forensics

Security checks for vulnerabilities and agentic risk

Overview

The skill is purpose-aligned for LLM endpoint forensics, but it uses local provider API keys in outbound requests to configurable endpoints without enough destination safeguards.

Install only if you are comfortable with the skill reading selected OpenClaw provider credentials and sending authenticated probes to the configured endpoints. Use disposable or least-privileged API keys, verify each base URL before running it, avoid unknown gateways, and be especially careful with --deep because it sends additional test prompts.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/llm_provider_forensics.py:94
Finding

Unvalidated Provider URLs Enable Credential Disclosure and Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: scripts/llm_provider_forensics.py:94-100, 111-114, 125-131, 146-148, 174-178, 427-444
Vulnerability Type: Unvalidated credential-bearing outbound requests / SSRF
Risk Level: High
Category: T09: Insecure Skill Coding Practices

Vulnerable Code

python
def openai_call(base_url, api_key, model, prompt, endpoint='responses', timeout=20, stream=False):
    headers = {'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json', 'User-Agent': 'Mozilla/5.0'}
    path = '/responses' if endpoint == 'responses' else '/chat/completions'
    body = {'model': model, 'input': prompt, 'max_output_tokens': 256} if endpoint == 'responses' else {'model': model, 'messages': [{'role': 'user', 'content': prompt}], 'max_tokens': 256, 'temperature': 0}
    if stream:
        body['stream'] = True
    ok, http, lat, raw = _request(base_url.rstrip('/') + path, headers=headers, body=body, timeout=timeout)
python
def anthropic_call(base_url, api_key, model, prompt, timeout=20):
    headers = {'x-api-key': api_key, 'anthropic-version': '2023-06-01', 'content-type': 'application/json', 'User-Agent': 'Mozilla/5.0'}
    body = {'model': model, 'max_tokens': 256, 'messages': [{'role': 'user', 'content': prompt}]}
    ok, http, lat, raw = _request(base_url.rstrip('/') + '/v1/messages', headers=headers, body=body, timeout=timeout)
python
def gemini_call(base_url, api_key, model, prompt, timeout=20):
    base = base_url.rstrip('/')
    q = urllib.parse.urlencode({'key': api_key})
    body = {'contents': [{'parts': [{'text': prompt}]}]}
    last = {'ok': False, 'http': None, 'latency_s': None, 'text': None, 'usage': None, 'object': None, 'endpoint': None, 'raw_preview': ''}
    for p in [f'/v1beta/models/{model}:generateContent?{q}', f'/v1/models/{model}:generateContent?{q}']:
        ok, http, lat, raw = _request(base + p, headers={'Content-Type': 'application/json', 'User-Agent': 'Mozilla/5.0'}, body=
...[truncated 4003 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require secure transport

    • Accept only https:// URLs by default.
    • Reject plaintext HTTP unless a narrowly scoped, explicit development override is provided.
    • Display a prominent warning and never attach production credentials when an insecure override is used.
  2. Validate destinations before sending credentials

    • Parse URLs with urllib.parse.urlsplit.
    • Reject embedded user information, malformed hosts, unexpected schemes, and unsupported ports.
    • Resolve all destination addresses and block loopback, private, link-local, multicast, unspecified, reserved, and cloud metadata ranges by default.
    • Revalidate the destination after DNS resolution and immediately before connection to reduce DNS-rebinding risk.
  3. Control redirects

    • Disable automatic redirects for authenticated probes, or validate every redirect target using the same scheme, hostname, port, and IP policy.
    • Strip credentials whenever a redirect changes the origin.
    • Set a strict redirect limit.
  4. Bind credentials to expected hosts

    • Associate each credential with an approved provider hostname or explicit user-maintained allowlist.
    • Require interactive confirmation before sending a key to an unrecognized or changed host.
    • Use disposable, least-privileged audit credentials for unknown gateways.
  5. Reduce Gemini query-string exposure

    • Permit query-parameter authentication only where required by the selected protocol.
    • Avoid printing, persisting, or including complete request URLs in exceptions and logs.
    • Redact key, api_key, authorization headers, and equivalent secret fields from all diagnostics.
  6. Harden configuration handling

    • Validate all selected provider records before initiating any request.
    • Reject missing, ambiguous, or conflicting URL and protocol fields.
    • Document that provider configuration files contain sensitive credentials and should have restrictive filesyste ...[truncated 309 chars]
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 101)May include surrounding context.

md
python3 scripts/llm_provider_forensics.py --config /root/.openclaw/openclaw.json --providers omgteam ypemc vpsai --model gpt-5.4 --deep

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill instructs the agent to probe provider endpoints and load local reference files, which implies network and file-read capabilities, but it does not declare any explicit tool scope or permission boundaries. That omission can let the skill run with whatever ambient agent privileges are available, increasing the chance of unintended outbound requests, broader file access, or misuse in environments with powerful default tools.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code actively sends user-supplied prompts, model IDs, and request metadata to external provider endpoints during probing, including optional deep tests that contain policy-sensitive prompts. In a forensics skill this behavior is expected, but without explicit disclosure, consent gates, or data-minimization controls, operators may unintentionally exfiltrate sensitive prompts or identifiers to third-party services.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated summary schema hard-codes Chinese keys such as '已配置但不可用' and '当前可用' with no option for the user to choose language or locale. This is a natural-language policy violation because it forces a specific language in output without documented opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script loads API keys from configuration and immediately uses them in outbound Authorization headers, but contains no in-code safeguards around secret sourcing, redaction, or user notice. While using API keys is necessary for the tool's purpose, this increases the risk of accidental credential misuse, testing against unintended endpoints, or unsafe handling in surrounding automation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The phrase "Chinese-first deployment context" encodes a locale-oriented preference in natural language. Because the file does not explain that this is merely an observational fingerprinting signal or provide user choice/justification, it can read as a locale-policy assumption.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.