Back to skill

Security audit

Baidu Wenku AI picture book of video

Security checks for vulnerabilities and agentic risk

Overview

The skill’s core picture-book function is legitimate, but it contains an undocumented proxy mode that can redirect submitted story content and task IDs to an arbitrary environment-provided URL.

Review before installing. Use only in a trusted environment, avoid submitting confidential or regulated story content, prefer BAIDU_API_KEY from the environment instead of command-line arguments, and do not run it with DUMATE_SCHEDULER_URL or DUMATE_SESSION_ID set unless you trust that exact scheduler service.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ai_picture_book_task_create.py:11
Finding

Unvalidated Scheduler URL Redirects User Story Content and Session Identifier

Content
View full analysis

Vulnerability Details

File Location: scripts/ai_picture_book_task_create.py, lines 11-48
Vulnerability Type: Environment-controlled request destination / sensitive-data exposure
Risk Level: Medium

Vulnerable Code

python
def resolve_sandbox_url(api_key: str, original_url: str) -> Tuple[str, Dict[str, str]]:
    """若当前在沙盒环境中,将目标 URL 替换为代理 URL,并返回需要附加的 headers。"""
    session_id = os.environ.get("DUMATE_SESSION_ID")
    scheduler_url = os.environ.get("DUMATE_SCHEDULER_URL")

    headers = {
        "Content-Type": "application/json",
    }
    if not session_id or not scheduler_url:
        if not api_key:
            raise ValueError("未设置 API Key,请通过环境变量 BAIDU_API_KEY 设置或使用")
        headers.update({
            "Authorization": f"Bearer {api_key}",
            "X-Appbuilder-From": "openclaw",
        })
        return original_url, headers

    parsed = urlparse(original_url)
    proxy_url = f"{scheduler_url}/api/qianfanproxy{parsed.path}"
    if parsed.query:
        proxy_url += f"?{parsed.query}"

    headers.update({
        "Host": parsed.netloc,
        "X-Dumate-Session-Id": session_id,
        "X-Appbuilder-From": "desktop",
    })
    return proxy_url, headers

def ai_picture_book_task_create(api_key: str, method: int, content):
    url = f"{BASE_URL}/tools/ai_picture_book/task_create"
    url, headers = resolve_sandbox_url(api_key, url)
    params = {
        "method": method,
        "input_type": "1",
        "input_content": content,
    }
    response = requests.post(url, headers=headers, json=params)

Technical Analysis

The normal request destination is the fixed HTTPS origin https://qianfan.baidubce.com. However, when both DUMATE_SESSION_ID and DUMATE_SCHEDULER_URL are present, resolve_sandbox_url() silently replaces that destination with a URL constructed from DUMATE_SCHEDULER_URL.

The scheduler URL is not validated ...[truncated 2181 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the environment-controlled proxy path unless it is required for supported deployments.
  2. If proxying is required, parse DUMATE_SCHEDULER_URL and require:
    • An https scheme.
    • An exact hostname from a hardcoded trusted allowlist.
    • An approved port.
    • No embedded username or password.
    • No query string or fragment.
  3. Build the proxy URL using safe URL-joining logic rather than direct string concatenation.
  4. Document proxy mode, its trusted destination, and every transmitted data field in SKILL.md.
  5. Replace reusable session identifiers with narrowly scoped, short-lived proxy tokens where supported.
  6. Add an explicit timeout to requests.post() to prevent indefinite blocking.
  7. Fail closed if proxy configuration is incomplete or does not exactly match the trusted deployment configuration.
  8. Add tests confirming that HTTP, unknown hosts, unexpected ports, and credential-bearing URLs are rejected.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ai_picture_book_task_query.py:11
Finding

Unvalidated Scheduler URL Redirects Task Query Identifiers and Session Identifier

Content
View full analysis

Vulnerability Details

File Location: scripts/ai_picture_book_task_query.py, lines 11-47
Vulnerability Type: Environment-controlled request destination / sensitive-data exposure
Risk Level: Medium

Vulnerable Code

python
def resolve_sandbox_url(api_key: str, original_url: str) -> Tuple[str, Dict[str, str]]:
    """若当前在沙盒环境中,将目标 URL 替换为代理 URL,并返回需要附加的 headers。"""
    session_id = os.environ.get("DUMATE_SESSION_ID")
    scheduler_url = os.environ.get("DUMATE_SCHEDULER_URL")

    headers = {
        "Content-Type": "application/json",
    }
    if not session_id or not scheduler_url:
        if not api_key:
            raise ValueError("未设置 API Key,请通过环境变量 BAIDU_API_KEY 设置或使用")
        headers.update({
            "Authorization": f"Bearer {api_key}",
            "X-Appbuilder-From": "openclaw",
        })
        return original_url, headers

    parsed = urlparse(original_url)
    proxy_url = f"{scheduler_url}/api/qianfanproxy{parsed.path}"
    if parsed.query:
        proxy_url += f"?{parsed.query}"

    headers.update({
        "Host": parsed.netloc,
        "X-Dumate-Session-Id": session_id,
        "X-Appbuilder-From": "desktop",
    })
    return proxy_url, headers

def ai_picture_book_task_query(api_key: str, task_id: str):
    url = f"{BASE_URL}/tools/ai_picture_book/query"
    url, headers = resolve_sandbox_url(api_key, url)
    task_ids = task_id.split(",")
    params = {
        "task_ids": task_ids,
    }
    response = requests.post(url, headers=headers, json=params)

Technical Analysis

The task-query script accepts the request destination from DUMATE_SCHEDULER_URL whenever a DUMATE_SESSION_ID is also present. It performs no scheme, hostname, port, or origin validation.

An arbitrary server can therefore receive all supplied task IDs and the Dumate session identifier. An HTTP destination is accepted, so these values can also be expo ...[truncated 1538 chars]

Remediation
View remediation

Remediation Suggestions

  1. Restrict proxy destinations to exact, explicitly trusted HTTPS origins.
  2. Reject non-HTTPS schemes, embedded credentials, fragments, unknown ports, and non-allowlisted hosts.
  3. Do not trust DUMATE_SCHEDULER_URL merely because it exists in the environment.
  4. Use a short-lived, narrowly scoped proxy credential instead of a broadly reusable session identifier.
  5. Validate the structure and expected fields of proxy responses before presenting them as trusted task results.
  6. Add an explicit request timeout.
  7. Document the proxy route and data disclosure in the Skill documentation.
  8. Prefer environment-only API-key handling; remove or discourage --api_key because command-line secrets may be visible in shell history or process listings.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ai_picture_book_poll.py:25
Finding

Polling Requests Can Be Redirected to an Arbitrary Scheduler Endpoint

Content
View full analysis

Vulnerability Details

File Location: scripts/ai_picture_book_poll.py, lines 25-61
Vulnerability Type: Environment-controlled request destination / repeated sensitive-data exposure
Risk Level: Medium

Vulnerable Code

python
def resolve_sandbox_url(api_key: str, original_url: str) -> Tuple[str, Dict[str, str]]:
    """若当前在沙盒环境中,将目标 URL 替换为代理 URL,并返回需要附加的 headers。"""
    session_id = os.environ.get("DUMATE_SESSION_ID")
    scheduler_url = os.environ.get("DUMATE_SCHEDULER_URL")

    headers = {
        "Content-Type": "application/json",
    }
    if not session_id or not scheduler_url:
        if not api_key:
            raise ValueError("未设置 API Key,请通过环境变量 BAIDU_API_KEY 设置或使用")
        headers.update({
            "Authorization": f"Bearer {api_key}",
            "X-Appbuilder-From": "openclaw",
        })
        return original_url, headers

    parsed = urlparse(original_url)
    proxy_url = f"{scheduler_url}/api/qianfanproxy{parsed.path}"
    if parsed.query:
        proxy_url += f"?{parsed.query}"

    headers.update({
        "Host": parsed.netloc,
        "X-Dumate-Session-Id": session_id,
        "X-Appbuilder-From": "desktop",
    })
    return proxy_url, headers

def query_task(api_key: str, task_id: str) -> Dict[str, Any]:
    """Query task status."""
    url = f"{BASE_URL}/tools/ai_picture_book/query"
    url, headers = resolve_sandbox_url(api_key, url)
    params = {"task_ids": [task_id]}

    response = requests.post(url, headers=headers, json=params, timeout=5)

Technical Analysis

The polling script uses the same unvalidated proxy mechanism as the creation and manual-query scripts. Because polling repeats the query up to the configured number of attempts, the task ID and DUMATE_SESSION_ID may be transmitted repeatedly to an arbitrary destination.

The destination may use HTTP, and no trusted-host allowlist is enforced. The attacker-co ...[truncated 1561 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove proxy support if it is not essential to the declared Skill functionality.
  2. Otherwise, permit only exact trusted HTTPS scheduler origins from a hardcoded allowlist.
  3. Reject scheduler URLs with HTTP, unexpected ports, credentials, query strings, fragments, or non-allowlisted hostnames.
  4. Validate response schemas and ensure returned video URLs use approved HTTPS domains before presenting them as successful results.
  5. Use short-lived, least-privilege proxy credentials rather than a general session identifier.
  6. Avoid resending identifiers after malformed or unauthenticated responses.
  7. Document proxy mode and its data flows in SKILL.md.
  8. Add security tests for destination validation, redirect behavior, forged responses, and disallowed video URL origins.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (6)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares runtime capabilities requiring environment access and network/API communication, but it does not explicitly constrain tool scope with permissions or allowed-tools. This weakens least-privilege controls and can allow an agent runtime to invoke broader capabilities than users or reviewers would expect, increasing the chance of unintended data access or outbound requests.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs users to submit story content to create picture-book tasks, but it does not clearly warn that the content is transmitted to Baidu's external service for processing. Users may unknowingly send sensitive or proprietary text to a third party, creating privacy, confidentiality, and compliance risks; in this content-creation context, that omission is especially relevant because stories may include unpublished, personal, or client material.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/ai_picture_book_poll.py (reported line 61)May include surrounding context.

python
params = {
        "task_ids": task_ids,
    }
    response = requests.post(url, headers=headers, json=params)
    response.raise_for_status()
    result = response.json()
    datas = []

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/ai_picture_book_task_create.py (reported line 48)May include surrounding context.

python
params = {
        "task_ids": task_ids,
    }
    response = requests.post(url, headers=headers, json=params)
    response.raise_for_status()
    result = response.json()
    datas = []

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/ai_picture_book_task_query.py (reported line 47)May include surrounding context.

python
params = {
        "task_ids": task_ids,
    }
    response = requests.post(url, headers=headers, json=params)
    response.raise_for_status()
    result = response.json()
    datas = []

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Natural-language strings in the docstring, CLI description, help text, and error message are presented only in Chinese. This can violate language/locale policy when a skill forces a specific language without user opt-in and no region-specific justification is provided in the file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.