Back to skill

Security audit

Adversarial Engine

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real multi-model debate engine, but it exposes unauthenticated network interfaces and can automatically run model-generated Python on the host under a misleading sandbox label.

Do not install this as-is on a shared, internet-reachable, or sensitive machine. Treat it as Review-needed: disable generated-code execution, remove and rotate the embedded API key, bind services to localhost with authentication, add job limits, isolate WebSocket sessions, and avoid enabling knowledge-base retrieval unless the files are approved for external model processing.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (6)

T09 · Insecure Skill Coding Practices

Error
Location
engine.py:148
Finding

Model-Generated Python Executes Without Effective Isolation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
engine.py:39
Finding

Hard-Coded External API Credential in Source Code

Content
View full analysis
str: """获取API Key(优先使用路由池)""" if HAS_KEY_ROUTER: router = get_key_router() return router.request_key(model) return DEFAULT_API_KEY ``` ### Technical Analysis A credential-shaped API key is stored directly in two source files. Anyone who obtains the project archive, repository contents, deployment image, backup, or generated logs containing the source can recover the credential. The fallback behavior makes the credential operational rather than an unused example. It is selected automatically whenever the external key-router module cannot be imported. ### Attack Path 1. An attacker obtains read access to the Skill package or repository. 2. The attacker extracts `DEFAULT_API_KEY` from `engine.py` or `async_engine.py`. 3. The attacker sends independent requests to the configured DashScope endpoint using the key. 4. Requests consume the key owner's quota and may be attributed to the legitimate owner. ### Impact Assessment Potential impact includes: - Unauthorized use of paid model services. - Financial loss and quota exhaustion. - Service disruption for legitimate users. - Abuse attributed to the credential owner. - Continued compromise across every installation containing the same key. The exact backend permissions of the key cannot be established from the source, but its use in the Authorization header confirms that it is treated as a secret. ]]>
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
engine.py:109
Finding

Local Knowledge-Base Content Is Sent to an External Model Provider Without Filtering

Content
View full analysis
= top_k: return results ``` Retrieved content is inserted into the prompt: ```python knowledge_text = "" if knowledge: knowledge_text = "\n\n=== 知识库检索结果 ===\n" + "\n".join(knowledge[:2]) ``` The prompt is then sent to an external endpoint: ```python payload = { "model": model, "messages": [ {"role": "system", "content": system_prompt}, {"role": "user", "content": user_prompt} ], "max_tokens": 2000, "temperature": 0.7 } resp = requests.post(BASE_URL, headers=headers, json=payload, timeout=90) ``` ### Technical Analysis When vector search is enabled, the engine recursively scans `/home/admin/.openclaw/workspace/kb`, selects Markdown files by a simple keyword match, and inserts excerpts into external LLM requests. There is no file-level sensitivity policy, directory allowlist, secret detection, redaction, data-loss-prevention check, or per-request confirmation. The search operates over every matching Markdown file below the configured directory. External model calls are necessary for the declared multi-model debate feature, but sending arbitrary local knowledge-base excerpts exceeds minimum privilege when the content has not been explicitly ...[truncated 911 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
server.py:101
Finding

Unauthenticated Public Interfaces Can Start Unbounded Debate and Code-Execution Jobs

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
server.py:43
Finding

Global WebSocket Broadcast Discloses Debate Data Across Users

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
server.py:75
Finding

Permissive Cross-Origin Configuration and Missing WebSocket Origin Validation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (26)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the actual exposed behavior is operating HTTP/WebSocket listeners with connection management and message broadcast, but the skill is described mainly as a debate engine, users and deployers may not realize they are launching a network service. Hidden service exposure expands the attack surface, can leak session content to connected clients, and may create unauthorized access paths on the host environment.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

If the actual exposed behavior is operating HTTP/WebSocket listeners with connection management and message broadcast, but the skill is described mainly as a debate engine, users and deployers may not realize they are launching a network service. Hidden service exposure expands the attack surface, can leak session content to connected clients, and may create unauthorized access paths on the host environment.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

If the actual exposed behavior is operating HTTP/WebSocket listeners with connection management and message broadcast, but the skill is described mainly as a debate engine, users and deployers may not realize they are launching a network service. Hidden service exposure expands the attack surface, can leak session content to connected clients, and may create unauthorized access paths on the host environment.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

If the actual exposed behavior is operating HTTP/WebSocket listeners with connection management and message broadcast, but the skill is described mainly as a debate engine, users and deployers may not realize they are launching a network service. Hidden service exposure expands the attack surface, can leak session content to connected clients, and may create unauthorized access paths on the host environment.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The component is labeled a 'sandbox', but it only writes code to a temp file and runs it with python3 under /tmp, which is not a real sandbox. This misleading abstraction can cause operators or downstream code to trust dangerous behavior, increasing the chance that arbitrary code runs under false assumptions of safety.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This skill is for debate/review, yet it executes LLM-produced Python locally as part of normal flow. That creates an unnecessary remote-to-local execution path: prompt content influences model output, model output becomes code, and that code is run on the machine, enabling arbitrary command execution and data theft.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill automatically executes code extracted from the engineer model response when enable_code_sandbox is on, with no confirmation dialog, policy review, or safety interlock. In this context, normal user input can indirectly steer generated code, so the absence of consent and gating makes arbitrary local execution especially dangerous.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The class is presented as a 'safe' code sandbox, but it actually writes arbitrary code to a temp file and runs it with the system Python interpreter. This mismatch is dangerous because operators may trust the feature as isolated when it is not, increasing the likelihood that hostile or prompt-injected code is executed with local privileges.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill executes LLM-generated Python automatically with no user confirmation, safety interstitial, or trust boundary. In this context, the content is especially risky because the entire system is an adversarial multi-model debate engine, making hostile or manipulative outputs more likely than in a normal assistant flow.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The engine enables code execution as part of a debate/review workflow even though arbitrary local execution is not necessary to fulfill the advertised purpose. This expands the attack surface substantially: any prompt injection, adversarial topic, or compromised upstream model response can lead directly to code execution on the machine.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises capabilities that imply file access, network communication, and code execution, but it declares no explicit tool scope or permission boundary. In a skill that includes Python sandboxing, WebSocket services, and retrieval/persistence, missing permissions disclosure can cause the host or user to unknowingly grant broad capabilities, increasing the risk of code execution, data exfiltration, and unintended system access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill openly references code execution, network pushing, retrieval, and persistence but provides no user-facing warning about the operational and privacy implications. In this context, users may unknowingly trigger generated-code execution or share sensitive content that is transmitted or retained, which is a meaningful safety and consent failure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The workflow explicitly says conclusions are saved and solidified into a knowledge base, but it does not warn that user content may be retained. This is dangerous because debate sessions may include proprietary designs, prompts, code, or internal reasoning that users do not expect to persist beyond the session.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

A default external API key is hardcoded in source, exposing a reusable secret to anyone who can read the file or logs derived from it. Embedded credentials are easily leaked through source control, backups, artifacts, or debugging output, and they enable unauthorized use of the external service and possible billing or data exposure.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The code invokes python3 on a temporary file whose contents come from model output, which creates direct local code-execution from untrusted generated content. Using subprocess without isolation, privilege dropping, or syscall/resource restrictions means the called code can read files, make network calls, spawn processes, or tamper with the host environment.

Content

Scanner excerpt · async_engine.py (reported line 112)May include surrounding context.

python
temp_path = f.name
        
        try:
            result = subprocess.run(
                ['python3', temp_path],
                capture_output=True,
                text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The function sends user topic and prompt content to an external LLM provider without any visible user-facing notice, consent flow, or data-classification guard. While outbound API use is expected for an LLM engine, topics, debate history, and potentially sensitive content may be transmitted off-box without transparency or redaction.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

manifest 描述聚焦于多模型对抗、代码沙箱验证、向量检索增强与自动收敛;其中“向量检索”在实现中并非真正向量检索,而是对本地知识库目录进行文件系统遍历和关键字匹配读取。再加上 L334-L374、L567-L610 将完整辩论内容持久化到 SQLite,这些本地数据访问/存储行为超出了描述中对“对抗辩论系统”的核心语义,且未在描述中体现数据持久化特性。

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The code invokes subprocess.run on Python source extracted from LLM output, which means untrusted model-generated code is executed on the host. Using a separate process and a timeout does not provide isolation: the code can still read local files, make network calls, consume resources, or abuse inherited permissions.

Content

Scanner excerpt · engine.py (reported line 183)May include surrounding context.

python
temp_path = f.name
        
        try:
            result = subprocess.run(
                ['python3', temp_path],
                capture_output=True,
                text=True,

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

The engine sends prompts and conversation content to an external LLM endpoint, which is expected for cloud inference but still constitutes data exfiltration risk. Because prompts may include prior rounds, knowledge-base excerpts, and possibly code/results, sensitive local information can be transmitted off-host if the topic or retrieved documents contain it.

Content

Scanner excerpt · engine.py (reported line 248)May include surrounding context.

python
}
    
    try:
        resp = requests.post(BASE_URL, headers=headers, json=payload, timeout=90)
        
        if resp.status_code != 200:
            release_api_key(model, rate_limit_hit=(resp.status_code == 429))

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The system prompt instructs the models to always reply in Chinese, imposing a specific language policy regardless of user preference. This is a natural-language policy issue because the file provides no opt-in, fallback, or documented reason for the locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file includes natural-language descriptions such as the module summary and author-facing text only in Chinese. Under the policy, forcing a specific language without user opt-in or justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest describes a multi-model adversarial debate/review engine, but this file implements a network-facing FastAPI server with open WebSocket and HTTP endpoints plus static HTML serving. Providing a public application server/UI is not an obvious requirement of the declared skill purpose itself, which is the debate engine behavior rather than web hosting capability.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The top-level docstring presents the skill entirely in Chinese, including its description and authoring context, with no indication that language choice is optional or region-specific. Under the policy rule, forcing a specific language or locale without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The natural-language interface elements in the manifest—description and triggers—are exclusively Chinese, and the document does not offer an alternative language or note that the skill is intentionally region/language-specific. Per SQP-3, forcing a specific language without user opt-in can be a language policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The system prompt explicitly instructs the model to reply in Chinese. This is a natural-language locale policy issue because the file imposes a language choice unconditionally and does not provide a user opt-in or explain a region-specific need.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.