Back to skill

Security audit

Research-engine

Security checks for vulnerabilities and agentic risk

Overview

This research skill is mostly coherent, but it needs review because research topics can cause local files to be written outside the intended folder and browsing history is retained locally.

Review before installing. Use this only where external web research is allowed, avoid entering confidential topics or secrets, set RESEARCH_DIR to a location you control, and fix or avoid topic values containing slashes or path traversal until filename sanitization is added.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
research_engine.py:208
Finding

Arbitrary File Write Outside the Configured Research Directory

Content
View full analysis

Vulnerability Details

File Location: research_engine.py, lines 208-211
Vulnerability Type: Path traversal and unsafe user-controlled file path construction
Risk Level: High

python
filename = f"{topic.replace(' ', '_')}_{datetime.now().strftime('%Y%m%d_%H%M')}.md"
filepath = os.path.join(RESEARCH_DIR, filename)
with open(filepath, 'w', encoding='utf-8') as f:
    f.write(report)

Technical Analysis

The public run_research() function uses the caller-controlled topic value to construct the report filename. When invoked through the command-line interface, the same value originates directly from command-line arguments.

Replacing spaces with underscores does not remove absolute-path prefixes, path separators, or parent-directory components such as ../. Consequently:

  • If topic begins with an absolute path, os.path.join(RESEARCH_DIR, filename) discards RESEARCH_DIR and returns the attacker-controlled absolute path.
  • If topic contains ../, path resolution can traverse outside RESEARCH_DIR.
  • The resulting path is opened in write mode, creating a new file or truncating an existing file with the same generated name.

The timestamp suffix prevents selection of an entirely arbitrary final filename, but it does not enforce the intended directory boundary. An attacker can still choose the destination directory and filename prefix. Existing writable destination directories are required because the code does not create parent directories for the generated path.

Attack Path

  1. An attacker obtains the ability to invoke the command-line interface or call run_research() with a chosen topic.
  2. The attacker supplies an absolute or traversal-based topic, for example:
    bash
    python3 research_engine.py /tmp/attacker/report
    
    or:
    bash
    python3 research_engine.py ../../../tmp/attacker/report
    
  3. The topic is converted into a filename without removi ...[truncated 1306 chars]
Remediation
View remediation

Remediation Suggestions

  1. Do not use the raw research topic as a filesystem path. Generate the physical filename from a UUID or another application-controlled identifier.
  2. If a readable topic-derived filename is required, apply a strict allowlist and retain only safe characters such as ASCII letters, digits, underscores, and hyphens.
  3. Explicitly reject absolute paths, path separators, null bytes, . components, and .. components.
  4. Resolve both the base directory and candidate output path, then verify that the candidate remains beneath the base directory before opening it.
  5. Use exclusive creation mode ('x') when overwriting existing reports is unnecessary.
  6. Limit filename length to prevent filesystem errors and resource abuse.

Example hardening approach:

python
import re
from pathlib import Path
from uuid import uuid4

base_dir = Path(RESEARCH_DIR).resolve()
safe_topic = re.sub(r"[^A-Za-z0-9_-]+", "_", topic).strip("_")[:80]
safe_topic = safe_topic or "research"

filename = (
    f"{safe_topic}_{datetime.now().strftime('%Y%m%d_%H%M')}_"
    f"{uuid4().hex[:8]}.md"
)
filepath = (base_dir / filename).resolve()

if filepath.parent != base_dir:
    raise ValueError("Invalid report path")

with filepath.open("x", encoding="utf-8") as f:
    f.write(report)

Where supported, Path.is_relative_to() may also be used to enforce directory containment. Filename sanitization and resolved-path validation should both be applied as defense in depth.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (8)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill states that all research reports are saved to a fixed local path under /home/vken/.openclaw/workspace/research/, but it does not present this as a user-facing warning or explain the storage implications. Silent persistence to a hard-coded directory can expose sensitive research content, cause cross-user data leakage on shared systems, and create predictable locations for discovery or tampering.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly describes scheduled behavior such as hourly searches, daily analysis, and periodic self-optimization, but it does not warn users that enabling the skill may trigger automatic external network activity without per-run user awareness. This can lead to unexpected data egress, privacy issues, policy violations, or resource consumption, especially in environments where outbound access should be tightly controlled.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill persistently stores research topics and browsing metadata to local files without any visible consent, warning, retention policy, or opt-out. Research topics can contain sensitive business plans, incident details, credentials pasted by mistake, or other private material, so silent retention increases privacy and compliance risk if the workspace is later accessed, backed up, or shared.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code writes user research topics and browsing history in plain natural language to persistent markdown files. That creates a durable data exposure surface: anyone with filesystem access, backups, logs, or synced workspace visibility may learn sensitive user interests, queries, and activity patterns, and the natural-language format makes the contents immediately readable and searchable.

Content

No source excerpt is available for this finding.

Tainted flow: 'filepath' from os.environ.get (line 229, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · research_engine.py (reported line 230)May include surrounding context.

python
# 6. 保存报告
    filename = f"{topic.replace(' ', '_')}_{datetime.now().strftime('%Y%m%d_%H%M')}.md"
    filepath = os.path.join(RESEARCH_DIR, filename)
    with open(filepath, 'w', encoding='utf-8') as f:
        f.write(report)
    
    print(f"\n✅ 研究报告已保存: {filepath}")

Tainted flow: 'records_file' from os.environ.get (line 271, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · research_engine.py (reported line 264)May include surrounding context.

python
content += "\n---\n"
    
    # 追加到文件
    with open(records_file, 'a', encoding='utf-8') as f:
        f.write(content)
    
    return len(records)

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language description is entirely in Chinese and presents the skill in a single language without any indication of multilingual support, user opt-in, or region-specific scope. This can violate language/locale policy expectations when skills should not implicitly force one language unless clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The module docstring presents the skill's purpose and behavior entirely in Chinese, and the CLI usage/help text is also Chinese. This creates a language policy concern because the skill effectively forces one language without opt-in, fallback, or justification that it is intended only for a Chinese-speaking or region-specific environment.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.