Back to skill

Security audit

AI Market Research

Security checks for vulnerabilities and agentic risk

Overview

This looks like a market-research helper, but it needs Review because it persists research data and has a confirmed unsafe report-writing bug that can write outside its intended folder.

Install only in a constrained workspace or sandboxed OpenClaw process until the report path handling is fixed. Avoid sensitive internal topics or URLs unless you are comfortable with generated reports and summaries being saved locally and to agentmemory. Treat WeChat/channel delivery and cron examples as opt-in only after reviewing destinations and report contents, and treat the current implementation as a mock/demo rather than a fully integrated market-research pipeline.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
engine.py:212
Finding

Unsanitized Report Parameters Permit Arbitrary File Writes

Content
View full analysis

Vulnerability Details

File Location: engine.py, lines 212-220
Vulnerability Type: Path traversal and unrestricted file write
Risk Level: High

Vulnerable Code

python
async def _generate_report(self) -> Path:
    """生成最终报告文件"""
    logger.info("📝 生成报告中...")
    now = datetime.now()
    timestamp = now.strftime('%Y-%m-%d %H:%M')
    filename = f"{self.topic.replace(' ', '_')}_{self.depth}_{now.strftime('%Y%m%d_%H%M')}.{self.output_format}"
    report_path = OUTPUT_DIR / filename

    # 渲染模板
    template = self._load_template()
    content = self._render_report(template)

    report_path.write_text(content, encoding='utf-8')

Technical Analysis

The topic, depth, and output_format values are obtained from attacker-controllable JSON supplied through standard input. These values are incorporated directly into the output filename without an allowlist, path-component sanitization, canonicalization, or containment check.

Replacing spaces in topic does not remove absolute path prefixes, .. traversal components, or directory separators. In particular, when topic begins with /, the resulting filename is an absolute path. Python's pathlib discards the preceding OUTPUT_DIR when an absolute path is joined:

python
Path("/trusted/output") / "/tmp/report_standard_20260916_1200.markdown"

This resolves to the attacker-selected /tmp/... path rather than a location under /trusted/output. Relative traversal components can similarly escape OUTPUT_DIR when the resulting parent directories exist.

The report body also contains the attacker-controlled topic, so exploitation writes partially attacker-controlled content. The timestamp makes targeting an existing file less convenient, but it does not prevent unauthorized file creation outside the intended directory. Predictable timestamps and attacker-controlled suffix components can also increase overwrite opportunities.

...[truncated 1938 chars]

Remediation
View remediation

Remediation Suggestions

  1. Use a server-generated filename. Do not use user input as a filesystem path component. Generate a random identifier or digest and store the human-readable topic only inside the report.

    python
    from uuid import uuid4
    
    allowed_formats = {"markdown": "md", "html": "html", "json": "json"}
    extension = allowed_formats.get(self.output_format)
    if extension is None:
        raise ValueError("Unsupported output format")
    
    filename = f"research_{uuid4().hex}.{extension}"
    
  2. Strictly validate enumerated parameters. Allow only documented values for depth and output_format:

    python
    if self.depth not in {"quick", "standard", "deep"}:
        raise ValueError("Invalid depth")
    if self.output_format not in {"markdown", "html", "json"}:
        raise ValueError("Invalid output format")
    
  3. If the topic must appear in the filename, convert it to a safe slug. Allow only a conservative set of characters, impose a length limit, and reject empty results. Do not merely remove spaces.

    python
    import re
    
    slug = re.sub(r"[^A-Za-z0-9_-]+", "_", self.topic).strip("_")[:80]
    if not slug:
        slug = "research"
    
  4. Enforce canonical path containment before writing.

    python
    base = OUTPUT_DIR.resolve()
    report_path = (base / filename).resolve()
    
    if report_path.parent != base:
        raise ValueError("Report path escapes the output directory")
    
  5. Prevent unintended overwrites. Use exclusive file creation where replacement is unnecessary:

    python
    with report_path.open("x", encoding="utf-8") as handle:
        handle.write(content)
    
  6. Run the Skill with least privilege. Restrict the process account to the report directory and required memory interfaces. Avoid granting write access to application code, startup configuration, credentials, or shared executable directories.

...[truncated 191 chars]

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (23)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document content is written in Chinese from the outset, which effectively imposes a specific language on readers. There is no indication that this skill is intended only for a Chinese-speaking audience, nor any opt-in or alternative language option documented here.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · PUBLISHING.md (reported line 63)May include surrounding context.

md
### 创建 Release

1. 进入仓库 → Releases → Create a new release
2. Tag 版本:`v0.1.0`
3. Release title: `AI Market Research Skill v0.1.0`
4. Description: 复制 `CHANGELOG.md` 中对应版本内容

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · PUBLISHING.md (reported line 63)May include surrounding context.

md
### 创建 Release

1. 进入仓库 → Releases → Create a new release
2. Tag 版本:`v0.1.0`
3. Release title: `AI Market Research Skill v0.1.0`
4. Description: 复制 `CHANGELOG.md` 中对应版本内容

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The feature list presents current capabilities such as using crawl4ai for structured extraction, integrating trendradar for hotspot tracking, and generating SWOT/competitor analysis. However, the roadmap later says '真实 MCP 调用' and 'product-research 深度集成' are still planned for future versions, which directly contradicts the earlier documentation about what the skill currently does.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README promotes automatic crawling, trend monitoring, historical comparison, and channel-oriented reporting without warning that data may be collected from external sources, processed, stored, and potentially redistributed. In an agent skill context, this omission can cause users to feed sensitive topics or internal URLs into a workflow that performs outbound requests and retains artifacts, increasing privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The scheduled task example automatically delivers generated research to WeChat, an external messaging channel, without warning that report contents may include sensitive internal prompts, collected source material, or confidential analysis. In an automation setting, users may enable this as-is and unintentionally exfiltrate business-sensitive information on a recurring basis.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The default configuration includes automatic push to a messaging channel, which normalizes outbound data sharing without explicit consent or review. Default-on external transmission is especially risky for research workflows because reports can aggregate scraped content, internal context, and comparative history, making accidental disclosure more likely and more damaging.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The requirements explicitly request persistent memory read/write access and report file generation, but the document does not warn users that the skill may retain collected data or modify local storage. In a market-research workflow, gathered web content, prompts, and generated analysis may contain sensitive business information, so silent persistence increases the risk of unintended retention, later reuse, or disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The configuration exposes an auto_push_to_channel option that can automatically send generated reports to an external destination, but the requirements do not warn users about outbound data transfer. Because this skill aggregates crawled external data, analysis results, and possibly historical memory, automatic pushing can leak sensitive research or internal business context to third-party channels without sufficiently informed consent.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill advertises capabilities that include writing reports, artifacts, and memory entries, but it does not declare any explicit tool scope such as permissions or allowed-tools. In an agent environment, this weakens least-privilege controls and can let a caller invoke file-writing side effects without clear user-visible authorization boundaries.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill automatically persists research outputs and findings into agentmemory/vector storage, but this side effect is not surfaced as a prominent warning in the skill description or invocation contract. Research tasks may include proprietary URLs, sensitive conclusions, or user-provided topics, so silent persistence increases the risk of unintentional retention and later disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill performs external data collection and platform monitoring across websites and social platforms, yet it does not prominently warn users about the resulting network activity and privacy implications. In practice, this can trigger scraping of third-party content, collection of personal data, or contact with external services without informed operator consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill can push generated results to a configured WeChat recipient, but this outbound exfiltration path is not called out as a prominent warning. If reports contain sensitive market intelligence or internal analysis, automatic messaging can disclose that information to unintended recipients or external accounts.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document presents crawl4ai and trendradar integration as active behavior while the roadmap says real MCP calls are still future work, creating a misleading security and operational model. This can cause users or downstream agents to rely on network collection, monitoring, or persistence behavior that may not actually exist, or to misjudge what data is transmitted and what controls are needed.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring describes a full orchestration pipeline that calls crawl4ai, trendradar, and a product-research analysis framework and stores results to memory. In the implementation, trendradar and crawl4ai are replaced with simulated/mock data paths (L76-L77, L93-L109, L143-L146), source auto-discovery is unimplemented (L134-L139), and the analysis step explicitly uses a simplified placeholder instead of the stated framework (L176-L181).

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Several docstrings/comments claim operational behavior such as calling crawl4ai, trendradar, and the product-research analysis framework (for example L71, L112, L176-L179). The actual code instead uses simulation helpers, mock return values, or simplified placeholder logic (L76-L77, L93-L109, L141-L151, L180-L181), which materially contradicts the stated intent of those routines.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill writes a report containing user-supplied topic, collected source content, trends, and historical references to a fixed local path without explicit user consent or a visibility control. In an agent environment, silent persistence can expose sensitive research interests or gathered data to other local users, later processes, backups, or unintended retention policies.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The engine stores a research summary and topic into agentmemory by default, again without explicit user acknowledgement. In this context, topics can reveal confidential business strategy, customer names, or unreleased initiatives, and memory storage may make that data retrievable by unrelated future tasks or agents.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file mixes English headings with Chinese substantive entries, effectively requiring readers to understand Chinese to use the changelog details. Under the policy, forcing a specific language without opt-in or justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file contains contributor-facing instructions primarily in Chinese, but does not indicate that the skill or project is intentionally region-specific or offer an alternative language option. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

SQP-3 applies to language or locale policy violations when a skill effectively forces a specific language without user opt-in. This README presents its core description and examples in Chinese and does not indicate multilingual support or that Chinese is a documented, justified requirement.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The feature list describes historical comparison as an automatic capability, only noting that it requires agentmemory. Later, the troubleshooting section states that when agentmemory is not installed the skill warns and skips the comparison, meaning the advertised behavior is not actually guaranteed. This is a documentation-level contradiction about operational behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring, log messages, and final user-facing output are written entirely in Chinese, which imposes a specific language without any visible opt-in or locale selection. This can violate language/locale policy when the skill is intended for general use.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.