Back to skill

Security audit

data-analysis-partner

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its data-reporting purpose, but generated HTML reports can run injected JavaScript from malicious spreadsheet or user-supplied content.

Install only if you are comfortable reviewing generated reports before opening or sharing them. Avoid using untrusted CSV/Excel files, avoid sensitive data unless reports stay in a private folder, and prefer pinned Python dependencies plus a locally bundled ECharts library.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
index.js:302
Finding

Stored JavaScript Injection in Generated HTML Reports

Content
View full analysis
{ const typeLabel = { numeric: '数值', categorical: '分类', datetime: '时间', text: '文本', boolean: '布尔', empty: '空列', }[c.type] || `${c.type}`; const missingCell = c.missing_pct > 20 ? `${c.missing_pct}%` : c.missing_pct > 0 ? `${c.missing_pct}%` : `0%`; return ` ${c.name} ${typeLabel} ${missingCell} ${c.unique.toLocaleString()} ${c.sample.join(" / ")} `; }) .join(""); const insightItems = insights .map((ins) => { const color = ins.level === "warning" ? "#fff3e0" : "#e8f4fd"; const border = ins.level === "warning" ? "#ffa940" : "#4e9bff"; return `
${ins.icon} ${ins.text}
`; }) .join(""); const chartDivs = charts .map((chart, i) => { const option = buildEChartsOption(chart); if (!option) return ""; const optionJson = JSON.stringify(option); return `
Remediation
View remediation
`, `"`, and `'`. 2. Treat filenames, requirements, column names, cell samples, insight text, chart labels, and type labels as untrusted input. 3. Do not interpolate serialized objects directly into executable inline scripts. 4. Place chart configuration in an `application/json` element or separate JSON file, parse it with `JSON.parse()`, and escape characters significant to the HTML parser. At minimum, replace `<`, `>`, `&`, U+2028, and U+2029 with Unicode escapes before embedding. 5. Prefer creating report elements through safe DOM APIs and assigning untrusted strings with `textContent`. 6. Add a restrictive Content Security Policy. Avoid `unsafe-inline`; move scripts to a reviewed local file or use hashes/nonces. 7. Add regression tests using payloads in filenames, requirements, column names, and cell values, including HTML event handlers and `` sequences. 8. Consider bundling ECharts locally so that a restrictive network policy can be applied to generated reports. ]]>

T08 · Insecure Dependencies

Warning
Location
SKILL.md:93
Finding

Unpinned Third-Party Python Dependency Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (14)

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation claims the report is self-contained while later admitting it performs external CDN requests, creating a trust and privacy gap for users handling local data. In the context of a data-analysis skill, users may reasonably open reports containing sensitive business data, so undocumented network access is more dangerous than in a purely public-content skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The README represents the output as a self-contained HTML report, but elsewhere discloses that ECharts is fetched from cdn.jsdelivr.net at view time. This is a security-relevant misrepresentation because opening the report triggers external network access, which can leak usage metadata and violates offline/self-contained expectations in sensitive environments.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README suggests invoking the skill with the phrase "帮我分析一下这个销售数据" followed by a file path, which overlaps with common conversational language and does not clearly define when the skill should or should not activate. The surrounding text says OpenClaw will automatically call the tool, but provides no negative examples or explicit trigger boundaries to prevent unintended invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description is written as “智能数据分析 Skill” and the file consistently specifies the skill behavior only in Chinese, without offering a language choice or stating that output language follows user preference. This can violate language/locale policy where a skill should not impose a specific language absent user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger conditions are broad enough to activate on many generic requests like 'analyze this file' or 'generate a report,' which can cause the agent to invoke this skill when the user did not specifically intend this workflow. In context, that can lead to unintended processing of uploaded files, unnecessary dependency installation, or generation of HTML reports with external CDN references, expanding both privacy and execution risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The report generation explicitly formats timestamps using the zh-CN locale and the user-facing report text is written in Chinese throughout the file. This enforces a specific language/locale experience without any opt-in, user selection mechanism, or documented regional justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The report is described as self-contained, but it loads ECharts from a public CDN at runtime. This creates a supply-chain and privacy risk: opening the report triggers a network request and the report's integrity depends on third-party content remaining trustworthy and available.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Opening the generated HTML causes the browser to fetch ECharts from a remote CDN, but this behavior is not disclosed here. That can leak metadata such as IP address and access timing, and it turns a local report into a networked artifact with external dependency risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This skill writes an HTML file to a user-controlled or default directory without any visible disclosure in this file. While expected for report generation, silent file creation can surprise users, expose sensitive analyzed data at rest, or place files in shared locations such as Downloads.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill passes user-supplied file paths and analysis requirements into a Python subprocess with no user-facing disclosure in this file. Although spawnSync is used safely without a shell, the subprocess may process untrusted files and large inputs, increasing risk from parser bugs, malicious spreadsheets, or denial-of-service conditions in the Python data stack.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest’s user-facing description and tool parameter descriptions are written entirely in Chinese, which imposes a language/locale expectation without offering a user choice or documenting that the skill is region-specific. This matches the policy concern for language/locale constraints that are not opt-in or clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module docstring and all user-facing messages in the script are written in Chinese, which indicates the skill is designed to operate in a single language. The file does not provide any opt-in, locale selection, or justification that the skill is intentionally region-specific, so it conflicts with the language-choice policy.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The insight text states that for datasets over 100,000 rows the skill is using sampling for analysis, but elsewhere the script still reports summary-style conclusions such as duplicate counts and missingness based only on the sampled dataframe. Because basic_info is computed after df is replaced by a 50,000-row sample, these outputs describe the sample while the surrounding messaging can be read as describing the full dataset.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The primary skill description is presented in Chinese and the usage example is also Chinese, but the README does not indicate that the language is optional or user-selectable. This can violate language/locale policy when a skill effectively assumes a specific language without explicit opt-in or documented regional scope.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
index.js:18