Back to skill

Security audit

Data Analyst

Security checks for vulnerabilities and agentic risk

Overview

This is a local data-analysis skill with noteworthy install and file-output hygiene issues, but the artifacts do not show hidden data theft, persistence, or destructive behavior.

Install this in a virtual environment, do not run sudo pip, and use it only on data you are authorized to analyze. Expect it to create or overwrite summary, cleaned-data, chart, and report files beside the input file, so avoid shared or untrusted directories and review reports before emailing or uploading them.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T08 · Insecure Dependencies

Warning
Location
install.sh:26
Finding

Unpinned Dependency Installation Enables Supply-Chain Substitution

Content
View full analysis
=1.3.0 numpy>=1.21.0 openpyxl>=3.0.0 matplotlib>=3.4.0 seaborn>=0.11.0 ``` ### Technical Analysis The installation script requests packages by name without exact versions or package hashes. Consequently, the installed code depends on the state of the configured Python package index at installation time. Python packages may execute packaging hooks during installation, and their modules subsequently run with the privileges of the user invoking the skill. The absence of an isolated virtual environment also means the installation modifies the active Python environment. The package names are established packages, and the audit found no direct evidence of typosquatting or a currently malicious dependency. The security issue is the lack of reproducibility and integrity enforcement. Relevant compromise scenarios include: - A compromised upstream publisher account or malicious future package release. - A user or system configured to use an untrusted `PIP_INDEX_URL` or additional package index. - Dependency resolution selecting a vulnerable or behaviorally incompatible future release. - Installation into a privileged or shared Python environment, increasing the effect of package compromise. ### Attack Path 1. An attacker compromises a package publisher, package-index account, or package source configured on the target. 2. The attacker publishes a malicious version satisfying the unrestricted dependency request. 3. A user runs `install.sh`. 4. `pip3 install` resolves and downloads the attacker-controlled release. 5. ...[truncated 776 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
tools/analyze.py:274
Finding

Predictable Analysis Outputs Permit Symlink-Based File Clobbering

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
test.sh:13
Finding

Test Suite Uses Predictable Shared Temporary Files

Content
View full analysis
/tmp/test_data.csv << EOF id,name,age,salary,department,join_date 1,Alice,25,50000,Engineering,2022-01-15 2,Bob,30,60000,Marketing,2021-06-20 3,Charlie,,55000,Engineering,2022-03-10 4,Diana,28,62000,Marketing,2021-11-05 5,Eve,35,75000,Engineering,2020-08-12 6,Frank,30,60000,Marketing,2021-06-20 7,Grace,29,58000,Sales,2022-02-28 8,Henry,32,68000,Engineering,2021-04-15 9,Ivy,27,52000,Sales,2022-05-20 10,Jack,,65000,Marketing,2021-09-30 EOF ``` Cleanup uses predictable wildcard patterns: ```bash rm -f /tmp/test_data*.csv /tmp/test_data*.json /tmp/test_data*.png /tmp/test_data*.md ``` ### Technical Analysis The test suite creates files directly in the globally shared `/tmp` directory using fixed names. Shell redirection follows symbolic links. On a multi-user system, another local user can create `/tmp/test_data.csv` as a symbolic link before the test begins. The installer automatically invokes this test suite: ```bash bash "$SCRIPT_DIR/test.sh" ``` The analyzer then creates additional predictable files such as: - `/tmp/test_data_summary.json` - `/tmp/test_data_cleaned.csv` - `/tmp/test_data_report.md` - `/tmp/test_data_*.png` These paths are also vulnerable to pre-positioned symbolic links. The final wildcard cleanup can remove unrelated files matching `/tmp/test_data*` that belong to the invoking user or the same automated environment. ### Attack Path 1. An attacker with access to the shared temporary directory predicts that the test suite will run. 2. The attacker creates a symbolic link at one of the fixed paths: ```bash ln -s /path/writable/by-victim/target /tmp/test_data.csv ``` 3. The victim runs `install.sh`, which invokes `test.sh`. 4. Shell redirection follows the link and ov ...[truncated 729 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
tools/report.py:53
Finding

Unescaped Dataset Content Is Embedded in Generated Markdown Reports

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (38)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · INSTALL.md (reported line 69)May include surrounding context.

python3 ~/.openclaw/skills/data-analyst/tools/analyze.py /tmp/test.csv

Cleanup

rm /tmp/test.csv

text

## Troubleshooting

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Even though the path is scoped to the skill directory, 'rm -rf ~/.openclaw/skills/data-analyst' is a destructive command that permanently deletes files and can cause data loss if the path is mistyped, expanded unexpectedly, or copied into the wrong environment. In install documentation, destructive shell commands deserve special caution because users often paste them without review.

Content

Scanner excerpt · INSTALL.md (reported line 184)May include surrounding context.

bash
# Remove skill
rm -rf ~/.openclaw/skills/data-analyst

# Remove dependencies (optional)
pip3 uninstall pandas openpyxl matplotlib seaborn

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Even though the path is scoped to the skill directory, 'rm -rf ~/.openclaw/skills/data-analyst' is a destructive command that permanently deletes files and can cause data loss if the path is mistyped, expanded unexpectedly, or copied into the wrong environment. In install documentation, destructive shell commands deserve special caution because users often paste them without review.

Content

Scanner excerpt · INSTALL.md (reported line 184)May include surrounding context.

bash
# Remove skill
rm -rf ~/.openclaw/skills/data-analyst

# Remove dependencies (optional)
pip3 uninstall pandas openpyxl matplotlib seaborn

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · test.sh (reported line 80)May include surrounding context.

sh
# Cleanup
echo ""
echo "🧹 Cleaning up..."
rm -f /tmp/test_data*.csv /tmp/test_data*.json /tmp/test_data*.png /tmp/test_data*.md

echo ""
echo "✅ All tests passed!"

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · INSTALL.md (reported line 59)May include surrounding context.

Test 3: Quick Analysis

bash
# Create test file
echo "name,age,city
Alice,25,NYC
Bob,30,LA

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · INSTALL.md (reported line 96)May include surrounding context.

bash
# Solution: Use python instead
# Or install Python 3
sudo apt-get install python3  # Ubuntu/Debian
brew install python3          # macOS

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
84% confidence
Finding

Using 'sudo pip3 install' is dangerous because pip executes package installation logic with root privileges, increasing the blast radius of a compromised dependency or typo-squatted package. In an installation guide for a skill, encouraging root-level Python package installs is unnecessarily risky when user-local or virtualenv installs are available.

Content

Scanner excerpt · INSTALL.md (reported line 137)May include surrounding context.

bash
# System-wide install
sudo pip3 install pandas openpyxl matplotlib seaborn

# Or user install
pip3 install --user pandas openpyxl matplotlib seaborn

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file includes an irreversible removal command (rm -rf ~/.openclaw/skills/data-analyst) but does not explicitly warn the user that it will permanently delete the installed skill directory. Under the markdown-specific warning criteria, destructive actions affecting user/system state should be disclosed clearly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The marketing copy encourages users to upload and analyze CSV data but does not disclose that submitted datasets may contain personal, confidential, or regulated information and may be processed by the service. This can mislead users into sharing sensitive data without informed consent, creating privacy, compliance, and trust risks even though the file itself is only promotional content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The social and email templates repeatedly promote automated CSV analysis as frictionless and 'free to try' without any caution that uploaded files may expose personal, business-sensitive, or regulated data. Because these are outward-facing acquisition materials, the omission increases the chance that non-technical users submit sensitive datasets under incomplete privacy expectations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The file makes a strong privacy/security claim that 'all processing is local' and 'we never see your data,' while also advertising account signup, contact forms, email support, API access, and enterprise flows that imply some remote service interaction. Even if the core analysis may run locally, this wording can mislead users into sharing sensitive data under false assumptions about data exposure, retention, or transmission.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · PROJECT_SUMMARY.md (reported line 127)May include surrounding context.

md
### Immediate (This Week)

1. **Register ClawhHub Account**
   - [ ] Create account on clawhub.ai
   - [ ] Verify email
   - [ ] Set up profile

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README promotes automatic cleaning and creation of output files without clearly warning users that running the tool may alter datasets or generate derived artifacts on disk. In a data-analysis context, this can cause unintended data loss, overwrites, privacy exposure through copied data, or misuse of cleaned/exported files when users assume the operation is read-only.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises executable shell usage and file-writing behavior through its documented commands and install scripts, but it does not declare any explicit tool scope such as permissions or allowed-tools. That creates an authorization boundary problem: an agent may invoke shell/file operations more broadly than reviewers or runtime policy expect, increasing the chance of unintended command execution or filesystem modification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The report-generation trigger phrase '生成报告' is broad enough to overlap with many unrelated requests, making accidental activation plausible. In this skill, accidental activation matters because the skill is positioned to process local files and run analysis commands, so a generic request could unintentionally trigger code-assisted operations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The visualization trigger phrases include very generic terms like '画图', '图表', and '可视化', which can match ordinary conversation outside a true data-analysis request. Overbroad activation can cause the skill to engage unexpectedly and then run local tooling on files or invoke shell-backed workflows when the user did not clearly intend that behavior.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · examples/demo_sales_analysis.md (reported line 11)May include surrounding context.

md
1. Clean the data
2. Analyze sales trends
3. Generate visualizations
4. Create a report for stakeholders

## Sample Data

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · examples/demo_sales_analysis.md (reported line 208)May include surrounding context.

md
2. **Automation**
   - Schedule daily/weekly reports
   - Set up alerts for anomalies
   - Create dashboard

3. **Advanced Analytics**
   - Sales forecasting

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation recommends piping a generated report directly into email without an adjacent warning that reports may contain sensitive or regulated data. In a data-analysis skill, reports commonly include row counts, sample values, quality issues, and derived insights that can expose internal or personal information if shared broadly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The cloud upload example normalizes sending generated reports to S3 without warning about privacy, bucket permissions, retention, or data residency requirements. Because this skill processes Excel/CSV/JSON business datasets, generated reports may contain confidential findings or sensitive fields that become exposed through misconfigured or overly broad cloud storage.

Content

No source excerpt is available for this finding.

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is uploaded to cloud storage (S3 / GCS / Azure Blob). This may be a legitimate backup or exfiltration to an external bucket. Manual review is recommended.

Content

Scanner excerpt · references/report_generation.md (reported line 222)May include surrounding context.

md
# Upload to cloud
{baseDir}/tools/analyze.py data.csv --report
aws s3 cp report.md s3://reports/

# Convert to presentation
pandoc report.md -o slides.pptx

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The tool is presented as an analysis utility, but it automatically writes a derived summary file beside the input dataset without requiring explicit opt-in. In an agent/automation context, this can unexpectedly persist potentially sensitive metadata, alter a read-only workflow assumption, and write into directories the user did not intend to modify.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Automatically saving a summary JSON without advance disclosure is a real security/privacy concern because summaries can still contain sensitive schema details, column names, value distributions, and missingness patterns. In enterprise data analysis workflows, silently creating such artifacts next to the source file can violate user expectations, leak metadata, or leave recoverable traces on shared systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

When --clean is used, the tool writes a cleaned CSV to disk as a side effect, but the interface does not clearly foreground that this will create a new data file. In a skill meant for automated data handling, this can propagate sensitive records into additional files, increase data exposure surface, and create retention/compliance issues if run in shared or monitored directories.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.destructive_delete_command

Documentation contains a destructive delete command without an explicit confirmation gate.

Warn
Code
suspicious.destructive_delete_command
Location
INSTALL.md:184