Back to skill

Security audit

Data Analyzer

Security checks for vulnerabilities and agentic risk

Overview

This skill is a basic user-directed data analyzer with some documentation and output-safety shortcomings, but no evidence of hidden execution, exfiltration, persistence, or privilege abuse.

Install only if you are comfortable with a simple local data-analysis script. Use trusted datasets for HTML reports, prefer JSON output for untrusted inputs, choose output filenames carefully to avoid overwriting existing files, and be aware that several advertised features are not actually included.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/analyze_data.py:140
Finding

Stored HTML Injection in Generated Analysis Reports

Content
View full analysis
Generated: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}

Source: {self.input_file}

Summary

Total Records: {self.results.get('basic_stats', {}).get('count', 'N/A')}

Columns: {', '.join(map(str, self.results.get('basic_stats', {}).get('columns', [])))}

Grouped Analysis

""" # Add grouped data grouped = self.results.get('grouped', {}) for key, value in grouped.items(): html += f" \n" ``` ### Technical Analysis The report generator directly interpolates the input path, dataset column names, group keys, and grouped values into an HTML document without applying context-appropriate HTML escaping. An attacker who controls a CSV, Excel, or JSON dataset can place HTML markup or JavaScript-capable elements in a column name or grouped field value. When the analyzer generates an HTML report, the supplied markup is stored verbatim in the output document. It can then be interpreted by a browser when the report is opened. For example, a malicious grouping value could contain an image element with an error event handler. Because the value is inserted directly between `
GroupCount
{key}{value}
` tags, the browser treats it as markup rather than plain text. The input filename is also inserted without escaping. This provides another injection surface where an attacker can influence the filename used to deliver a dataset, subject to filesystem naming restrictions. ### Attack Path 1. An attacker creates or modifies a supported dataset. 2. The attacker inserts malicious HTML into a column name or a ...[truncated 1265 chars]
Remediation
View remediation
{safe_key}{safe_value}\n" ``` Additional hardening should include: 1. Use a template engine with automatic HTML escaping instead of assembling documents through string concatenation. 2. Escape values according to their exact output context, including text nodes and attributes. 3. Consider adding a restrictive Content Security Policy to generated reports, such as disallowing inline scripts and remote content. 4. Add regression tests containing script elements, event-handler attributes, malformed tags, encoded payloads, and malicious column names. 5. Treat filenames and filesystem paths as untrusted display values and escape them as well. ]]>

T08 · Insecure Dependencies

Note
Location
README.md:15
Finding

Unpinned Package Execution in the Documented Installation Command

Content
View full analysis
Remediation
View remediation
install yinan-data-analyzer ``` Additional supply-chain hardening should include: 1. Document the expected package registry and publisher identity. 2. Use a lockfile or another reproducible installation mechanism where supported. 3. Verify package integrity through registry integrity metadata, checksums, signatures, or provenance attestations. 4. Review package lifecycle scripts and release changes before updating the pinned version. 5. Avoid instructing users to run the installation command with elevated privileges. 6. Consider installing the reviewed CLI as a project dependency and invoking the lockfile-resolved binary. ]]>
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (4)

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The code generally aligns with the high-level theme of analyzing CSV/Excel/JSON data and generating simple reports, so the primary purpose is related. However, the declared description materially overstates capabilities. The implementation supports file loading, basic stats, grouping, and HTML/JSON report output, but there is no chart generation, no pivot table functionality, and only limited statistical analysis. Reporting is present but basic rather than the more comprehensive automated reporting implied by the description. Additionally, the CLI exposes a CSV output format choice without implementing that branch. This is best classified as a description-behavior mismatch due to significant missing declared features rather than any hidden undeclared behavior.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
96% confidence
Finding

The README instructs users to run npx clawhub install yinan-data-analyzer without pinning a specific version, which can cause users to fetch and execute whatever package version is current at install time. If the upstream package, dependency chain, or publisher account is compromised, users could install malicious code, and the data-analysis context makes this plausible because the skill is intended to process business datasets and may be used in trusted environments.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill documentation indicates file read/write behavior through its scripts, but it does not declare any explicit tool scope or permissions. In an agent environment, undocumented file capabilities can cause the skill to be granted broader filesystem access than reviewers or policy layers expect, increasing the chance of unintended data exposure or file modification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The examples instruct users to generate output files but do not warn that existing files may be overwritten. In a file-writing skill, this can lead to accidental destruction or replacement of reports or other user data, especially when an agent autonomously selects output paths.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.