Back to skill

Security audit

sn-da-excel-workflow

Security checks for vulnerabilities and agentic risk

Overview

This Excel workflow is mostly aligned with spreadsheet analysis, but it over-activates and writes sensitive spreadsheet-derived data to predictable local files without enough scoping or safeguards.

Review this skill before installing if you handle sensitive spreadsheets. Use it only on intended Excel/CSV analysis tasks, avoid multilingual cleaning unless you explicitly want non-Chinese text removed, and run it in an isolated workspace where fixed output and cache filenames cannot overwrite important files or expose data across tasks.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
capability/excel-result-export/formatted-export/SKILL.md:4
Finding

Unpinned Automatically Provisioned Python Dependencies

Content
View full analysis

Vulnerability Details

File Location: capability/excel-result-export/formatted-export/SKILL.md:4
Vulnerability Type: Unpinned third-party dependencies
Risk Level: Medium

Vulnerable Code

yaml
metadata: "{\"nanobot\": {\"requires\": {\"pip\": [\"pandas\", \"pyarrow\", \"openpyxl\"]}}}"

Technical Analysis

The Skill declares automatically provisioned Python packages without exact version constraints or integrity hashes. Consequently, dependency resolution may install releases that differ from those present when the Skill was reviewed.

Package installation and import can execute package-controlled Python code. If a listed package, its distribution channel, or the configured package index is compromised, the Agent may install and execute malicious code under the runtime account. No evidence indicates that the named packages are currently malicious; the risk arises from unconstrained future dependency resolution and the absence of integrity verification.

Attack Path

  1. An attacker compromises a listed package release, its distribution account, or an untrusted package index used by the runtime.
  2. The attacker publishes a malicious release that still satisfies the unconstrained dependency declaration.
  3. The Skill environment resolves and installs that release.
  4. Malicious installation hooks or imported package initialization code executes with the Agent runtime's permissions.
  5. The malicious package can access data and resources available to that runtime, including spreadsheets being processed and writable output directories.

Impact Assessment

Successful exploitation could allow arbitrary code execution with the privileges of the Agent runtime. The accessible scope may include input spreadsheet contents, generated reports, environment variables available to the process, and files writable by the runtime account. This declaration does not itself grant elevated system privileges, so imp ...[truncated 58 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin every dependency to a reviewed exact version, such as package==version.
  • Maintain a lockfile generated from an approved dependency set.
  • Require cryptographic hashes for downloaded distributions.
  • Install packages exclusively from a trusted, explicitly configured package index.
  • Prefer prebuilt, immutable runtime images containing audited dependencies rather than installing packages when the Skill runs.
  • Continuously scan locked dependencies for known vulnerabilities and review changes before updating versions.
  • Run dependency installation and Skill execution inside a least-privilege sandbox without unnecessary secrets or filesystem access.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:57
Finding

Predictable Shared Parquet Cache Paths

Content
View full analysis

Vulnerability Details

File Locations:

  • SKILL.md:57-61
  • capability/excel-reading/multi-sheet-reading/SKILL.md:30-33
  • capability/excel-reading/large-excel-reading/SKILL.md:25-29

Vulnerability Type: Predictable shared temporary files and unsafe cache reuse
Risk Level: Medium

Vulnerable Code

SKILL.md:57-61:

python
parquet_path = "/tmp/_auto_parquet.parquet"
df = pd.read_excel(file_path, sheet_name=target_sheet)
df.to_parquet(parquet_path, engine="pyarrow")
del df; gc.collect()
df = pd.read_parquet(parquet_path)

capability/excel-reading/multi-sheet-reading/SKILL.md:30-33:

python
df = pd.read_excel(file_path, sheet_name=sheet_names[0])
parquet_path = '/tmp/temp_data.parquet'
df.to_parquet(parquet_path, engine='pyarrow')
df = pd.read_parquet(parquet_path)

capability/excel-reading/large-excel-reading/SKILL.md:25-29:

python
parquet_path = "/mnt/data/temp.parquet"
df = pd.read_excel(file_path, sheet_name=0)
df.to_parquet(engine='pyarrow', path=parquet_path)
df = pd.read_parquet(parquet_path)

Technical Analysis

These workflows use constant, globally predictable filenames for intermediate Parquet data. The files are written and subsequently read without creating a private per-task directory, validating file ownership, rejecting symbolic links, or ensuring exclusive creation.

Concurrent Skill invocations can overwrite each other's caches, causing cross-task data corruption or disclosure. If another process can write to /tmp or the relevant /mnt/data directory, it may pre-create the target as a symbolic link or replace the cache between the write and read operations. The exact feasibility of symbolic-link exploitation depends on directory permissions, operating-system protections, and the behavior of the Parquet writer, but predictable shared paths unnecessarily expose this attack surface.

Attack Path

Cross-task collision:

1 ...[truncated 1336 chars]

Remediation
View remediation

Remediation Suggestions

  • Create an unpredictable, private directory for every task using tempfile.TemporaryDirectory().
  • Generate cache filenames with tempfile.NamedTemporaryFile() or a cryptographically unpredictable identifier.
  • Set restrictive permissions so only the current runtime user can access the cache directory and files.
  • Avoid placing intermediate sensitive data in shared output locations such as /mnt/data.
  • Reject symbolic links and verify that the cache remains a regular file owned by the expected runtime user before reading it.
  • Prevent time-of-check/time-of-use replacement by retaining secure file handles or using exclusive creation where supported.
  • Delete intermediate caches in a finally block, including after failed processing.
  • Use a unique task identifier in cache paths and prevent concurrent tasks from sharing mutable intermediate state.
  • Consider avoiding disk caching entirely when the dataset can be processed safely in memory.
Vulnerability Patterns
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (108)

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill declares extremely broad trigger phrases such as generic 'data analysis', 'data cleaning', 'data statistics', and instructs the agent to proactively use this skill instead of normal reasoning. This can cause inappropriate tool routing for many ordinary requests, increasing the chance of unnecessary file access, over-collection of user data, and suppression of safer or more specific workflows.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The document mandates CJK font setup for chart generation and the workflow description is framed around Chinese text handling, including preserving only Chinese characters in cleaning examples. This can imply a fixed language/locale behavior without offering the user a choice or stating that Chinese-specific handling is optional based on user needs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language description and all instructional text are written in Chinese, and there is no indication that the skill is region-specific or that users may choose another language. Under the policy, forcing a specific language without opt-in is a locale/language policy violation.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
80% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · capability/excel-cell-coloring/duplicate-value-coloring/SKILL.md (reported line 48)May include surrounding context.

Step2 生成对比报告,并使用 openpyxl 对异常值(非标准系数)进行红色高亮标记。

python
from openpyxl import Workbook
from openpyxl.styles import PatternFill

output_path = 'comparison_report.xlsx'
wb = Workbook()

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language description specifies the skill in Chinese and does not indicate that the user may choose another language or locale. Under SQP-3, forcing a specific language without opt-in is a policy violation unless the locale constraint is documented and justified, which is not present here.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
80% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · capability/excel-cell-coloring/outlier-coloring/SKILL.md (reported line 95)May include surrounding context.

md
# 标注异常单元格
for err in error_cells:
    ws_error[err['coord']].fill = red_fill

output_path = "highlighted_report.xlsx"
wb.save(output_path)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest description is entirely in Chinese and describes the skill behavior in a way that implies Chinese-language reporting and labels, but it does not offer any language choice or explain that the skill is intentionally region-specific. This is a natural-language locale policy concern because it imposes a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Natural-language policy review applies to all file types. The description forces a specific language/locale presentation for the skill without any opt-in, fallback, or explanation that the skill is intended only for a Chinese-language context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file forces a specific language for the skill’s user-facing description and step instructions. Under the policy, language constraints should offer user choice or be clearly documented as region-specific and justified; neither is present here.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description says this skill compares two categories and generates charts, which suggests a focused comparison capability. However, Step1 adds a broader Excel-workflow behavior by reading all sheet names and counting total rows to decide processing strategy, which goes beyond the described comparison-analysis scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description, column labels, chart titles, exclusion values, and reporting text are all hard-coded in Chinese, and Step4 explicitly configures Chinese font support. There is no indication that the skill is region-specific or that users can opt into another language, which may violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This markdown skill includes user-facing natural-language description and step instructions entirely in Chinese. The file does not indicate that the skill is region-specific or provide the user with a language/locale choice, which can violate a language/locale policy requiring opt-in or documented justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The parent skill manifest emphasizes multi-Sheet Excel reading, large-file optimization, data cleaning, conditional extraction, cross-Sheet aggregation, and Excel/CSV export with download links. This documented skill instead reads one Excel sheet via a fixed path, performs trend/grade/forecast calculations, saves a PNG locally, and prints console summaries, which is a narrower and different behavior than the manifest-described workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill metadata and steps are written to operate in Chinese, including explicitly configuring Chinese fonts to ensure chart rendering. This imposes a specific language/locale behavior without user opt-in or an alternative, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description and all user-facing instructions are written only in Chinese, and the code explicitly configures Chinese fonts to ensure Chinese chart rendering. This imposes a specific language/locale choice without offering the user an alternative or documenting it as a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest name and description describe counting total rows across multiple sheets and choosing a processing strategy based on dataset size. In the actual code, L15-L18 reads only one specified sheet (Sheet1), and the remaining steps perform filtering, deduplication, and export; there is no multi-sheet traversal or scale/threshold logic anywhere in the file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill description hard-codes cleaning logic that retains only Chinese characters, which can silently destroy or alter user data in multilingual datasets without clear opt-in or contextual justification. In a generic Excel cleaning workflow, this is dangerous because it can cause irreversible integrity loss, misclassification, or compliance issues when identifiers, codes, names, or non-Chinese content are stripped.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill writes a local Excel file and exposes a file:// download path without any warning, consent flow, or safeguards around overwriting existing files. In an agent setting, implicit file creation can cause unintended local state changes, data leakage through path disclosure, or destructive overwrites if output paths are reused or later made user-controllable.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The natural-language description and instructional content are entirely in Chinese, which can impose a language requirement on users without opt-in. The policy allows locale constraints only when they are explicitly offered as a choice or clearly justified, which is not present here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly writes files to the local filesystem (for example, saving a PNG chart and an XLSX report) without any warning, consent flow, or indication of where files will be created. In an agent setting, silent file creation can overwrite existing files, leak sensitive derived data to persistent storage, or surprise users who expected read-only analysis.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest description is entirely in Chinese and presents the skill's behavior and outputs in that locale without offering any user opt-in or alternative language. Under the stated policy, forcing a specific language or locale is a natural-language policy violation unless the constraint is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill metadata description is written entirely in Chinese, and the rest of the file also uses Chinese headings and output strings. This imposes a specific language on users without any opt-in or justification that the skill is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description at L003 says the skill performs outlier detection and distribution-based quality assessment. Step 5 adds handle_outliers, which can cap values or delete rows, introducing active data-cleaning/modification behavior beyond pure assessment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The description specifies text normalization that extracts only Chinese characters, and the implementation reinforces this by retaining only the Unicode Chinese range. This imposes a language-specific processing policy without indicating user opt-in or clarifying that the skill is intentionally limited to Chinese-language data.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

文档将能力描述为面向 Excel 数据的分类统计、交叉分析与可视化,但代码实际数据入口仅使用 pd.read_excel 读取单个 Excel 文件,未体现更广义表格输入支持。另一方面,代码会将分析结果保存为 PNG 图片,却没有把处理后的结果导出为 Excel/CSV,这与上层工作流强调的数据导出链路存在能力落差。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.