Back to skill

Security audit

File Splitter

Security checks for vulnerabilities and agentic risk

Overview

This skill coherently splits user-selected JSON, Markdown, and text files, with the main caution that output files can overwrite predictable existing paths if the output folder is unsafe.

Install only if you intend to let the agent split local JSON, Markdown, or text files. Use a new or trusted output folder, prefer dry-run first, and avoid shared or attacker-writable output directories because existing predictable output names may be replaced.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/split_files.py:116
Finding

Predictable Output Files Permit Overwrite and Symbolic-Link Attacks

Content
View full analysis

Vulnerability Details

File Location: scripts/split_files.py, lines 116–123 and 186–194
Vulnerability Type: Unsafe file creation and symbolic-link following
Risk Level: Medium

Vulnerable Code

JSON output path:

python
seq = str(idx + 1).zfill(seq_digits)
out_name = f"{base}{seq}{ext}"
out_path = os.path.join(output_folder, out_name)

if dry_run:
    print(f"  [DRY] {out_name} ({len(chunk)} 条)")
else:
    with open(out_path, "w", encoding="utf-8") as f:
        json.dump(chunk, f, ensure_ascii=False, indent=2)

Markdown and text output path:

python
seq = str(idx + 1).zfill(seq_digits)
out_name = f"{base}{seq}{ext}"
out_path = os.path.join(output_folder, out_name)

chunk_bytes = len(chunk.encode("utf-8"))
if dry_run:
    print(f"  [DRY] {out_name} ({fmt_size(chunk_bytes)})")
else:
    with open(out_path, "w", encoding="utf-8") as f:
        f.write(chunk)

Technical Analysis

The output filenames are deterministically derived from the source filename and a sequential number. The script opens each destination using Python's "w" mode without first ensuring that the path does not exist and is not a symbolic link.

The "w" mode silently truncates an existing regular file. It also follows symbolic links under normal filesystem semantics. Consequently, a party able to create entries in the selected output directory can predict the generated filename and place a symbolic link there before execution. When the script opens that path, it writes generated chunk content to the symlink target.

The destination path is joined to the requested output folder, but no canonical-path or file-type validation is performed immediately before creation. There is also no use of exclusive creation, no atomic no-follow operation, and no explicit overwrite confirmation.

Attack Path

  1. The attacker obtains write access to an output directory that a victim will use with the skill.
  2. The attacker learns or predicts the source basename and for ...[truncated 1278 chars]
Remediation
View remediation

Remediation Suggestions

  1. Refuse to overwrite existing destinations by default. Use exclusive creation:
python
with open(out_path, "x", encoding="utf-8") as f:
    json.dump(chunk, f, ensure_ascii=False, indent=2)
  1. Add an explicit --overwrite option if replacement is required. Keep safe, non-overwriting behavior as the default and clearly warn users before replacement.

  2. Reject symbolic links and other unexpected file types. Where supported, open files using low-level flags such as O_CREAT | O_EXCL | O_WRONLY | O_NOFOLLOW, then wrap the resulting descriptor with os.fdopen.

  3. Resolve and validate the destination path against a canonical output directory before writing. This should supplement, not replace, no-follow and exclusive-creation controls because path checks alone can be vulnerable to time-of-check/time-of-use races.

  4. For intended replacement operations, write to a securely and exclusively created temporary file in the same directory, flush and synchronize it as appropriate, and then atomically replace the final destination only after validating overwrite policy.

  5. Ensure the output directory has restrictive permissions and is not writable by untrusted users. Document that shared or attacker-controlled output directories are unsafe without these protections.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documents and invokes a script that reads from an input folder and writes split output files, but it does not declare any explicit tool scope such as permissions or allowed-tools. That mismatch can lead to overbroad or implicit file-system access at runtime, making it harder to enforce least privilege and increasing the chance of unintended file reads or writes if the skill is invoked in sensitive contexts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger list includes generic terms like "chunk" and "segment," which are common in unrelated conversations and can cause accidental or overly broad skill invocation. Because this skill performs file read/write operations, unintended activation could expose local data to unnecessary processing or create files in contexts where the user did not explicitly request file manipulation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This file is a code file, so SQP-3 applies to natural-language strings embedded in docstrings and help text. The script presents its usage instructions exclusively in Chinese and does not provide any opt-in, fallback, or justification for a locale-specific constraint, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.