Back to skill

Security audit

Clean CSV Toolkit

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent local CSV toolkit, but several data-writing commands can accidentally erase source files when the output path points to an input file.

Review before installing. Use this only on approved local datasets, avoid printing or saving previews/reports that contain PII or secrets, and never point an output path at the same file as any input; keep backups or write to a separate directory until the aliasing issue is fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/filter.py:309
Finding

Destructive input truncation through output-path aliasing in filter.py

Content
View full analysis

Vulnerability Details

File Location: scripts/filter.py:309-322
Vulnerability Type: Input/output path aliasing and unsafe non-atomic file replacement
Risk Level: High

Vulnerable Code

python
with open(out_path, "w", encoding="utf-8", newline="") as fout:
    if fmt == "jsonl":
        writer = fout
        def emit(row):
            fout.write(json.dumps(row, ensure_ascii=False) + "\n")
    else:
        delim = "\t" if fmt == "tsv" else ","
        csvw = csv.DictWriter(fout, fieldnames=out_cols, delimiter=delim,
                              extrasaction="ignore")
        csvw.writeheader()
        def emit(row):
            csvw.writerow(row)

    with open_table(in_path) as (_kind, _hdr, reader):

Technical Analysis

The output file is opened with mode "w" before the processing pass opens the input. Opening a file in this mode immediately truncates it. The script does not verify that out_path and in_path refer to different filesystem objects.

A textual path comparison alone would also be insufficient because the same file can be referenced through relative-path normalization, symbolic links, or hard links. Although safe_path() restricts path characters, it does not enforce input/output separation and therefore does not prevent this condition.

Attack Path

  1. The caller supplies an existing CSV file as both input and output:
    bash
    python3 scripts/filter.py data.csv data.csv --where "id > 0"
    
  2. The script initially verifies that data.csv is a file and reads its header.
  3. It opens out_path with "w", immediately truncating data.csv.
  4. It then reopens in_path for the processing pass.
  5. The input now contains no original records, so the source data is permanently replaced by an empty or header-only result.

The same result can be induced using a symbolic-link or hard-link output that aliases the input.

Impact Assessment

...[truncated 395 chars]

Remediation
View remediation

Remediation Suggestions

  1. Resolve input and output paths before processing and reject direct equality.
  2. If the output already exists, use os.path.samefile(in_path, out_path) to detect symbolic-link and hard-link aliases.
  3. Write output to a securely created temporary file in the destination directory.
  4. Flush and close the temporary file successfully before replacing the destination with os.replace().
  5. On failure, delete the temporary file and preserve the original input.
  6. Add regression tests covering identical paths, normalized relative paths, symbolic links, hard links, and failures during row processing.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/transform.py:533
Finding

Destructive input truncation through output-path aliasing in transform.py

Content
View full analysis

Vulnerability Details

File Location: scripts/transform.py:533-547
Vulnerability Type: Input/output path aliasing and unsafe non-atomic file replacement
Risk Level: High

Vulnerable Code

python
with open(out_path, "w", encoding="utf-8", newline="") as fout:
    if fmt == "jsonl":
        def emit(row, header=None):
            fout.write(json.dumps(
                {k: ("" if v is None else v) for k, v in row.items()},
                ensure_ascii=False) + "\n")
    else:
        delim = "\t" if fmt == "tsv" else ","
        writer = csv.DictWriter(fout, fieldnames=out_header, delimiter=delim,
                                extrasaction="ignore")
        writer.writeheader()
        def emit(row, header=None):
            writer.writerow({k: ("" if v is None else v) for k, v in row.items()})

    with open_table(in_path) as (_kind, header, reader):

Technical Analysis

The transformation command opens and truncates its output before reopening the input for the row-processing pass. There is no check that the two paths identify distinct filesystem objects.

The earlier schema-discovery pass does not protect the source file: after that pass closes, opening an aliased output with "w" erases the original content. The subsequent processing pass therefore reads the truncated file.

Attack Path

  1. Invoke the transformation with an output that aliases its input:
    bash
    python3 scripts/transform.py data.csv data.csv --add "total = price + tax"
    
  2. The script reads the input header to compute the output schema.
  3. It opens the output in truncating mode.
  4. Because the output and input are the same file, all original records are erased.
  5. The script reopens the now-truncated input and emits no transformed records.

Equivalent exploitation is possible using a symbolic link or hard link as the output path.

Impact Assessment

The vulnerability does n ...[truncated 293 chars]

Remediation
View remediation

Remediation Suggestions

Reject outputs that refer to the input using both canonical-path comparison and os.path.samefile() where applicable. Implement transformations through a temporary file created in the destination directory, close and validate it, and then atomically publish it with os.replace(). Preserve the temporary file only when explicitly requested for diagnostics, and ensure cleanup occurs after exceptions. Add tests for direct, symbolic-link, and hard-link aliases.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/concat.py:117
Finding

Destructive source truncation through output-path aliasing in concat.py

Content
View full analysis

Vulnerability Details

File Location: scripts/concat.py:117-131
Vulnerability Type: Output aliasing with one or more input files
Risk Level: High

Vulnerable Code

python
with open(out_path, "w", encoding="utf-8", newline="") as fout:
    if fmt == "jsonl":
        def emit(row):
            fout.write(json.dumps(row, ensure_ascii=False) + "\n")
    else:
        delim = "\t" if fmt == "tsv" else ","
        csvw = csv.DictWriter(fout, fieldnames=union_header, delimiter=delim,
                              extrasaction="ignore")
        csvw.writeheader()
        def emit(row):
            csvw.writerow(row)

    for path in in_paths:
        kept_from_file = 0
        with open_table(path) as (_kind, _hdr, reader):

Technical Analysis

concat.py performs a first pass to collect headers, then opens the output with "w" before beginning its second input pass. It does not reject an output that is identical to, or aliases, any path in in_paths.

Consequently, if the output aliases one of the source files, that source is truncated before the second pass reads it. Depending on input ordering and buffering, the result can omit the aliased source, replace it with a partial concatenation, or produce other corrupted output.

Attack Path

  1. Select one existing input file as the output:
    bash
    python3 scripts/concat.py january.csv january.csv february.csv
    
  2. The first pass reads headers from both inputs successfully.
  3. Opening january.csv as the output truncates it.
  4. The second pass attempts to read january.csv, but its records have already been erased.
  5. The command writes only remaining input data and replaces the original January dataset.

The output may also be a symbolic link or hard link to any input rather than using the same path string.

Impact Assessment

No additional system privileges are obtained. The attacker or mistaken cal ...[truncated 237 chars]

Remediation
View remediation

Remediation Suggestions

Before opening the output, compare it against every input:

  • Compare canonicalized paths.
  • Use os.path.samefile() for existing files to identify hard-link and symbolic-link aliases.
  • Reject the operation if any alias is detected, or explicitly support in-place behavior through a safe temporary-file workflow.
  • Write the complete concatenation to a temporary file in the output directory and publish it atomically only after every input has been read successfully.
  • Add tests where the output aliases the first, middle, and last input through direct paths, symbolic links, and hard links.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/merge.py:183
Finding

Destructive left-input truncation through output-path aliasing in merge.py

Content
View full analysis

Vulnerability Details

File Location: scripts/merge.py:183-192
Vulnerability Type: Join output aliasing a streamed input
Risk Level: High

Vulnerable Code

python
with open(out_path, "w", encoding="utf-8", newline="") as fout:
    if fmt == "jsonl":
        writer = fout
    else:
        delim = "\t" if fmt == "tsv" else ","
        writer = csv.DictWriter(fout, fieldnames=out_header, delimiter=delim,
                                extrasaction="ignore")
        writer.writeheader()

    with open_table(left_path) as (_kind, _hdr, left_reader):

Technical Analysis

The right input is indexed before output creation, but the left input is streamed only after the output is opened in truncating mode. If out_path aliases left_path, the left dataset is erased before the join loop reads it.

No canonical identity or samefile check is performed. An alias therefore does not need to use an identical path string. Although aliasing the right input is less directly destructive during processing because it is already indexed in memory, it can still overwrite that source and should also be rejected.

Attack Path

  1. Invoke a merge with the left input as the output:
    bash
    python3 scripts/merge.py users.csv orders.csv users.csv --on user_id --how left
    
  2. The script indexes the right file and reads the left header.
  3. It opens users.csv with "w", erasing its original contents.
  4. It then opens users.csv as the left-side stream.
  5. No original left rows remain, resulting in an empty or invalid join output while permanently destroying the source file.

A symbolic-link or hard-link output can trigger the same condition.

Impact Assessment

The issue does not provide privilege escalation and is confined to the executing user's filesystem permissions. It can, however, erase either source dataset and cause complete loss of the left-side records. It can also pr ...[truncated 73 chars]

Remediation
View remediation

Remediation Suggestions

Compare the output against both join inputs before any output file is opened. Use resolved-path comparison plus os.path.samefile() to cover symbolic links and hard links. Generate the join in a temporary file located in the destination directory, handle all read and write failures without touching existing destination data, and atomically replace the destination only after successful completion. Add regression tests for output aliases of both left and right inputs.

Vulnerability Patterns
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (9)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 29)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 310)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 314)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 323)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 327)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 331)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 376)May include surrounding context.

md
- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly supports previewing, profiling, diffing, deduplication reports, and inline markdown output of row data, all of which can surface raw records, sample values, emails, and other sensitive fields. Without a clear warning, users or downstream agents may unintentionally expose PII/secrets in console logs, chat responses, CI output, or saved reports during normal use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code performs safety-relevant filesystem changes by creating directories and writing a new output file. While file transformation is the tool's purpose, there is no inline confirmation, cautionary comment, or user-facing disclosure near the write path indicating that existing filesystem state may be modified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.