Back to skill

Security audit

Knowledge Graph - Etl Pipeline Generator

Security checks for vulnerabilities and agentic risk

Overview

This ETL skill is mostly purpose-aligned, but it includes unsafe code-generation and under-scoped guidance for sensitive data, credentials, logs, and production graph writes.

Review before installing or using this skill in real workflows. Do not execute generated Python from untrusted pipeline metadata without sanitizing or reviewing it. Use environment variables or a secret manager for credentials, prefer TLS endpoints, test loads in staging first, and redact or minimize any sensitive records written to logs, dead-letter queues, or external enrichment APIs.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/etl_generator.py:426
Finding

Generated Python Script Injection Through Unsanitized Pipeline Metadata

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
examples/example-pipelines.md:196
Finding

Hardcoded Database Password Pattern in Runnable Example

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
examples/example-pipelines.md:295
Finding

Plaintext Transport Recommended for Remote Graph Database Connections

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
references/pipeline-patterns.md:434
Finding

Sensitive Records and Connection Information May Be Written to Logs and Dead-Letter Queues

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
references/pipeline-patterns.md:218
Finding

Customer Identifiers Disclosed Through External API URL Paths

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (15)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill describes capabilities that involve reading local files and external data sources, but it does not declare any explicit tool scope or permission boundaries. In an agent environment, this can lead to over-broad file access assumptions and unsafe execution contexts where the skill may be used against unintended files or sensitive datasets.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The extract-stage guidance includes authentication and external source ingestion but does not warn users about secret handling, untrusted data, or validation of external endpoints. This increases the chance that users will embed credentials insecurely or ingest hostile data into downstream transformation and execution steps.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The load-stage section discusses bulk import, streaming, and batch loading into graph systems without warning about overwrite behavior, duplicate creation, schema corruption, or transactional side effects. In ETL contexts, load operations can materially alter production graph data, so missing safety guidance increases the risk of integrity loss or destructive writes.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · examples/example-pipelines.md (reported line 24)May include surrounding context.

md
location: data/customers.csv
      - name: orders_api
        type: api
        endpoint: https://api.ecommerce.com/orders
        auth: bearer_token
      - name: products_db
        type: database

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · examples/example-pipelines.md (reported line 118)May include surrounding context.

md
location: data/customers.csv
      - name: orders_api
        type: api
        endpoint: https://api.ecommerce.com/orders
        auth: bearer_token
      - name: products_db
        type: database

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file includes a healthcare pipeline that integrates patient, provider, treatment, and claims data from EHR and insurance systems. Although the example mentions anonymization and HIPAA-related processing, it does not provide any user-facing warning that working with this example involves highly sensitive health data and requires appropriate authorization, secure handling, and compliance controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guide recommends storing failed records in logs and dead-letter queues, including the original record and error details, but does not warn that these artifacts may contain PII, credentials, or other sensitive business data. In an ETL skill focused on graph/knowledge-graph ingestion, failed records commonly include customer and identity data, so insecure DLQ/logging guidance can lead to secondary data exposure through files, retention, or operator access.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module documentation says it provides functionality for designing and executing ETL pipelines, implying real end-to-end execution. In the implementation, _extract_api only logs a placeholder message and returns no data, while _execute_load merely logs and returns len(data) without loading into any graph database or knowledge graph, so the documented intent overstates what the code actually does.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/pipeline-patterns.md (reported line 38)May include surrounding context.

md
description="Fetch data from API and convert to RDF"
    )

    pipeline2.add_extract("api", "https://api.example.com/data")
    pipeline2.add_transform(
        operations=[
            "parse_json",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/pipeline-patterns.md (reported line 222)May include surrounding context.

md
description="Fetch data from API and convert to RDF"
    )

    pipeline2.add_extract("api", "https://api.example.com/data")
    pipeline2.add_transform(
        operations=[
            "parse_json",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/etl_generator.py (reported line 551)May include surrounding context.

python
description="Fetch data from API and convert to RDF"
    )

    pipeline2.add_extract("api", "https://api.example.com/data")
    pipeline2.add_transform(
        operations=[
            "parse_json",

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The markdown specifies the NLP pipeline with ner_model: en_core_web_lg, which imposes an English-language processing assumption. Under the policy, forcing a specific language without user opt-in or clear justification is a natural-language locale constraint that should be disclosed or made configurable.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The YAML generation emits source locations and other pipeline configuration details directly into output. In ETL contexts, these fields often contain sensitive file paths, internal endpoints, database names, or even embedded credentials in connection strings, so printing or exporting them without redaction can leak infrastructure details.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

to_python_script emits docstrings such as Extract stage - load data from source and Load stage - load to target system, which describe completed behavior. However, the generated functions only initialize empty data, add TODO comments, and count records, so the embedded documentation contradicts the generated script's actual no-op behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The example code prints generated YAML and script content, which may disclose local file paths, API endpoints, target database URIs, and other environment-specific details if adapted from real deployments. In a skill meant to automate ETL pipelines, this context makes accidental disclosure more plausible because operators commonly work with sensitive infrastructure metadata.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.