T09 · Insecure Skill Coding Practices
- Location
scripts/etl_generator.py:426- Finding
Generated Python Script Injection Through Unsanitized Pipeline Metadata
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This ETL skill is mostly purpose-aligned, but it includes unsafe code-generation and under-scoped guidance for sensitive data, credentials, logs, and production graph writes.
Review before installing or using this skill in real workflows. Do not execute generated Python from untrusted pipeline metadata without sanitizing or reviewing it. Use environment variables or a secret manager for credentials, prefer TLS endpoints, test loads in staging first, and redact or minimize any sensitive records written to logs, dead-letter queues, or external enrichment APIs.
scripts/etl_generator.py:426Generated Python Script Injection Through Unsanitized Pipeline Metadata
examples/example-pipelines.md:196Hardcoded Database Password Pattern in Runnable Example
examples/example-pipelines.md:295Plaintext Transport Recommended for Remote Graph Database Connections
references/pipeline-patterns.md:434Sensitive Records and Connection Information May Be Written to Logs and Dead-Letter Queues
references/pipeline-patterns.md:218Customer Identifiers Disclosed Through External API URL Paths
The skill describes capabilities that involve reading local files and external data sources, but it does not declare any explicit tool scope or permission boundaries. In an agent environment, this can lead to over-broad file access assumptions and unsafe execution contexts where the skill may be used against unintended files or sensitive datasets.
The extract-stage guidance includes authentication and external source ingestion but does not warn users about secret handling, untrusted data, or validation of external endpoints. This increases the chance that users will embed credentials insecurely or ingest hostile data into downstream transformation and execution steps.
The load-stage section discusses bulk import, streaming, and batch loading into graph systems without warning about overwrite behavior, duplicate creation, schema corruption, or transactional side effects. In ETL contexts, load operations can materially alter production graph data, so missing safety guidance increases the risk of integrity loss or destructive writes.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
location: data/customers.csv
- name: orders_api
type: api
endpoint: https://api.ecommerce.com/orders
auth: bearer_token
- name: products_db
type: database
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
location: data/customers.csv
- name: orders_api
type: api
endpoint: https://api.ecommerce.com/orders
auth: bearer_token
- name: products_db
type: database
This markdown file includes a healthcare pipeline that integrates patient, provider, treatment, and claims data from EHR and insurance systems. Although the example mentions anonymization and HIPAA-related processing, it does not provide any user-facing warning that working with this example involves highly sensitive health data and requires appropriate authorization, secure handling, and compliance controls.
The guide recommends storing failed records in logs and dead-letter queues, including the original record and error details, but does not warn that these artifacts may contain PII, credentials, or other sensitive business data. In an ETL skill focused on graph/knowledge-graph ingestion, failed records commonly include customer and identity data, so insecure DLQ/logging guidance can lead to secondary data exposure through files, retention, or operator access.
The module documentation says it provides functionality for designing and executing ETL pipelines, implying real end-to-end execution. In the implementation, _extract_api only logs a placeholder message and returns no data, while _execute_load merely logs and returns len(data) without loading into any graph database or knowledge graph, so the documented intent overstates what the code actually does.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
description="Fetch data from API and convert to RDF"
)
pipeline2.add_extract("api", "https://api.example.com/data")
pipeline2.add_transform(
operations=[
"parse_json",
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
description="Fetch data from API and convert to RDF"
)
pipeline2.add_extract("api", "https://api.example.com/data")
pipeline2.add_transform(
operations=[
"parse_json",
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
description="Fetch data from API and convert to RDF"
)
pipeline2.add_extract("api", "https://api.example.com/data")
pipeline2.add_transform(
operations=[
"parse_json",
The markdown specifies the NLP pipeline with ner_model: en_core_web_lg, which imposes an English-language processing assumption. Under the policy, forcing a specific language without user opt-in or clear justification is a natural-language locale constraint that should be disclosed or made configurable.
The YAML generation emits source locations and other pipeline configuration details directly into output. In ETL contexts, these fields often contain sensitive file paths, internal endpoints, database names, or even embedded credentials in connection strings, so printing or exporting them without redaction can leak infrastructure details.
to_python_script emits docstrings such as Extract stage - load data from source and Load stage - load to target system, which describe completed behavior. However, the generated functions only initialize empty data, add TODO comments, and count records, so the embedded documentation contradicts the generated script's actual no-op behavior.
The example code prints generated YAML and script content, which may disclose local file paths, API endpoints, target database URIs, and other environment-specific details if adapted from real deployments. In a skill meant to automate ETL pipelines, this context makes accidental disclosure more plausible because operators commonly work with sensitive infrastructure metadata.
No suspicious patterns detected.