T09 · Insecure Skill Coding Practices
- Location
scripts/governance.py:316- Finding
Uncontrolled Transmission of Warehouse Metadata and Sample Values to External LLM Services
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill has a coherent data-governance purpose, but it can send warehouse samples and metadata to external LLM endpoints and automatically modify Hive comments at scale.
Review before installing. Use only with approved warehouse credentials, restrict the target databases/tables, disable Hive writeback and automatic approval unless you have a review and rollback process, and replace the default external Level2 model endpoint with an approved internal/local model or sanitize prompts so raw samples and DML do not leave your environment.
scripts/governance.py:316Uncontrolled Transmission of Warehouse Metadata and Sample Values to External LLM Services
scripts/governance.py:224Hive DDL Injection Through Unescaped LLM-Generated Comments
The skill instructs collection of table schema, sample values, and view definitions, which can expose sensitive business logic, PII, credentials embedded in SQL, or regulated data. It does so without an explicit warning, scoping guidance, or consent checkpoint, making over-collection and accidental disclosure more likely.
The skill describes capabilities that write files and depend on environment/configuration discovery, but it does not declare an explicit tool scope such as allowed-tools or permissions. That creates a mismatch between documented behavior and enforceable boundaries, increasing the chance the agent uses broader file-writing capability than the user expects.
The trigger phrases are broad and conversational, so the skill could activate during ordinary discussion of data governance rather than an explicit request to run it. In this skill, unintended activation is riskier because activation can lead to schema discovery, sampling data, and possible metadata modification in Hive.
The skill normalizes automatic writeback to Hive COMMENT as part of routine operation, but does not clearly warn that this modifies production metadata and may affect downstream users, governance processes, lineage tools, or NL2SQL behavior. In environments with shared warehouses, incorrect or premature comments can propagate widely and be difficult to audit.
The configuration sends metadata content to an external LLM endpoint at api.deepseek.com. In a data-governance skill, table names, column names, comments, and schema context can contain sensitive business information, so transmitting them to a third-party service creates confidentiality and data-governance risk even if the API key is loaded securely from an environment variable.
level2:
type: openai_compatible
base_url: https://api.deepseek.com/v1
model: deepseek-chat
api_key: ${DEEPSEEK_API_KEY} # 从环境变量读取
concurrency: 5
The auto_approve threshold enables automatic writeback of generated comments into Hive without guaranteed human review. Because the skill performs governance actions on production metadata, incorrect or hallucinated annotations can be mass-written, degrading metadata quality, misleading users and downstream NL2SQL systems, and causing integrity issues at scale.
# model: bge-m3
thresholds:
auto_approve: 85 # >= 85% 自动写回 Hive COMMENT
level2_review: 70 # 70-84% 先 Self-Consistency,不过才升级 Level2
self_consistency_n: 5 # Self-Consistency 采样次数
self_consistency_min_agreement: 0.6 # 5次中至少3次一致才通过
Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.
def _can_import(mod):
try:
__import__(mod)
return True
except ImportError:
return False
The prompt requires output JSON containing a '15字以内中文注释', which mandates Chinese output regardless of user preference. This is a natural-language locale constraint and the file does not provide any opt-in, alternative language selection, or documented regional justification.
The prompt-building path packages raw field samples, sibling fields, DML-derived expressions, and existing comments, then sends them to external LLM endpoints via call_llm. In a data-governance context, those samples can contain sensitive business data, identifiers, or regulated fields, so this exceeds minimal metadata processing and creates a real data-exfiltration/privacy risk.
The code sends schema context and data samples to external model endpoints through the LLM path without any user-facing disclosure or consent in the execution flow. Because this skill is intended for enterprise warehouse governance, the context makes undisclosed outbound transmission materially more dangerous than in a toy/demo setting.
The Level1 and Level2 system messages are written entirely in Chinese and orient the task around generating Chinese metadata annotations. Because this file applies the locale choice globally and automatically, it constitutes a language-policy concern unless users can opt in or configure another locale.
The code auto-approves LLM-generated comments once a confidence threshold is met, allowing model output to advance without human validation. In combination with optional Hive writeback, this creates an autonomous metadata modification pipeline that can propagate inaccurate or unsafe annotations at scale.
reviewed_by = "auto"
review_status = "pending"
if conf >= thresholds["auto_approve"]:
review_status = "approved"
counts["auto"] += 1
Level2 review can also auto-approve based solely on model-reported confidence, further entrenching autonomous decision-making without independent verification. Because these decisions can trigger persistent Hive COMMENT updates, the integrity impact is real even though the changes are 'only metadata.'
r2 = infer_l2(cfg, prompt, r1)
l2_c = r2.get("comment", comment)
l2_conf = float(r2.get("confidence", conf))
if l2_conf >= thresholds["auto_approve"]:
final_comment = l2_c
conf = l2_conf
reviewed_by = "level2"
Approved comments are written back automatically to Hive during scan, causing persistent metadata changes without an explicit confirmation step at execution time. In this workflow, LLM-generated outputs can therefore modify production catalog metadata in bulk, creating integrity and operational risk if comments are wrong, misleading, or policy-noncompliant.
The initialization flow suggests and defaults a remote API endpoint for the Level2 model, normalizing outbound transmission to an external service in an enterprise data-governance tool. By itself the string is not exfiltration, but in context it steers operators toward a configuration that later transmits schema context and possible raw samples off-platform.
user = input("Hive 用户名 [hive]: ").strip() or "hive"
l1_url = input("Level1 模型 API 地址 [http://localhost:8000/v1]: ").strip() or "http://localhost:8000/v1"
l1_mod = input("Level1 模型名称 [qwen2.5-7b]: ").strip() or "qwen2.5-7b"
l2_url = input("Level2 模型 API 地址 [https://api.deepseek.com/v1]: ").strip() or "https://api.deepseek.com/v1"
l2_mod = input("Level2 模型名称 [deepseek-chat]: ").strip() or "deepseek-chat"
with open(tmpl_path) as f:
The config-template replacement path persists the public external model endpoint into the generated configuration, making later external transmission more likely without additional scrutiny. In this skill's enterprise warehouse setting, that increases the chance that sensitive schema and sample data will be sent to third-party services.
content = content.replace("username: hive", f"username: {user}", 1)
content = content.replace("http://localhost:8000/v1", l1_url, 1)
content = content.replace("qwen2.5-7b", l1_mod, 1)
content = content.replace("https://api.deepseek.com/v1", l2_url, 1)
content = content.replace("deepseek-chat", l2_mod, 1)
with open(cfg_path, "w") as f:
Core invocation guidance and examples are presented entirely in Chinese, including trigger phrases that assume Chinese-language user input, with only a minimal English phrase included. This can amount to an implicit language constraint without opt-in or documented justification.
No suspicious patterns detected.