Back to skill

Security audit

warehouse-meta

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent data-governance purpose, but it can send warehouse samples and metadata to external LLM endpoints and automatically modify Hive comments at scale.

Review before installing. Use only with approved warehouse credentials, restrict the target databases/tables, disable Hive writeback and automatic approval unless you have a review and rollback process, and replace the default external Level2 model endpoint with an approved internal/local model or sanitize prompts so raw samples and DML do not leave your environment.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/governance.py:316
Finding

Uncontrolled Transmission of Warehouse Metadata and Sample Values to External LLM Services

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/governance.py:224
Finding

Hive DDL Injection Through Unescaped LLM-Generated Comments

Content
View full analysis
bool: try: conn.cursor().execute( f"ALTER TABLE {db}.{table} CHANGE `{field}` `{field}` COMMENT '{comment}'" ) return True except Exception as e: print(f" ⚠️ 写回失败 {table}.{field}:{e}") return False ``` The model controls both the generated comment and its reported confidence: ```python try: r1 = infer_l1(cfg, prompt) comment = r1.get("comment", "") conf = float(r1.get("confidence", 0)) except Exception: comment, conf = "", 0 final_comment = comment reviewed_by = "auto" review_status = "pending" if conf >= thresholds["auto_approve"]: review_status = "approved" counts["auto"] += 1 ``` Approved output is automatically passed to the DDL function: ```python # 写回 Hive COMMENT(仅 pyhive 模式且已批准) if connector_mode == "pyhive" and review_status == "approved" and hive_conn: if cfg.get("writeback", {}).get("hive_comment", True): pyhive_writeback(hive_conn, table_data["db"], table_data["table"], field_name, final_comment) ``` ### Technical Analysis `pyhive_writeback` constructs a Hive DDL statement through direct string interpolation. The generated comment is placed inside a single-quoted SQL literal without escaping or validation. Database and table identifiers are also interpolated without validation, while field names are enclosed in backticks but are not checked for embedded backticks. The comment originates from an LLM response. That response is untrusted because its prompt includes externally sourced schema metadata, existing comments, DDL, and sampled warehouse values. A malicious value or ...[truncated 2661 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (17)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs collection of table schema, sample values, and view definitions, which can expose sensitive business logic, PII, credentials embedded in SQL, or regulated data. It does so without an explicit warning, scoping guidance, or consent checkpoint, making over-collection and accidental disclosure more likely.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill describes capabilities that write files and depend on environment/configuration discovery, but it does not declare an explicit tool scope such as allowed-tools or permissions. That creates a mismatch between documented behavior and enforceable boundaries, increasing the chance the agent uses broader file-writing capability than the user expects.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad and conversational, so the skill could activate during ordinary discussion of data governance rather than an explicit request to run it. In this skill, unintended activation is riskier because activation can lead to schema discovery, sampling data, and possible metadata modification in Hive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill normalizes automatic writeback to Hive COMMENT as part of routine operation, but does not clearly warn that this modifies production metadata and may affect downstream users, governance processes, lineage tools, or NL2SQL behavior. In environments with shared warehouses, incorrect or premature comments can propagate widely and be difficult to audit.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The configuration sends metadata content to an external LLM endpoint at api.deepseek.com. In a data-governance skill, table names, column names, comments, and schema context can contain sensitive business information, so transmitting them to a third-party service creates confidentiality and data-governance risk even if the API key is loaded securely from an environment variable.

Content

Scanner excerpt · references/config_template.yaml (reported line 22)May include surrounding context.

yaml
level2:
    type: openai_compatible
    base_url: https://api.deepseek.com/v1
    model: deepseek-chat
    api_key: ${DEEPSEEK_API_KEY}   # 从环境变量读取
    concurrency: 5

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

The auto_approve threshold enables automatic writeback of generated comments into Hive without guaranteed human review. Because the skill performs governance actions on production metadata, incorrect or hallucinated annotations can be mass-written, degrading metadata quality, misleading users and downstream NL2SQL systems, and causing integrity issues at scale.

Content

Scanner excerpt · references/config_template.yaml (reported line 37)May include surrounding context.

yaml
# model: bge-m3

thresholds:
  auto_approve: 85        # >= 85% 自动写回 Hive COMMENT
  level2_review: 70       # 70-84% 先 Self-Consistency,不过才升级 Level2
  self_consistency_n: 5   # Self-Consistency 采样次数
  self_consistency_min_agreement: 0.6   # 5次中至少3次一致才通过

Dynamic import via __import__()

Medium
Category
Dangerous Code Execution
Confidence
75% confidence
Finding

Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

Content

Scanner excerpt · scripts/governance.py (reported line 87)May include surrounding context.

python
def _can_import(mod):
    try:
        __import__(mod)
        return True
    except ImportError:
        return False

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The prompt requires output JSON containing a '15字以内中文注释', which mandates Chinese output regardless of user preference. This is a natural-language locale constraint and the file does not provide any opt-in, alternative language selection, or documented regional justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The prompt-building path packages raw field samples, sibling fields, DML-derived expressions, and existing comments, then sends them to external LLM endpoints via call_llm. In a data-governance context, those samples can contain sensitive business data, identifiers, or regulated fields, so this exceeds minimal metadata processing and creates a real data-exfiltration/privacy risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code sends schema context and data samples to external model endpoints through the LLM path without any user-facing disclosure or consent in the execution flow. Because this skill is intended for enterprise warehouse governance, the context makes undisclosed outbound transmission materially more dangerous than in a toy/demo setting.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The Level1 and Level2 system messages are written entirely in Chinese and orient the task around generating Chinese metadata annotations. Because this file applies the locale choice globally and automatically, it constitutes a language-policy concern unless users can opt in or configure another locale.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
92% confidence
Finding

The code auto-approves LLM-generated comments once a confidence threshold is met, allowing model output to advance without human validation. In combination with optional Hive writeback, this creates an autonomous metadata modification pipeline that can propagate inaccurate or unsafe annotations at scale.

Content

Scanner excerpt · scripts/governance.py (reported line 594)May include surrounding context.

python
reviewed_by   = "auto"
        review_status = "pending"

        if conf >= thresholds["auto_approve"]:
            review_status = "approved"
            counts["auto"] += 1

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
92% confidence
Finding

Level2 review can also auto-approve based solely on model-reported confidence, further entrenching autonomous decision-making without independent verification. Because these decisions can trigger persistent Hive COMMENT updates, the integrity impact is real even though the changes are 'only metadata.'

Content

Scanner excerpt · scripts/governance.py (reported line 615)May include surrounding context.

python
r2      = infer_l2(cfg, prompt, r1)
                    l2_c    = r2.get("comment", comment)
                    l2_conf = float(r2.get("confidence", conf))
                    if l2_conf >= thresholds["auto_approve"]:
                        final_comment = l2_c
                        conf          = l2_conf
                        reviewed_by   = "level2"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Approved comments are written back automatically to Hive during scan, causing persistent metadata changes without an explicit confirmation step at execution time. In this workflow, LLM-generated outputs can therefore modify production catalog metadata in bulk, creating integrity and operational risk if comments are wrong, misleading, or policy-noncompliant.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The initialization flow suggests and defaults a remote API endpoint for the Level2 model, normalizing outbound transmission to an external service in an enterprise data-governance tool. By itself the string is not exfiltration, but in context it steers operators toward a configuration that later transmits schema context and possible raw samples off-platform.

Content

Scanner excerpt · scripts/governance.py (reported line 992)May include surrounding context.

python
user   = input("Hive 用户名 [hive]: ").strip() or "hive"
        l1_url = input("Level1 模型 API 地址 [http://localhost:8000/v1]: ").strip() or "http://localhost:8000/v1"
        l1_mod = input("Level1 模型名称 [qwen2.5-7b]: ").strip() or "qwen2.5-7b"
        l2_url = input("Level2 模型 API 地址 [https://api.deepseek.com/v1]: ").strip() or "https://api.deepseek.com/v1"
        l2_mod = input("Level2 模型名称 [deepseek-chat]: ").strip() or "deepseek-chat"

        with open(tmpl_path) as f:

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The config-template replacement path persists the public external model endpoint into the generated configuration, making later external transmission more likely without additional scrutiny. In this skill's enterprise warehouse setting, that increases the chance that sensitive schema and sample data will be sent to third-party services.

Content

Scanner excerpt · scripts/governance.py (reported line 1002)May include surrounding context.

python
content = content.replace("username: hive", f"username: {user}", 1)
        content = content.replace("http://localhost:8000/v1", l1_url, 1)
        content = content.replace("qwen2.5-7b", l1_mod, 1)
        content = content.replace("https://api.deepseek.com/v1", l2_url, 1)
        content = content.replace("deepseek-chat", l2_mod, 1)

        with open(cfg_path, "w") as f:

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

Core invocation guidance and examples are presented entirely in Chinese, including trigger phrases that assume Chinese-language user input, with only a minimal English phrase included. This can amount to an implicit language constraint without opt-in or documented justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.