Back to skill

Security audit

ai-data-delivery

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly disclosed local data-delivery toolkit, but its deployment helper can stop and start services and contains a shell fallback that contradicts its shell-free safety claim.

Review before installing in sensitive environments. Use it in a project sandbox, only point endpoint-based tools at services you are authorized to test, and do not run rebuild.py unless you are comfortable with it stopping a service by port and starting a detached process. Redact sensitive prompts, SQL, reasoning, and reports before sharing outputs. The shell fallback and full environment inheritance should be fixed or operationally controlled.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (80)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Calling remote QA engines over HTTP and producing gate-like regression results extends the skill from static semantic work into active networked evaluation and decision support. This is risky because prompts, model metadata, or business-sensitive test cases may be transmitted externally, and the outputs may be treated as authoritative despite being generated by an undeclared remote interaction path.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · examples/onboarding/run_demo.py (reported line 22)May include surrounding context.

python
def run(script, *args):
    command = [sys.executable, str(SCRIPTS / script), *map(str, args)]
    result = subprocess.run(command, capture_output=True, text=True, encoding="utf-8",
                            env=dict(os.environ, PYTHONIOENCODING="utf-8"))
    if result.returncode:
        raise RuntimeError(f"{script}: {result.stdout}\n{result.stderr}")
    print("PASS " + script)

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file contains an explicit security assurance that commands are executed with shell-free argument arrays, but the actual code later uses sh -c. This mismatch is dangerous because it can cause reviewers and operators to underestimate risk and deploy the tool in sensitive CI/CD or service-management workflows, where even a small shell-execution foothold can have outsized consequences.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
75% confidence
Finding

The child process inherits the full parent environment via dict(os.environ, PYTHONIOENCODING="utf-8"). In CI/CD or release-gate contexts, environment variables often contain secrets, tokens, proxy settings, or Python startup controls; passing them wholesale to subsidiary scripts expands the trust boundary and can expose sensitive values to less-trusted code or alter execution through inherited variables.

Content

Scanner excerpt · scripts/release_gate.py (reported line 21)May include surrounding context.

python
def run_gate(name, command):
    result = subprocess.run([sys.executable, str(HERE / command[0]), *command[1:]],
                            capture_output=True, text=True, encoding="utf-8", errors="replace",
                            env=dict(os.environ, PYTHONIOENCODING="utf-8"))
    return {"gate": name, "status": "pass" if result.returncode == 0 else "fail",
            "detail": "\n".join((result.stdout + result.stderr).splitlines()[-5:])}

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises capabilities and workflows that imply file, shell, environment, and possible network access, but it does not declare any explicit tool scope or permissions boundary. That creates an authorization ambiguity where an agent runtime may permit broader actions than the user expects, increasing the chance of unintended file modification, command execution, or data exposure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The primary description is written entirely in Chinese and presents the skill's behavior and scope in that language, with no indication that other languages are supported or that the user may choose their preferred locale. Under the policy, language constraints should be optional or clearly justified; this file does not provide such an opt-in or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Multiple values and labels in this manifest are written in Chinese, which imposes a specific language on users or downstream consumers. The file does not indicate that the skill is China-specific, nor does it offer any language/locale opt-in or alternative.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · examples/onboarding/run_demo.py (reported line 21)May include surrounding context.

python
def run(script, *args):
    command = [sys.executable, str(SCRIPTS / script), *map(str, args)]
    result = subprocess.run(command, capture_output=True, text=True, encoding="utf-8",
                            env=dict(os.environ, PYTHONIOENCODING="utf-8"))
    if result.returncode:
        raise RuntimeError(f"{script}: {result.stdout}\n{result.stderr}")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The entire skill description is written in Chinese and does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-speaking or region-specific audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This Python file contains natural-language instructions and descriptions entirely in Chinese, and the script also emits Chinese-only output. Under the language/locale policy rule, forcing a specific language without user opt-in is a policy violation unless the locale restriction is explicitly justified, which is not stated here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

SQP-3 applies to all file types, including JSON. The repeated natural-language prompts and labels are exclusively in Chinese, but the file does not document that the skill is China/Chinese-specific or that users can opt into the locale, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code includes natural-language instructions and descriptions entirely in Chinese, which imposes a specific language on users without any opt-in or documented justification. The stated policy flags language or locale constraints unless the skill offers a choice or clearly documents a justified regional scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This manifest uses Chinese-only natural-language values for the domain, field labels, instructions, concepts, metrics, and ontology text. Under the policy, forcing a specific language without user opt-in or a documented region-specific justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring is written entirely in Chinese and does not offer any language choice or indicate that the locale restriction is intentional and justified. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This JSON file encodes all user-facing question examples and labels exclusively in Chinese, indicating a fixed language/locale assumption. The policy requires flagging language constraints when the skill does not offer user choice or document a justified region-specific limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest sets the domain and downstream semantic definitions in Chinese, and the AI instructions are also written only in Chinese. There is no natural-language indication that the locale is optional, user-selected, or justified as a region-specific tool, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The module docstring presents the skill entirely in Chinese and frames the dataset semantics using Chinese labels, with no indication that language choice is optional. Under the policy, locale-specific behavior should be opt-in or clearly justified as region-specific; this file does not state such a constraint explicitly.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
mocks/run_mock.py:42