Back to skill

Security audit

alibabacloud-dataphin-skills

Security checks for vulnerabilities and agentic risk

Overview

This Dataphin automation suite matches its stated purpose, but it needs Review because it combines broad administrative access with unsafe credential, installer, and network-security guidance.

Before installing, treat this as a powerful Dataphin administration suite. Use the narrow per-sub-skill RAM policy instead of the suite-wide Resource "*" policy, configure credentials outside the agent conversation, avoid commands that expose secrets as arguments, prefer HTTPS with valid certificates or a trusted private CA, and do not run remote installer scripts directly without verification.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T03 · Remote Payload Retrieval and Execution

Error
Location
references/cli-installation-guide.md:19
Finding

Unverified Remote Installer Scripts Are Executed Directly by Bash

Content
View full analysis
Remediation
View remediation
" curl --fail --show-error --location \ --output "$ARCHIVE" \ "https://trusted.example/releases/${VERSION}/${ARCHIVE}" printf '%s %s\n' "$EXPECTED_SHA256" "$ARCHIVE" | sha256sum --check - tar -xzf "$ARCHIVE" install -m 0755 aliyun "$HOME/.local/bin/aliyun" ``` ]]>

T09 · Insecure Skill Coding Practices

Error
Location
references/dataservice/call-data-service-api/scripts/call-data-service-api.py:125
Finding

Credential-Bearing Dataphin API Requests Use Plaintext HTTP by Default

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
references/assets/manage-asset-attributes/scripts/dataphin_asset_api.py:110
Finding

TLS Certificate Verification Is Disabled by Default for Standalone API Operations

Content
View full analysis
``` ### Technical Analysis The Python client sets `verify_ssl=False` by default, disables the corresponding wa ...[truncated 2110 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:135
Finding

Long-Lived Cloud Secrets Are Solicited in Agent Conversations and Placed in Command Arguments

Content
View full analysis
\ --access-key-secret \ --region ``` For standalone deployments, it similarly constructs: ```bash aliyun configure set --profile dataphin-standalone \ --mode AK \ --access-key-id \ --access-key-secret \ --endpoint ``` ### Technical Analysis The Skill explicitly makes collection of the AccessKey ID and AccessKey secret through the Agent conversation part of its primary credential workflow. It then interpolates the secret into command-line arguments. This creates multiple unnecessary exposure surfaces: - Conversation history and Agent telemetry may retain the secret. - Tool-call logs may record the complete command. - Shell history may preserve manually copied commands. - Process inspection may expose command-line arguments to other processes or users permitted to inspect them. - Error reporting or debugging output may inadvertently reproduce the command. The project does mention configuring credentials independently as an optional alternative, but secure out-of-band configuration should be mandatory rather than optional. Long-lived cloud secrets are not required in the conversational context for the declared Dataphin functionality. ### Attack Path 1. The Agent follows the Skill and asks the user to provide an AccessKey ID and secret. 2. The user enters the long-lived secret into the conversation. 3. The Agent builds an `aliyun configure set` command containing the secret as an argument. 4. The secret may be retained in conversation record ...[truncated 1242 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (563)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description says this skill is a Dataphin suite入口 that routes user scenarios to many concrete sub-skills spanning nearly the full Dataphin product surface. The supplied code does not do routing, scenario detection, or orchestration. Instead, it is a standalone implementation for one narrow function: calling already-published Dataphin Data Service APIs through the gateway, including signature generation, POST requests, async job status/result polling, closeJob, and SSE stream handling. This is a materially different primary purpose from a suite router. While '数据服务 API 调用' appears within the broad declared coverage, the code chunk itself is not an entry skill for all those scenarios and exposes a specific undeclared operational capability profile centered on authenticated gateway invocation. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description says this skill is a Dataphin suite entry point that routes many different business scenarios to specialized sub-skills across a large set of domains. The supplied code chunk does something materially different: it is a test file for a particular data-service API client script. Its primary purpose is regression testing of request signatures, paths, transport handling, async pagination/job closing, server-sent events, and CLI behavior. It even enforces offline execution by blocking socket access. While 'data service API' is one topic mentioned in the description, this code is not implementing a suite router or the broad multi-domain business workflow coverage claimed. Instead, it is narrowly focused on validating a client implementation. That is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents this skill as a broad Dataphin suite entry point that routes requests into many business-domain sub-skills covering end-to-end Dataphin workflows. The supplied code does none of that. It is a standalone support script for generating a skill User-Agent string from a local manifest and session ID, with validation logic for manifest name, semantic version, and session ID format. There is no evidence of routing, keyword handling, Dataphin API access, cloud/resource management, or any of the listed business capabilities. This is a clear description-behavior mismatch because the actual primary purpose is unrelated utility initialization rather than Dataphin workflow orchestration.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents this skill as a broad Dataphin operational router covering many platform workflows. The supplied code does not implement any Dataphin routing, API calls, business-flow handling, or user-triggered operations. Instead, it is a regression test file for a validator module, focused on checking structure and safety constraints of eval definitions and reference documents. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents an orchestration/router skill for a broad Dataphin business suite. The supplied code instead is a repository utility script for offline evaluation file validation. Its behavior is limited to reading local JSONC/reference files, checking naming/assertion conventions, validating prompt contents for offline eval cases, and reporting static coverage. There is no evidence of Dataphin API interaction, request routing to sub-skills, business keyword dispatch, cloud/platform resource operations, or any of the listed domain capabilities. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The standalone flow mandates --skip-secure-verify, which disables TLS certificate validation and enables man-in-the-middle interception of credentials, API requests, and responses. This is especially dangerous here because the same workflow handles AccessKey secrets, tenant identifiers, and administrative Dataphin operations over the network.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 211)May include surrounding context.

md
数据文件:[`references/config/openapi-2.0-versions.json`](./references/config/openapi-2.0-versions.json)(383 条 union,字段 `min_version` + `channel`)。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 254)May include surrounding context.

md
数据文件:[`references/config/openapi-2.0-versions.json`](./references/config/openapi-2.0-versions.json)(383 条 union,字段 `min_version` + `channel`)。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 256)May include surrounding context.

md
数据文件:[`references/config/openapi-2.0-versions.json`](./references/config/openapi-2.0-versions.json)(383 条 union,字段 `min_version` + `channel`)。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 392)May include surrounding context.

md
| 即席查询 / 执行 SQL / 临时跑代码 / 建表 / execute-ad-hoc-task | [execute-ad-hoc-task](./references/dev/execute-ad-hoc-task/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 397)May include surrounding context.

md
| 业务日期 / bizdate / T-1 / 今天日期 / 昨天 | [get-bizdate](./references/dev/get-bizdate/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 403)May include surrounding context.

md
| 数据同步 / 数据搬运 / pipeline / reader-writer / MySQL→MaxCompute | [create-pipeline-task](./references/pipeline/create-pipeline-task/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 427)May include surrounding context.

md
| 知识图谱 Schema / 本体模型 / 实体类型 / 关系类型 / Schema 导入导出 / Schema 发布 | [manage-kg-schema](./references/knowledge-graph/manage-kg-schema/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 431)May include surrounding context.

md
| 质量监控 / 质量规则 / 数据质量校验 / 质量告警 / 质量试跑 / 质量调度 / 监控对象 | [configure-quality-rule](./references/assets/configure-quality-rule/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 432)May include surrounding context.

md
向量化 / Embedding 入库 / create-work-flow-by-json | [create-unstructured-workflow](./references/unstructured-data/create-unstructured-workflow/SKILL.md) |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 467)May include surrounding context.

md
详见 [相关命令索引](./references/related-commands.md)。

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/dataplan/create-project/references/acceptance-criteria.md (reported line 17)May include surrounding context.

md
## 安全验收

- [ ] 不伪造 `create-project`、`update-project`、`delete-project` 等不存在的公开 CLI 命令。
- [ ] 不直接调用页面内部 `/api/project/basic`、`/api/project/update` 或 `DELETE /api/project/{projectId}`。
- [ ] 白名单更新必须先回读旧值,并确认合并后的新值。
- [ ] 项目成员初始化建议路由到 `manage-project-member`,避免本 Skill 混入成员全生命周期。
- [ ] 任何未来写操作都必须先 HITL 确认命令全文、影响范围和回滚方案。

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/dataplan/create-project/references/related-commands.md (reported line 26)May include surrounding context.

md
|---|---|---|
| `/api/project/basic` | 创建 Basic 模式项目 | 仅作为语义参考,不执行 |
| `/api/project/update` | 更新项目信息 | 仅作为语义参考,不执行 |
| `DELETE /api/project/{projectId}` | 删除项目 | 仅作为语义参考,不执行 |
| `/api/datacatalog/project/search` | 页面项目搜索与创建后验证 | 公开替代为 `get-project-by-name` / `list-projects` |
| `/api/v1/schedule/resource/config/list` | 获取调度资源组 | 仅作为语义参考,不执行 |
| `/api/project/relation` | 查询项目依赖关系 | 公开替代为 `check-project-has-dependency` |

External Script Fetching

High
Category
Supply Chain
Confidence
98% confidence
Finding

The skill instructs users to fetch and immediately execute a remote shell script via curl ... | bash, which bypasses integrity verification, review, and pinning. If the hosting domain, CDN path, TLS trust chain, or upstream script is compromised, the user could execute arbitrary code on the local machine running the agent or operator workflow.

Content

Scanner excerpt · references/dataplan/manage-project-member/SKILL.md (reported line 59)May include surrounding context.

md
**Pre-check: Aliyun CLI >= 3.4.8 required**
> Run `aliyun version` to verify >= 3.4.8. If not installed or version too low,
> run `curl -fsSL https://aliyuncli.alicdn.com/setup.sh | bash` to install/update,
> or see `references/cli-installation-guide.md` for installation instructions.

**Pre-check: Aliyun CLI plugin update required**

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

The markdown provides a copy-pastable curl command using -k and cleartext HTTP to probe a host, encouraging unsafe tool usage with attacker-influenced network parameters. Even if framed as discovery, this normalizes disabling transport protections and could expose app identifiers, enable spoofed responses, or train users to bypass certificate failures instead of fixing trust configuration.

Content

Scanner excerpt · references/dataservice/call-data-service-api/references/pre-call-discovery.md (reported line 102)May include surrounding context.

数据服务调用路径为 /{methodType}/{apiId}?appKey=<AppKey>&env=PROD。用 GET 探一下候选域名即可判断是否命中网关(GET 会被拒,但拒的方式能证明域名对不对):

bash
curl -sS -k -m 8 "http://<候选host>/list/<apiId>?appKey=<AppKey>&env=PROD"

| 返回 | 结论 |

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill instructs users to set runtime.ignore_ssl = True, which disables TLS certificate validation and makes HTTPS connections vulnerable to man-in-the-middle interception or endpoint spoofing. In this skill, those connections carry Alibaba Cloud access credentials and API management actions, so an attacker on the network path could steal secrets or tamper with create/publish requests.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This skill enables ad-hoc SQL, DDL, and Shell execution but does not clearly foreground the risk of destructive or irreversible actions such as DROP, TRUNCATE, ALTER, UPDATE/DELETE without predicates, or harmful shell commands. In this context, the missing warning is significant because the skill is explicitly designed to execute one-off code directly against production-capable data systems.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
99% confidence
Finding

The command downloads a shell script from an external URL and immediately executes it with sudo, creating a classic remote code execution and supply-chain risk. In the context of an agent skill, this is more dangerous because users may follow the snippet verbatim with elevated trust, leading to full host compromise if the remote source is malicious or compromised.

Content

Scanner excerpt · references/dev/manage-resource-file/references/ossutil-mc-install.md (reported line 21)May include surrounding context.

官方一键脚本,自动识别架构并安装到 /usr/local/bin:

bash
sudo -v ; curl https://gosspublic.alicdn.com/ossutil/install.sh | sudo bash

依赖 unzip 或 7z 解压;macOS 自带 unzip,一般无需额外安装。

Chaining Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

The pipeline chains network retrieval directly into privileged shell execution, reducing opportunities for inspection and increasing the chance that users execute attacker-controlled content. This kind of chaining is especially risky in automation or skill documentation because it encourages copy-paste execution with minimal scrutiny.

Content

Scanner excerpt · references/dev/manage-resource-file/references/ossutil-mc-install.md (reported line 21)May include surrounding context.

官方一键脚本,自动识别架构并安装到 /usr/local/bin:

bash
sudo -v ; curl https://gosspublic.alicdn.com/ossutil/install.sh | sudo bash

依赖 unzip 或 7z 解压;macOS 自带 unzip,一般无需额外安装。

External Script Fetching

High
Category
Supply Chain
Confidence
99% confidence
Finding

This repeats the same unsafe pattern on Linux: fetching an external installer script and piping it directly into a privileged shell. A compromise of the CDN, DNS, TLS trust chain, or published script contents could yield arbitrary root-level command execution.

Content

Scanner excerpt · references/dev/manage-resource-file/references/ossutil-mc-install.md (reported line 39)May include surrounding context.

2a. 一键脚本(x86_64 / ARM64 自动识别,安装到 /usr/bin)

bash
sudo -v ; curl https://gosspublic.alicdn.com/ossutil/install.sh | sudo bash

需 unzip 或 7z:sudo yum install -y unzip 或 sudo apt-get install -y unzip。

Static analysis

Detected: suspicious.dynamic_code_execution, suspicious.insecure_tls_verification

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
references/dataservice/call-data-service-api/tests/test_client.py:15

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
references/knowledge-graph/manage-kg-schema/scripts/import-schema.py:184

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
references/dataservice/call-data-service-api/scripts/call-data-service-api.py:130