Back to skill

Security audit

huawei-cloud-ascend-models-deploy

Security checks for vulnerabilities and agentic risk

Overview

The skill is coherent for Ascend model deployment, but it asks agents to run unverified remote shell scripts and detached deployment jobs on user servers.

Review this before installing if the agent will have SSH access to important servers. Only use it when you trust the Huawei OBS-hosted deployment scripts or can pre-stage and verify them yourself, run with the least-privileged account practical, confirm every generated command manually, and keep a separate plan to stop or clean up background deployment processes.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (32)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The description overstates the implemented functionality. This script does align with part of the declared purpose: automated model matching, listing supported models, returning deployment script info, and generating deployment commands for categories such as LLM, VL, Embedding, and Rerank. However, the broader declared skill claims capabilities for inference testing, deployment log viewing, status monitoring, and prerequisite checks, none of which appear in the supplied code. Additionally, the description emphasizes single-machine and dual-machine deployment support, while this code only maps to single-machine script URLs and generates a shell command; there is no explicit dual-machine handling logic. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill documents extensive shell-based remote operations, including command generation and SSH-driven deployment flows, but does not declare any explicit tool scope such as allowed-tools or permissions. That creates a governance gap where an agent may invoke shell-capable execution without a clearly constrained boundary, increasing the chance of unintended or overly powerful command execution.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest trigger list includes generic words such as "deploy", "test", "model list", "inference", and especially "LLM", which can appear in many ordinary conversations outside this specific Huawei Ascend deployment context. Although some terms are domain-specific, the overall trigger set is broad enough to risk unintended invocation because it does not clearly constrain context or provide exclusions.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
93% confidence
Finding

This deployment template downloads a remote script with wget and immediately marks it executable and runs it, creating a classic remote code execution and supply-chain risk. If the remote host, path, or retrieved content is compromised, the agent can execute attacker-controlled shell code on the target server with the privileges of the invoking user.

Content

Scanner excerpt · SKILL.md (reported line 351)May include surrounding context.

LLM / Embedding / Rerank Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/deploy-large-models.sh && chmod 755 /home/modelarts-agent/deploy-large-models.sh && sh /home/modelarts-agent/deploy-large-models.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

VL Multimodal Command Template:

Session Persistence

Medium
Category
Rogue Agent
Confidence
86% confidence
Finding

Using nohup with background execution causes deployment processes to persist independently of the user session and agent lifecycle. In this context, persistent unattended execution can hide failures, make rollback harder, and allow risky commands or compromised scripts to continue running after the interactive session ends.

Content

Scanner excerpt · SKILL.md (reported line 351)May include surrounding context.

LLM / Embedding / Rerank Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/deploy-large-models.sh && chmod 755 /home/modelarts-agent/deploy-large-models.sh && sh /home/modelarts-agent/deploy-large-models.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

VL Multimodal Command Template:

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
93% confidence
Finding

This command repeats the same insecure pattern for VL deployment: download remote shell content, chmod it, and execute it without integrity verification. In a high-privilege deployment context, that exposes the host to supply-chain compromise and arbitrary code execution.

Content

Scanner excerpt · SKILL.md (reported line 356)May include surrounding context.

VL Multimodal Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/single-machine/deploy-qwen3-vl-model.sh && chmod 755 /home/modelarts-agent/deploy-qwen3-vl-model.sh && sh /home/modelarts-agent/deploy-qwen3-vl-model.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

OpenSource Command Template:

Session Persistence

Medium
Category
Rogue Agent
Confidence
86% confidence
Finding

This backgrounded VL deployment persists beyond the controlling session, which is risky when combined with remote script execution and model-serving startup. If something goes wrong or the script is malicious, it can continue operating without immediate visibility or containment.

Content

Scanner excerpt · SKILL.md (reported line 356)May include surrounding context.

VL Multimodal Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/single-machine/deploy-qwen3-vl-model.sh && chmod 755 /home/modelarts-agent/deploy-qwen3-vl-model.sh && sh /home/modelarts-agent/deploy-qwen3-vl-model.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

OpenSource Command Template:

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
93% confidence
Finding

The OpenSource deployment template also executes a remotely fetched script immediately after download. This is dangerous because any compromise of the hosting bucket, DNS path, or content delivery channel turns the deployment flow into arbitrary code execution on the remote server.

Content

Scanner excerpt · SKILL.md (reported line 361)May include surrounding context.

OpenSource Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/open_source/deploy-ai-models.sh && chmod 755 /home/modelarts-agent/deploy-ai-models.sh && sh /home/modelarts-agent/deploy-ai-models.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

3. Dual-machine Deployment

Session Persistence

Medium
Category
Rogue Agent
Confidence
86% confidence
Finding

The OpenSource deployment command similarly detaches execution from the session, reducing operator control over what the script does after launch. Detached persistence is especially dangerous in a deployment skill because it can leave unauthorized services or processes running indefinitely.

Content

Scanner excerpt · SKILL.md (reported line 361)May include surrounding context.

OpenSource Command Template:

bash
nohup bash -c 'export model_name=${model} && export required_cards=${cards} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/open_source/deploy-ai-models.sh && chmod 755 /home/modelarts-agent/deploy-ai-models.sh && sh /home/modelarts-agent/deploy-ai-models.sh ${model} ${cards} ${port}' > /home/modelarts-agent/deploy_${model}.log 2>&1 &

3. Dual-machine Deployment

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
94% confidence
Finding

The dual-machine head-node command downloads and executes a remote script across distributed infrastructure, compounding the blast radius of a compromised artifact. A malicious or altered script could establish persistence, tamper with cluster configuration, or compromise both deployment nodes.

Content

Scanner excerpt · SKILL.md (reported line 372)May include surrounding context.

Head Node Command Template:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/dual-machine/qwen3-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-235b-a22b.sh && sh /home/modelarts-agent/qwen3-235b-a22b.sh head ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_head.log 2>&1 &

Worker Node Command Template:

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

Detached execution on the dual-machine head node increases operational risk because a persistent process can alter cluster state after the interactive session ends. In distributed deployments, this can complicate incident response and prolong the effects of a bad command or compromised script.

Content

Scanner excerpt · SKILL.md (reported line 372)May include surrounding context.

Head Node Command Template:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/dual-machine/qwen3-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-235b-a22b.sh && sh /home/modelarts-agent/qwen3-235b-a22b.sh head ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_head.log 2>&1 &

Worker Node Command Template:

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
94% confidence
Finding

The worker-node variant has the same verified-integrity gap as the head-node command, enabling arbitrary code execution on cluster worker systems. In a multi-node deployment, this broadens compromise from one host to the whole deployment environment.

Content

Scanner excerpt · SKILL.md (reported line 377)May include surrounding context.

Worker Node Command Template:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/dual-machine/qwen3-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-235b-a22b.sh && sh /home/modelarts-agent/qwen3-235b-a22b.sh worker ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_worker.log 2>&1 &

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The worker-node detached process introduces the same persistence issue on another host, making it harder to contain or inspect active deployment actions. Multi-node persistence increases the chance of orphaned or unauthorized services across the cluster.

Content

Scanner excerpt · SKILL.md (reported line 377)May include surrounding context.

Worker Node Command Template:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/dual-machine/qwen3-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-235b-a22b.sh && sh /home/modelarts-agent/qwen3-235b-a22b.sh worker ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_worker.log 2>&1 &

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
94% confidence
Finding

The VL head-node dual-machine command again chains remote download to immediate execution, this time for multimodal deployment infrastructure. Because the skill is explicitly designed to operate over SSH on deployment servers, this pattern materially increases the risk of full server compromise.

Content

Scanner excerpt · SKILL.md (reported line 387)May include surrounding context.

VL Head Node Command:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/dual-machine/qwen3-vl-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-vl-235b-a22b.sh && sh /home/modelarts-agent/qwen3-vl-235b-a22b.sh head ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_head.log 2>&1 &

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The VL head-node command combines nohup persistence with remote script execution on cluster infrastructure, which materially raises the risk profile. Even if intended for convenience, detached long-running processes reduce oversight and can sustain compromise or misconfiguration.

Content

Scanner excerpt · SKILL.md (reported line 387)May include surrounding context.

VL Head Node Command:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/dual-machine/qwen3-vl-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-vl-235b-a22b.sh && sh /home/modelarts-agent/qwen3-vl-235b-a22b.sh head ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_head.log 2>&1 &

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
94% confidence
Finding

The VL worker-node command creates the same download-and-execute exposure on the second node, extending supply-chain and arbitrary code execution risk to the wider cluster. If exploited, an attacker could compromise distributed model infrastructure and any accessible data or credentials on those hosts.

Content

Scanner excerpt · SKILL.md (reported line 393)May include surrounding context.

VL Worker Node Command:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/dual-machine/qwen3-vl-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-vl-235b-a22b.sh && sh /home/modelarts-agent/qwen3-vl-235b-a22b.sh worker ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_worker.log 2>&1 &

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The VL worker-node detached process extends persistence to the secondary node, increasing cluster-wide exposure if the deployment command misbehaves. In a remote-execution skill, this is more dangerous than ordinary background jobs because it is designed to operate over SSH on powerful infrastructure.

Content

Scanner excerpt · SKILL.md (reported line 393)May include surrounding context.

VL Worker Node Command:

bash
nohup bash -c 'export ray_head_ip=${head_ip} && export model_name=${model} && export port=${port} && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-vl-model/dual-machine/qwen3-vl-235b-a22b.sh && chmod 755 /home/modelarts-agent/qwen3-vl-235b-a22b.sh && sh /home/modelarts-agent/qwen3-vl-235b-a22b.sh worker ${head_ip} ${model} ${port}' > /home/modelarts-agent/deploy_${model}_worker.log 2>&1 &

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 419)May include surrounding context.

md
Service URL: http://${IP}:${PORT}/v1/chat/completions

Example request:
curl -X POST http://${IP}:${PORT}/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"${model}","messages":[{"role":"user","content":"hello"}],"max_tokens":256}'

Unbounded Output

Medium
Category
Output Handling
Confidence
88% confidence
Finding

The instruction to return model responses with 'no truncation' can cause the agent to emit arbitrarily large outputs, including prompt echoes, sensitive data present in responses, or unexpectedly massive payloads. In an orchestration context, unbounded output can amplify denial-of-service risk, token exhaustion, and accidental disclosure of logs or model-generated secrets.

Content

Scanner excerpt · SKILL.md (reported line 477)May include surrounding context.

md
| finish_reason | stop |

Model Response:
[Extract full content, no truncation]

Raw Response:
[Full JSON, no truncation]

Unbounded Output

Medium
Category
Output Handling
Confidence
88% confidence
Finding

Requiring full raw JSON output with no truncation can expose internal metadata, long completions, or embedded sensitive content and can overwhelm downstream systems. Because this skill also deals with deployment and testing logs, unbounded raw output increases both disclosure and resource-consumption risk.

Content

Scanner excerpt · SKILL.md (reported line 480)May include surrounding context.

[Extract full content, no truncation]

Raw Response: [Full JSON, no truncation]

text

#### LLM Chat Completions

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 485)May include surrounding context.

LLM Chat Completions

bash
curl -s -X POST http://${IP}:${PORT}/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"${model}","messages":[{"role":"user","content":"${prompt}"}],"max_tokens":1024,"temperature":0.7}'

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file documents enable_thinking with true as the default and states it will output the reasoning process, but it does not warn readers that such output may expose internal reasoning-like content or create privacy/safety concerns if surfaced to end users. Under the markdown-file criteria for missing user warnings, behavior that affects privacy or user-visible output integrity should include an explicit warning.

Content

No source excerpt is available for this finding.

File System Enumeration

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code scans file system directories looking for sensitive files. This could be reconnaissance for credential theft.

Content

Scanner excerpt · references/prerequisites.md (reported line 52)May include surrounding context.

df -h /home

Check directory

ls -la /home/modelarts-agent

text

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/task-deploy-model.md (reported line 34)May include surrounding context.

LLM/Embedding/Rerank:

bash
nohup bash -c 'export model_name=Qwen3-14B && export required_cards=1 && export port=8080 && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/deploy-large-models.sh && chmod 755 /home/modelarts-agent/deploy-large-models.sh && sh /home/modelarts-agent/deploy-large-models.sh Qwen3-14B 1 8080' > /home/modelarts-agent/deploy_Qwen3-14B.log 2>&1 &

VL Multimodal:

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/task-deploy-model.md (reported line 39)May include surrounding context.

LLM/Embedding/Rerank:

bash
nohup bash -c 'export model_name=Qwen3-14B && export required_cards=1 && export port=8080 && wget -P /home/modelarts-agent/ https://documentation-samples-17.obs.cn-north-9.myhuaweicloud.com/solution-as-code-publicbucket/solution-as-code-module/quickly-deploy-llm-on-modelarts-lite-devserver/userdata/deploy-large-models/single-machine/deploy-large-models.sh && chmod 755 /home/modelarts-agent/deploy-large-models.sh && sh /home/modelarts-agent/deploy-large-models.sh Qwen3-14B 1 8080' > /home/modelarts-agent/deploy_Qwen3-14B.log 2>&1 &

VL Multimodal:

Static analysis

No suspicious patterns detected.