Back to skill

Security audit

alibabacloud-ecs-diagnose

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly coherent Alibaba Cloud ECS diagnostic tool, but it gives an agent broad cloud and guest-OS troubleshooting authority with some under-scoped automatic execution guidance.

Install only if you are comfortable letting the agent use Aliyun CLI credentials to inspect ECS resources and, in some cases, run diagnostic commands inside instances. Keep diagnosis read-only by default, do not use the curl-to-bash installer unless you manually reviewed it, and require explicit confirmation before any Cloud Assistant command or repair action such as opening ports, binding EIPs, rebooting instances, or resetting passwords.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (49)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description promises a broad ECS instance diagnostics capability covering many troubleshooting scenarios and specifically mentions GuestOS diagnostics via Cloud Assistant. The supplied code only implements one narrow workflow: EBS disk performance diagnosis using EBS RPC APIs (DescribeLensMonitorDisks, CreateDiagnoseReport, DescribeDiagnoseReport). It can list disks, run a performance diagnosis for a disk, poll for results, and print recommendations. While disk IO/performance is one of the declared trigger areas, the actual code does not match the description's primary purpose of comprehensive ECS diagnostics. This is a material scope and capability mismatch, not just an implementation detail.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
92% confidence
Finding

The skill includes a classic remote-script execution pattern using curl piped to bash, even if described as a fallback. Referencing this pattern in an operational skill materially increases the chance that users or agents adopt an unsafe installation path, enabling supply-chain compromise or arbitrary code execution if the remote content is tampered with.

Content

Scanner excerpt · SKILL.md (reported line 42)May include surrounding context.

md
> (download the official universal binary, verify it, and install into your user
> directory — no remote script, no sudo, no system-level writes; you can see the
> full scope of changes before running anything).
> The one-line remote installer `curl -fsSL https://aliyuncli.alicdn.com/setup.sh | bash`
> is a **manual fallback only** — it pipes a remote script into bash and may write
> into system directories via sudo, so do not run it unless you have reviewed it
> and explicitly accepted that risk.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 196)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 279)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 302)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 311)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 425)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 466)May include surrounding context.

md
> Region-traversal method: see `references/remote-connection-diagnose-design.md` §1.2.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill claims it never creates cloud resources automatically, yet it instructs automatic creation of EBS diagnose reports. This inconsistency can bypass user expectations about read-only operation and lead to unanticipated API-side state changes, billing effects, or audit issues.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
93% confidence
Finding

The document includes a classic curl-pipe-to-bash remote installer pattern. Even though it is framed as a manual fallback and explicitly warned against, embedding this command in a skill still presents a high-risk supply-chain and arbitrary code execution path if copied or later surfaced by an agent.

Content

Scanner excerpt · references/cli-installation-guide.md (reported line 13)May include surrounding context.

md
> Download the official universal binary below, verify the version, and install into a
> user-writable directory (e.g. `~/bin`). You can see the full scope of changes before
> running anything. The one-line remote installer
> `curl -fsSL https://aliyuncli.alicdn.com/setup.sh | bash` is a **manual fallback
> only** — it pipes a remote script into bash and may write into system directories via
> sudo, so review it before using.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The examples explicitly show opening public ingress, including 0.0.0.0/0 for HTTP and targeted SSH ingress, which can weaken perimeter controls far beyond what is needed for diagnosis. In the context of an agent skill, these documented actions could be operationalized to expose services to the internet, increasing the risk of unauthorized access, scanning, and compromise.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This diagnostics guide goes well beyond read-only troubleshooting and explicitly authorizes state-changing operations such as opening security groups, allocating and binding EIPs, starting/rebooting instances, and resetting passwords. In an agent skill, bundling these actions into the diagnostic workflow creates a strong risk of unauthorized or overly broad infrastructure modification, especially if the agent executes steps automatically or with weak confirmation boundaries.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The guide includes password reset commands for ECS instances, which is an account takeover capability rather than a diagnostic function. If misused, an agent could lock out legitimate administrators, seize control of systems, and trigger service interruption through the required reboot.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill explicitly instructs reading local reference files and checking CLI configuration, but it declares no tool scope or permissions boundary. That creates an authorization gap where an agent may access filesystem or environment-backed data without a clear least-privilege contract, increasing the chance of unintended local data exposure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill advertises activation on generic terms such as "instance", "server", "network", "slow", "diagnose", and "troubleshoot" without clear contextual constraints or exclusion examples. In a markdown skill description, such broad triggers can match many unrelated conversations and cause unintended invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instructions mandate that the empty-result template be output verbatim in English and not translated, regardless of the conversation language. This is a natural-language policy issue because it forces a specific language without user opt-in or a documented locale-choice mechanism.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The restart-diagnosis instructions require outputting a conclusion represented by a Chinese Unicode-escape template. This forces a specific language for user-facing output without opt-in, which violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill authorizes execution of guest OS commands based on implied consent from symptom keywords, without an explicit warning that commands will run inside the instance. This weakens informed consent for instance-side access and can expose sensitive system data or alter forensic expectations in production environments.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
96% confidence
Finding

The skill explicitly allows autonomous execution of system-level diagnostics 'without asking' when certain symptoms appear. Because those actions run commands on customer instances, this is an overreach of agent authority and can lead to unauthorized host access, sensitive data exposure, and surprise production interactions.

Content

Scanner excerpt · SKILL.md (reported line 451)May include surrounding context.

md
1. **Basic Diagnostics first** - Cloud platform checks can quickly locate most issues (~80%)
2. **Deep Diagnostics requires confirmation** - Always get user approval before executing system commands

> **Exception**: When the user's initial request explicitly describes symptoms that require system-level diagnosis (e.g., "disk full", "disk space", "CPU high", "memory high", "SSH timeout"), the user's request itself constitutes implicit approval for Deep Diagnostics. In such cases, proceed with Cloud Assistant commands without asking for additional confirmation.

1. **Security group focus** - ~70% of connectivity issues stem from security group misconfigurations
2. **Windows adaptation** - Use PowerShell commands and `RunPowerShellScript` type for Windows instances

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The notes instruct the agent to continue after API failure, contradicting earlier mandatory stop conditions for permission failures and empty instance lookups. In practice this can cause false-negative or misleading diagnostic results, masking access-control errors or missing-target conditions and prompting unsafe operator decisions.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 314)May include surrounding context.

cat ~/.aliyun/config.json aliyun configure get

NEVER ask user to input credentials directly

read -p "Enter your AccessKey ID: " AK

text

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 460)May include surrounding context.

  1. If yes, execute Deep Diagnostics commands
text

#### ❌ INCORRECT - Execute everything without asking

  1. Execute Basic Diagnostics

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 501)May include surrounding context.

  1. If yes, execute Deep Diagnostics commands
text

#### ❌ INCORRECT - Execute everything without asking

  1. Execute Basic Diagnostics

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 464)May include surrounding context.

text
1. Execute Basic Diagnostics
2. Automatically execute Deep Diagnostics
3. Present all results

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/cli-installation-guide.md (reported line 14)May include surrounding context.

md
> user-writable directory (e.g. `~/bin`). You can see the full scope of changes before
> running anything. The one-line remote installer
> `curl -fsSL https://aliyuncli.alicdn.com/setup.sh | bash` is a **manual fallback
> only** — it pipes a remote script into bash and may write into system directories via
> sudo, so review it before using.

### macOS

Static analysis

No suspicious patterns detected.