Back to skill

Security audit

huawei-cloud-skill-tester

Security checks for vulnerabilities and agentic risk

Overview

This skill is a real Huawei Cloud test runner, but it can execute untrusted tested-skill content with local and cloud credentials, so it needs careful review before use.

Use this only in an isolated workspace with a disposable Huawei Cloud account and narrowly scoped temporary credentials. Keep ALLOW_WRITES=0 unless you have reviewed every generated command, avoid --all-installed and default sibling scanning unless intended, and treat any skill being tested as untrusted code.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/tier1/phase-4-execute-tests.sh:141
Finding

Untrusted tested-Skill content is automatically executed with host and cloud credentials

Content
View full analysis

Vulnerability Details

File Location: scripts/tier1/phase-4-execute-tests.sh:141-170, 238-249, 332-375
Related Data-Flow Locations: scripts/tier1/phase-1-skill-analysis.sh:283-326; scripts/tier1/phase-3-gen-testcases.sh:164-205; scripts/tier2/phase-6-full-flow.sh:164-227
Vulnerability Type: Arbitrary command and code execution from untrusted Skill content
Risk Level: High

Complete Code Snippets

Phase 1 extracts executable content from the tested Skill's documentation and classifies writes using a limited keyword heuristic:

python
elif cl.startswith('bash ') and 'scripts/' in cl:
    # 自带脚本(bash scripts/xxx.sh ...)也是可执行命令, 提取为正例
    cmd_id += 1
    is_write = any(kw in cl.lower() for kw in [
        'create', 'delete', 'update', 'destroy',
        'activate', 'reclaim', 'cleanup'
    ])
    commands.append({
        'id': 'CMD-%02d' % cmd_id,
        'source': 'SKILL.md-bash-block',
        'description': cl[:80],
        'command': cl,
        'executor': 'script',
        'is_write': is_write
    })
elif cl.startswith('hcloud '):
    cmd_id += 1
    is_write = any(kw in cl.lower() for kw in [
        'create', 'delete', 'update', 'destroy'
    ])
    clean_cmd = re.sub(r'<[^>]+>', '', cl).strip()
    commands.append({
        'id': 'CMD-%02d' % cmd_id,
        'source': 'SKILL.md-bash-block',
        'description': clean_cmd[:80],
        'command': clean_cmd,
        'executor': 'cli',
        'is_write': is_write
    })

Phase 3 converts the extracted command into an executable test case:

python
functional_cases.append({
    'id': f'TC-F-{tc_f_id:02d}',
    'name': cmd.get('description', cmd_text[:60])
            if cmd.get('description') else f"命令-{tc_f_id:02d}",
    'type': '正向' if not is_write else '变更',
    'command': cmd_text,
    'expected': 'SDK调用成功并返回数据'
                if executor == 'sdk'
                else ('CLI命令执行成功'
                      if executor == 'cli'
                      else '脚本执行成功'),

...[truncated 6074 chars]
Remediation
View remediation

Remediation Suggestions

  1. Do not execute documentation-derived shell strings

    • Remove all bash -c execution of content extracted from a tested Skill.
    • Parse supported commands into argument arrays.
    • Allowlist the exact executable, service, operation, and permitted options.
    • Reject shell metacharacters, substitutions, redirections, pipelines, newlines, and additional command separators.
  2. Treat every tested Skill as hostile code

    • Run packaged scripts and generated SDK snippets only inside a disposable container or virtual machine.
    • Use a read-only mount for the tested Skill.
    • Do not mount the user's home directory, SSH configuration, cloud configuration, agent state, or unrelated workspace files.
    • Apply process, CPU, memory, filesystem, and execution-time limits.
  3. Remove credentials from untrusted process environments

    • Construct a minimal environment rather than passing os.environ.
    • Do not expose long-lived AK/SK values to tested code.
    • Use short-lived, narrowly scoped credentials for a disposable test account where live testing is necessary.
    • Route approved cloud operations through a trusted broker that validates service, operation, resource scope, region, and request parameters.
  4. Replace heuristic write detection

    • Do not infer authorization requirements from keywords such as create or delete.
    • Resolve supported operations against structured metadata and classify their effects explicitly.
    • Treat unknown or compound commands as unsafe and refuse execution.
  5. Enforce exact per-operation approval

    • Present the normalized executable, argument vector, target account, region, resource scope, and credential permissions.
    • Bind approval to a hash of the exact command or script.
    • Revalidate the hash immediately before execution.
    • Require approval for all untrusted scripts and code snippets, not only operations labeled as writes.
  6. **Separate static analysis f ...[truncated 544 chars]

Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (71)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The documented purpose frames the skill primarily as a testing framework, but the behavior includes bootstrap/dependency installation, remote downloads, and modification of the user's home environment. That mismatch can mislead operators into approving a test run that also performs software installation and persistent local changes, expanding supply-chain and host-integrity risk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The documented purpose frames the skill primarily as a testing framework, but the behavior includes bootstrap/dependency installation, remote downloads, and modification of the user's home environment. That mismatch can mislead operators into approving a test run that also performs software installation and persistent local changes, expanding supply-chain and host-integrity risk.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 65)May include surrounding context.

md
**Dependency**: Quality telemetry is collected automatically via `skill-quality-cli` (installed by `scripts/ensure_cli.sh` if absent).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 77)May include surrounding context.

md
**Dependency**: Quality telemetry is collected automatically via `skill-quality-cli` (installed by `scripts/ensure_cli.sh` if absent).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

md
> Full per-phase implementation specs (steps, pass criteria, JSON fields) are in `references/phase-details.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 115)May include surrounding context.

md
> Full per-phase implementation specs (steps, pass criteria, JSON fields) are in `references/phase-details.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 232)May include surrounding context.

md
> Full per-phase implementation specs (steps, pass criteria, JSON fields) are in `references/phase-details.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 115)May include surrounding context.

md
> Detailed steps, chain-verification rules, and JSON schemas: `references/phase-details.md` and `references/output-schema-spec.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 228)May include surrounding context.

md
> Detailed steps, chain-verification rules, and JSON schemas: `references/phase-details.md` and `references/output-schema-spec.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 275)May include surrounding context.

md
> Detailed steps, chain-verification rules, and JSON schemas: `references/phase-details.md` and `references/output-schema-spec.md`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 181)May include surrounding context.

md
bash scripts/tier1/phase-1-skill-analysis.sh --skill "huawei-cloud-bss-voucher-manage"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 182)May include surrounding context.

md
bash scripts/tier2/phase-5-orchestration.sh --skills "skill-a, skill-b"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 190)May include surrounding context.

md
bash scripts/tier2/phase-5-orchestration.sh --skills "skill-a, skill-b"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 193)May include surrounding context.

md
bash scripts/tier2/phase-6-full-flow.sh --skill "huawei-cloud-rds-intelligent-service"

External Script Fetching

High
Category
Supply Chain
Confidence
95% confidence
Finding

The script fetches a remote manifest and download URL, then retrieves and installs a tarball from that remote source. Although HTTPS and optional SHA-256 verification are used, the integrity model is weak because the manifest and checksum come from the same remote service, and if checksum tools are unavailable the script proceeds without verification at all.

Content

Scanner excerpt · scripts/install_cli.sh (reported line 28)May include surrounding context.

sh
API_URL="https://skillsapi.developer.myhuaweicloud.com/api/quality/cli/latest"

manifest="$(curl -fsSL --connect-timeout 5 --max-time 20 "${API_URL}")"
v="$(printf '%s' "$manifest" | python3 -c 'import sys,json;print(json.load(sys.stdin)["version"])')"
case "$(uname -m)" in
    x86_64|amd64) plat="linux-x86_64" ;;

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · scripts/tier1/phase-1-skill-analysis.sh (reported line 357)May include surrounding context.

sh
if ' ' not in clean_cmd and re.match(r'^[\w./{}\[\]$~-]+$', clean_cmd):
                continue
            if not (clean_cmd.startswith('hcloud ') or clean_cmd.startswith('python3 ')
                   or clean_cmd.startswith('bash ') or clean_cmd.startswith('curl ')
                   or clean_cmd.startswith('from ') or clean_cmd.startswith('sh ')):
                continue
            exe = 'cli'

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises executable behavior involving shell, filesystem access, environment inspection, and likely networked testing, but it does not declare an explicit tool/permission scope. That creates an authorization ambiguity where a caller may invoke a high-impact testing skill without clear least-privilege boundaries, increasing the chance of unexpected command execution, file writes, or credential-adjacent access.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains broad terms such as "verification", "e2e", "跑测试", and similar generic phrases that are likely to collide with ordinary user requests. Because this skill can execute shell scripts, inspect installed sibling skills, and potentially run live cloud actions, accidental triggering materially increases the chance of unintended high-impact operations.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
67% confidence
Finding

The skill intentionally persists phase JSON, archives prior runs, and preserves reports across executions to support chained testing and resumption. While operationally useful, this creates session persistence and residual-data risk because test artifacts, resource identifiers, execution logs, and possibly cloud response data remain on disk across runs and could be accessed later by other local users or reused unintentionally.

Content

Scanner excerpt · SKILL.md (reported line 38)May include surrounding context.

md
### Core Design Principles

1. **Chain Verification** — Before each Phase, check that the previous phase's JSON exists; if missing, refuse to execute
2. **Agent-proof** — Write operations require user confirmation for each item; automatic gate bypassing is not allowed
3. **Three-Track Layering** — Clear gates between Tiers; Tier 1 must be completed before entering Tier 2
4. **Batch Repeatable** — Supports `--skills "skill-a,skill-b"` or `--all-installed`
5. **Fallback Strategy** — When only 1 skill, Phase 5/6 automatically downgrade to single-skill lifecycle testing. **Sibling auto-scan is ON by default** (Phase 5/6 automatically looks for other `huawei-cloud-*` skills in the same directory as the skill under test to run orchestration combos). Use `--no-siblings` to opt-out.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 71)May include surrounding context.

md
1. **hcloud CLI** installed and authenticated (for Tier 2 CLI mode testing) — Reference: https://support.huaweicloud.com/qs-hcli/hcli_02_003.html
2. **Python 3.8+** + `huaweicloudsdk` packages (for SDK mode testing) — SDK Reference: https://console.huaweicloud.com/apiexplorer/#/sdkcenter
3. **Huawei Cloud AK/SK** — auto-scans all environment variables prefixed `HUAWEI` / `HW` / `HWC` (key-value pairs whose keys carry AK/SK markers). **If missing, the framework emits the env-var setup template to stderr and exits 77. NEVER ask the user to type AK/SK in chat; the user must set env vars in their shell profile out-of-band and re-run.** (full protocol: `references/agent-protocol.md`).
4. **Target Skill** must be under `$SKILL_INSTALL_DIR/` (auto-detected: `~/.agents/skills/` → `~/.hermes/skills/` → default) or a user-specified path
5. **jq** command (all JSON processing depends on it)
6. **API Reference**: https://console.huaweicloud.com/apiexplorer/#/openapi

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 290)May include surrounding context.

md
1. **hcloud CLI** installed and authenticated (for Tier 2 CLI mode testing) — Reference: https://support.huaweicloud.com/qs-hcli/hcli_02_003.html
2. **Python 3.8+** + `huaweicloudsdk` packages (for SDK mode testing) — SDK Reference: https://console.huaweicloud.com/apiexplorer/#/sdkcenter
3. **Huawei Cloud AK/SK** — auto-scans all environment variables prefixed `HUAWEI` / `HW` / `HWC` (key-value pairs whose keys carry AK/SK markers). **If missing, the framework emits the env-var setup template to stderr and exits 77. NEVER ask the user to type AK/SK in chat; the user must set env vars in their shell profile out-of-band and re-run.** (full protocol: `references/agent-protocol.md`).
4. **Target Skill** must be under `$SKILL_INSTALL_DIR/` (auto-detected: `~/.agents/skills/` → `~/.hermes/skills/` → default) or a user-specified path
5. **jq** command (all JSON processing depends on it)
6. **API Reference**: https://console.huaweicloud.com/apiexplorer/#/openapi

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 333)May include surrounding context.

md
1. **hcloud CLI** installed and authenticated (for Tier 2 CLI mode testing) — Reference: https://support.huaweicloud.com/qs-hcli/hcli_02_003.html
2. **Python 3.8+** + `huaweicloudsdk` packages (for SDK mode testing) — SDK Reference: https://console.huaweicloud.com/apiexplorer/#/sdkcenter
3. **Huawei Cloud AK/SK** — auto-scans all environment variables prefixed `HUAWEI` / `HW` / `HWC` (key-value pairs whose keys carry AK/SK markers). **If missing, the framework emits the env-var setup template to stderr and exits 77. NEVER ask the user to type AK/SK in chat; the user must set env vars in their shell profile out-of-band and re-run.** (full protocol: `references/agent-protocol.md`).
4. **Target Skill** must be under `$SKILL_INSTALL_DIR/` (auto-detected: `~/.agents/skills/` → `~/.hermes/skills/` → default) or a user-specified path
5. **jq** command (all JSON processing depends on it)
6. **API Reference**: https://console.huaweicloud.com/apiexplorer/#/openapi

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/lib/utils.sh (reported line 440)May include surrounding context.

sh
1. **hcloud CLI** installed and authenticated (for Tier 2 CLI mode testing) — Reference: https://support.huaweicloud.com/qs-hcli/hcli_02_003.html
2. **Python 3.8+** + `huaweicloudsdk` packages (for SDK mode testing) — SDK Reference: https://console.huaweicloud.com/apiexplorer/#/sdkcenter
3. **Huawei Cloud AK/SK** — auto-scans all environment variables prefixed `HUAWEI` / `HW` / `HWC` (key-value pairs whose keys carry AK/SK markers). **If missing, the framework emits the env-var setup template to stderr and exits 77. NEVER ask the user to type AK/SK in chat; the user must set env vars in their shell profile out-of-band and re-run.** (full protocol: `references/agent-protocol.md`).
4. **Target Skill** must be under `$SKILL_INSTALL_DIR/` (auto-detected: `~/.agents/skills/` → `~/.hermes/skills/` → default) or a user-specified path
5. **jq** command (all JSON processing depends on it)
6. **API Reference**: https://console.huaweicloud.com/apiexplorer/#/openapi

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

The skill design includes auto-execution behavior for parts of the pipeline and phase progression, and elsewhere states that the framework runs real-environment tests. Even with write gating, automatic execution of read operations and phase logic against live cloud resources can create unintended actions, information disclosure, or costly API usage if the skill is triggered accidentally or with overly broad target selection.

Content

Scanner excerpt · SKILL.md (reported line 309)May include surrounding context.

md
| User unsatisfied with derived orchestration scenarios | Manually edit the derived scenario or skip it; Phase 5 derivation is metadata-only, not executed |
| Phase 4 write op with no `ALLOW_WRITES=1` | Skipped with `status=skip`, no resource_changes recorded |
| Phase 4 hits missing business params (e.g. coupon_id) | Marked `status=warn`, surfaces as `manual_test_items` in the report; user must supply real data and retry |
| Cross-skill data flow mismatch | Logged in Phase 5 as `data_flow_tests` candidate; not auto-executed |
| Orphaned resources detected after E2E flow | Listed in `phase-6-summary.json` under `cleanup.manual_required` with concrete cleanup commands |

## Agent Protocol — Credential Request

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The acceptance criteria and instructions are written in Chinese and do not indicate any user language choice or opt-in. Under the stated policy, forcing a specific language/locale without offering a choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.