Back to skill

Security audit

huawei-cloud-cce-chaos-experiment

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real Huawei CCE chaos-drill skill, but it needs Review because it can stop cloud cluster nodes and its executor does not enforce the documented confirmation and validation gates.

Install only if you are prepared to run a controlled chaos experiment against Huawei Cloud CCE. Use least-privilege Huawei credentials, isolate KUBECONFIG to a temporary file instead of /root/.kube/config, run --dry-run first, manually verify validation.json is fresh and all_compatible=true, and require an explicit human approval step before any non-dry-run execution.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/execute_experiment.py:103
Finding

Destructive CCE node shutdown bypasses declared confirmation and validation gates

Content
View full analysis

Vulnerability Details

File Location: scripts/execute_experiment.py:103-106, 186-195
Supporting Location: scripts/generate_experiment.py:31-35, 120-129
Vulnerability Type: Missing authorization enforcement for a destructive cloud operation
Risk Level: High

Vulnerable Code

scripts/execute_experiment.py:103-106 defines dry-run support but provides no explicit confirmation or approval argument:

python
parser = argparse.ArgumentParser(description="Execute CCE AZ power outage experiment")
parser.add_argument("--experiment-dir", required=True, help="Experiment directory path")
parser.add_argument("--dry-run", action="store_true", help="Simulate without actual API calls")
parser.add_argument("--auto-rollback", action="store_true", help="Auto-rollback on failure")
args = parser.parse_args()

scripts/execute_experiment.py:186-195 performs the real shutdown whenever --dry-run is absent:

python
print(f"[2/6] Shutdown: executing BatchStopServers (os_stop={os_stop}) ...")

phase_start = now_iso()
if args.dry_run:
    print("  [DRY RUN] Skipping actual shutdown")
else:
    try:
        result = batch_stop_servers(instance_ids, region, os_stop)
        job_id = result.get("job_id", "")
        print(f"  ✓ BatchStopServers command sent (job_id: {job_id})")

The generated configuration explicitly declares that confirmation is required in scripts/generate_experiment.py:120-129:

python
"safety": {
    "max_duration_seconds": args.duration,
    "auto_rollback_on_failure": False,
    "require_confirmation": True,
    "check_pdb": True,
    "check_cross_az_capacity": True,
},

The validation file is loaded in scripts/generate_experiment.py:31-35, but its result is not checked before producing an executable experiment:

python
validation = None
if args.validation_file and os.path.exists(args.validation_file):
    with open(args.validation_f
...[truncated 3485 chars]
Remediation
View remediation

Remediation Suggestions

  1. Enforce explicit confirmation in the executor. For non-dry-run execution, require a value bound to the experiment, such as:

    bash
    python3 scripts/execute_experiment.py \
      --experiment-dir "$EXP_DIR" \
      --confirm-experiment "<experiment-name>"
    

    Reject execution unless the supplied value exactly matches the loaded experiment name or a newly generated approval token.

  2. Honor safety.require_confirmation. Read this field before any cloud mutation. If it is true, fail closed unless confirmation has been obtained through an explicit, auditable mechanism.

  3. Require successful validation. Refuse generation and execution unless a validation artifact exists and contains all_compatible: true. Do not merely load the file.

  4. Cryptographically or structurally bind validation to the action. Record and verify the validated region, cluster ID, AZ, node names, ECS instance IDs, discovery timestamp, and a digest of the target set. Reject stale or mismatched validation results.

  5. Revalidate immediately before shutdown. Confirm that every ECS instance still corresponds to a node in the selected cluster and AZ, and rerun the critical capacity and single-AZ checks.

  6. Use fail-closed target constraints. Reject empty IDs, duplicate IDs, unknown nodes, mismatched AZ labels, and instance IDs not returned by fresh discovery for the selected cluster.

  7. Make dry-run the safe default. Require a separate explicit option such as --execute in addition to confirmation before issuing BatchStopServers.

  8. Improve rollback guarantees. Enable rollback-on-failure by default, persist rollback state before shutdown, and prominently report any failed startup operation for immediate intervention.

Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (111)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description presents an end-to-end AZ power outage chaos drill for Huawei Cloud CCE, but the actual code chunk only performs prerequisite validation and kubectl installation. That is a materially different primary purpose from executing or orchestrating the drill itself. While environment checks can be a supporting component of such a skill, this code alone does not implement the described experiment workflow. Additionally, it contains an undeclared operational capability to download/install/remove kubectl binaries, and it checks for an hcloud CLI dependency that does not clearly align with the Huawei CCE description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The description promises an end-to-end AZ power outage chaos experiment, including fault injection (stopping all nodes in an AZ), observation of workload behavior, restoration, and recovery verification. The code provided only performs preparatory discovery: it queries Huawei Cloud CCE/ECS metadata and Kubernetes nodes/pods/workloads, then outputs a report. Its primary purpose is narrower than declared and lacks the central destructive/restorative actions of the drill. Accessed resources (CCE, ECS, Kubernetes) are consistent with the domain, but the actual implemented capability is only the discovery phase, so the declared description materially overstates behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a full fault-drill/chaos-experiment workflow affecting CCE nodes in an AZ. However, this code chunk is a read-only discovery utility. It invokes kubectl get on nodes, pods, services, ingress, and configmaps, correlates objects by namespace/service references, and saves the findings to a JSON file. That is at most a preparatory/supporting step for an outage drill, not the described end-to-end experiment itself. Because the primary purpose and key capabilities in the description are absent from the code, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents an end-to-end AZ power outage chaos experiment runner, including preparation, fault injection, recovery, and optional log analysis. The supplied code chunk is a monitoring utility only: it queries nodes and pods with kubectl, detects state changes, and records simple timelines/rescheduling events. While this is related to the observation phase of such a drill, it does not implement the main declared capabilities, especially node shutdown/startup and full experiment orchestration. Therefore the description materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a full AZ power outage chaos experiment workflow, including preparation, fault injection, observation, recovery, and optional log analysis. The supplied code chunk only implements the recovery/rollback portion: it starts instances and checks node/pod recovery status. While rollback is related to the broader experiment, this chunk’s actual behavior is materially narrower and its primary purpose differs from the declared end-to-end experiment capability. Therefore, the description does not accurately represent what this specific code chunk actually does.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill centers on shutting down all nodes in one AZ, which can disrupt workloads, reduce capacity, and potentially cause outages, yet the documentation does not present an up-front, explicit warning banner about service impact and cluster availability before operational steps begin. In the context of a chaos experiment skill, omission of a strong warning and authorization gate makes accidental or under-informed destructive use significantly more dangerous.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

The explicit instruction to obtain kubeconfig represents credential acquisition for cluster administration. In this skill's context, that credential enables visibility into cluster state and potentially destructive operations, so poor lifecycle controls around issuance and storage are security-relevant.

Content

Scanner excerpt · SKILL.md (reported line 81)May include surrounding context.

Step 1.5: Get Credentials & Discover Node AZ Distribution

bash
# Obtain kubeconfig
hcloud CCE CreateKubernetesClusterCert --cluster_id=<id> --cli-region=<region> --duration=30 --cli-output=json > /root/.kube/config

# Discover node AZ distribution

Credential Access

High
Category
Privilege Escalation
Confidence
98% confidence
Finding

The skill retrieves cluster credentials and writes them to /root/.kube/config, a privileged location containing sensitive authentication material. Storing generated kubeconfig in a default high-value path increases exposure to credential leakage, accidental reuse, overwriting existing admin config, and misuse by subsequent commands.

Content

Scanner excerpt · SKILL.md (reported line 82)May include surrounding context.

Step 1.5: Get Credentials & Discover Node AZ Distribution

bash
# Obtain kubeconfig
hcloud CCE CreateKubernetesClusterCert --cluster_id=<id> --cli-region=<region> --duration=30 --cli-output=json > /root/.kube/config

# Discover node AZ distribution
KUBECONFIG=/root/.kube/config kubectl get nodes --show-labels

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

Referencing KUBECONFIG directly is not inherently unsafe, but here it points to a root-owned credential store and is used to drive cluster interrogation. That makes it a valid credential-handling concern rather than a mere environment-variable reference.

Content

Scanner excerpt · SKILL.md (reported line 85)May include surrounding context.

hcloud CCE CreateKubernetesClusterCert --cluster_id= --cli-region= --duration=30 --cli-output=json > /root/.kube/config

Discover node AZ distribution

KUBECONFIG=/root/.kube/config kubectl get nodes --show-labels

text
Present node-to-AZ mapping to the user.

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

Referencing KUBECONFIG directly is not inherently unsafe, but here it points to a root-owned credential store and is used to drive cluster interrogation. That makes it a valid credential-handling concern rather than a mere environment-variable reference.

Content

Scanner excerpt · SKILL.md (reported line 85)May include surrounding context.

hcloud CCE CreateKubernetesClusterCert --cluster_id= --cli-region= --duration=30 --cli-output=json > /root/.kube/config

Discover node AZ distribution

KUBECONFIG=/root/.kube/config kubectl get nodes --show-labels

text
Present node-to-AZ mapping to the user.

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

This use of KUBECONFIG propagates administrative credentials into custom automation, which can inadvertently broaden access or persist sensitive context in generated artifacts. In a chaos experiment workflow, credential misuse can directly facilitate disruptive actions against production-like clusters.

Content

Scanner excerpt · SKILL.md (reported line 112)May include surrounding context.

Step 1.8: Discover Nodes & Pods in Target AZ

bash
KUBECONFIG=/root/.kube/config python3 scripts/discover_cce.py \
    --region <region> --cluster-id <id> --az <az> \
    --include-pods --output "$EXP_DIR/discovery.json"

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

This use of KUBECONFIG propagates administrative credentials into custom automation, which can inadvertently broaden access or persist sensitive context in generated artifacts. In a chaos experiment workflow, credential misuse can directly facilitate disruptive actions against production-like clusters.

Content

Scanner excerpt · SKILL.md (reported line 112)May include surrounding context.

Step 1.8: Discover Nodes & Pods in Target AZ

bash
KUBECONFIG=/root/.kube/config python3 scripts/discover_cce.py \
    --region <region> --cluster-id <id> --az <az> \
    --include-pods --output "$EXP_DIR/discovery.json"

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

Validation scripts inherit the same privileged KUBECONFIG context, compounding the credential-exposure pattern across the workflow. Reuse of broad credentials across discovery, validation, and execution phases is riskier than phase-specific least-privilege access.

Content

Scanner excerpt · SKILL.md (reported line 119)May include surrounding context.

Step 1.9: Validate Compatibility [CRITICAL GATE]

bash
KUBECONFIG=/root/.kube/config python3 scripts/validate_targets.py \
    --discovery-file "$EXP_DIR/discovery.json" \
    --cluster-id <id> --az <az> --region <region> \
    --output "$EXP_DIR/validation.json"

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

Validation scripts inherit the same privileged KUBECONFIG context, compounding the credential-exposure pattern across the workflow. Reuse of broad credentials across discovery, validation, and execution phases is riskier than phase-specific least-privilege access.

Content

Scanner excerpt · SKILL.md (reported line 119)May include surrounding context.

Step 1.9: Validate Compatibility [CRITICAL GATE]

bash
KUBECONFIG=/root/.kube/config python3 scripts/validate_targets.py \
    --discovery-file "$EXP_DIR/discovery.json" \
    --cluster-id <id> --az <az> --region <region> \
    --output "$EXP_DIR/validation.json"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 14)May include surrounding context.

md
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/cli-installation-guide.md (reported line 102)May include surrounding context.

md
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/cli-installation-guide.md (reported line 105)May include surrounding context.

md
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 16)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 33)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 46)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 48)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 49)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/collect_logs.py (reported line 50)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/discover_cce.py (reported line 24)May include surrounding context.

python
- [ ] Supports two modes: cluster discovery (no `--cluster-id`) and node/pod discovery (with `--cluster-id` + `--az`)
- [ ] `--include-pods` flag adds Pod listing for each target node
- [ ] ECS instance IDs are obtained via `hcloud ECS ListServersDetails` (private-IP → ECS-ID mapping), not from `spec.providerID`
- [ ] Uses `KUBECONFIG` environment variable for kubectl access

### Validation
- [ ] `validate_targets.py` checks all target nodes are `Ready`

Static analysis

No suspicious patterns detected.