Install
openclaw skills install @huaweicloudskill/huawei-cloud-cce-chaos-experimentHuawei Cloud CCE AZ power outage chaos experiment — end-to-end fault drill: discover clusters and node AZ distribution, validate cross-AZ capacity, shut down all CCE nodes in one AZ to simulate a power failure, observe Pod eviction and rescheduling to surviving AZs, then start nodes back and verify recovery, and optionally analyze logs for an impact report. Use this skill when the user wants to run or prepare a CCE fault drill / chaos experiment / AZ outage drill. Triggers: CCE, 故障演练, 混沌演练, AZ断电, 演练, 故障注入, chaos experiment, AZ power outage, fault drill, chaos drill.
openclaw skills install @huaweicloudskill/huawei-cloud-cce-chaos-experimentEnd-to-end CCE AZ power outage chaos experiment in three phases:
Phase 1: Prepare Phase 2: Execute Phase 3: Log Analysis
───────────── ────────────── ──────────────────
Discover clusters → Pre-check (nodes Ready) → Load experiment context
Select cluster Shutdown (BatchStop) Discover dependencies
Discover node AZs Monitor (Node/Pod) Collect logs (LTS + Pod)
Select target AZ Wait (duration) Analyze error patterns
Validate compatibility Rollback (BatchStart) Generate analysis report
Generate + deploy config Verify + report
Scenario: Shut down all CCE nodes in one AZ → Kubernetes marks nodes NotReady → Pods evicted and rescheduled to other AZs → start nodes back → verify recovery. Tests cluster HA under AZ-level failure.
HW_ACCESS_KEY, HW_SECRET_KEY, HW_REGION_NAME (optional, defaults to cn-north-4 if unset)Run environment check:
bash scripts/check_env.sh
If kubectl is missing, check_env.sh auto-installs it. Manual install:
ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -fsSL "https://dl.k8s.io/release/v1.28.0/bin/linux/${ARCH}/kubectl" -o /tmp/kubectl_download
sleep 1 && mv /tmp/kubectl_download /usr/local/bin/kubectl && chmod 755 /usr/local/bin/kubectl
Generates experiment configuration. Does NOT start the experiment.
bash scripts/check_env.sh
Check HW_REGION_NAME environment variable:
cn-north-4.echo $HW_REGION_NAME
Region is used in all subsequent --region parameters and hcloud --cli-region calls.
python3 scripts/discover_cce.py --region <region>
Output: JSON with all clusters (ID, name, status, flavor, version).
Present discovered clusters in a table. Ask the user to choose — never auto-select.
Available clusters can be selected directly.Hibernation clusters must be awakened first.# Obtain kubeconfig
hcloud CCE CreateKubernetesClusterCert --cluster_id=<id> --cli-region=<region> --duration=30 --cli-output=json > /root/.kube/config
# Discover node AZ distribution
KUBECONFIG=/root/.kube/config kubectl get nodes --show-labels
Present node-to-AZ mapping to the user.
Present AZ distribution. Ask the user to choose the target AZ.
Create the experiment directory before running discovery and validation, so all artifacts (discovery.json, validation.json, experiment.json, logs, reports) live inside one directory.
EXP_DIR="./experiments/$(date -u +%Y%m%d-%H%M%S)-cce-az-power-$(echo <az> | tr -d '-')"
mkdir -p "$EXP_DIR"
echo "$EXP_DIR"
All subsequent steps use $EXP_DIR as the output path.
KUBECONFIG=/root/.kube/config python3 scripts/discover_cce.py \
--region <region> --cluster-id <id> --az <az> \
--include-pods --output "$EXP_DIR/discovery.json"
KUBECONFIG=/root/.kube/config python3 scripts/validate_targets.py \
--discovery-file "$EXP_DIR/discovery.json" \
--cluster-id <id> --az <az> --region <region> \
--output "$EXP_DIR/validation.json"
Validation rules (see references/cce-validation-rules.md):
Decision: errors → stop and fix. warnings only → proceed to generate.
# Generate experiment configuration (use --experiment-dir to reuse $EXP_DIR)
python3 scripts/generate_experiment.py \
--discovery-file "$EXP_DIR/discovery.json" --validation-file "$EXP_DIR/validation.json" \
--cluster-id <id> --az <az> --region <region> \
--duration 300 --experiment-dir "$EXP_DIR"
# Deploy experiment (local mode, generates emergency rollback script)
bash scripts/deploy_experiment.sh --experiment-dir "$EXP_DIR/"
Experiment is prepared but NOT started. Proceed to Phase 2 when ready.
references/cce-az-power-scenario.md — scenario definition, fault impact chain, API mappingreferences/cce-validation-rules.md — validation rules with rationale and decision matrixreferences/experiment-template-guide.md — experiment.json schema with full examplePerforms actual node shutdown. Target CCE nodes will be stopped.
bash scripts/check_env.sh
# Normal execution
python3 scripts/execute_experiment.py --experiment-dir "$EXP_DIR/"
# Dry Run (simulate without API calls)
python3 scripts/execute_experiment.py --experiment-dir "$EXP_DIR/" --dry-run
# Auto-rollback on shutdown failure
python3 scripts/execute_experiment.py --experiment-dir "$EXP_DIR/" --auto-rollback
| Phase | Operation | Description |
|---|---|---|
| 1. Pre-check | kubectl get nodes | All target nodes must be Ready |
| 2. Shutdown | hcloud ECS BatchStopServers | Execute fault injection |
| 3. Monitor | Poll Node + Pod status | Wait for NotReady + Pod rescheduling |
| 4. Wait | sleep(duration) | Experiment duration |
| 5. Rollback | hcloud ECS BatchStartServers | Start all target nodes |
| 6. Verify | Poll Node + Pod status | Confirm Ready + Running + replicas met |
Produces execution-log.json with complete timeline.
python3 scripts/generate_report.py --log-file "$EXP_DIR/execution-log.json"
python3 scripts/rollback_experiment.py --experiment-dir "$EXP_DIR/"
# Or specify instance IDs directly
python3 scripts/rollback_experiment.py --ids "id1,id2" --region <region>
references/execution-workflow.md — 6-phase details, execution-log.json format, safety mechanismsreferences/monitoring-guide.md — Node/Pod state machines, polling mechanismNode: Ready ──关机──→ NotReady ──开机──→ Ready
Pod: Running → Terminating → Pending → Running (on new node in other AZ)
Collects and analyzes application/CCE logs within the experiment time window.
bash scripts/check_env.sh
python3 scripts/analyze_logs.py --experiment-dir "$EXP_DIR/" \
--output-dir "$EXP_DIR/log-analysis"
Auto-detects mode:
execution-log.json exists → post-hoc analysis (事后分析)experiment.json exists → real-time monitoring (实时监控)# Extract affected node names from discovery.json or use known nodes
python3 scripts/discover_dependencies.py \
--nodes "192.168.0.49,..." \
--output "$EXP_DIR/log-analysis/dependencies.json"
Scans Service/Ingress/ConfigMap references to identify directly impacted services.
Note: discover_dependencies.py queries the cluster directly via kubectl, not from discovery.json. Pass the affected node names with --nodes.
Tip: Step 3.2 (
analyze_logs.py) already callsdiscover_dependencies.pyinternally, so this step is optional — use it only for standalone dependency discovery.
python3 scripts/collect_logs.py --experiment-dir "$EXP_DIR/" \
--output-dir "$EXP_DIR/log-analysis/collected_logs"
Collects within experiment time window:
kubectl logs --since-timehcloud LTS ListLogskubectl get eventspython3 scripts/generate_analysis_report.py \
--analysis "$EXP_DIR/log-analysis/analysis-result.json" \
--context "$EXP_DIR/log-analysis/dependencies.json" \
--output "$EXP_DIR/log-analysis/analysis-report.md"
Output: Markdown report with:
references/analysis-workflow.md — 6-step analysis workflow, error pattern classification tablereferences/managed-service-logs.md — kubectl logs, LTS, CES commands, best practices| Pattern | Keywords | Severity | Description |
|---|---|---|---|
| ERROR | error | High | Application error |
| Exception | exception, stacktrace | High | Exception stack trace |
| Connection refused | connection refused | High | Service unavailable |
| Timeout | timeout, deadline exceeded | Medium | Timeout |
| 5xx HTTP | 500, 502, 503, 504 | High | Server error |
| Reconnect | reconnect, reconnected | Low | Reconnected successfully |
| Degraded | degrad | Medium | Degraded mode |
| Recovered | recover, restored | Low | Recovered |
| Script | Phase | Purpose |
|---|---|---|
scripts/check_env.sh | All | Environment verification (shared) |
scripts/discover_cce.py | 1 | Discover CCE clusters and AZ nodes |
scripts/validate_targets.py | 1 | Validate cross-AZ capacity and PDB |
scripts/generate_experiment.py | 1 | Generate experiment config |
scripts/deploy_experiment.sh | 1 | Deploy experiment (local mode, generate rollback script) |
scripts/execute_experiment.py | 2 | Main execution (6 phases) |
scripts/monitor_resources.py | 2 | Monitor Node/Pod status (called by execute) |
scripts/rollback_experiment.py | 2 | Emergency rollback |
scripts/generate_report.py | 2 | Generate Markdown execution report |
scripts/analyze_logs.py | 3 | Main analysis orchestration |
scripts/discover_dependencies.py | 3 | Discover CCE service dependencies |
scripts/collect_logs.py | 3 | Collect LTS + CCE Pod logs |
scripts/generate_analysis_report.py | 3 | Generate Markdown analysis report |
Interactive cluster and AZ selection — the agent must present discovered clusters and AZ distributions to the user and let the user choose. Never auto-select.
Validate before generating — cross-AZ capacity and PDB constraints are checked before any files are produced.
PDB warnings are informational for AZ outage — PDBs govern voluntary pod eviction, not direct node shutdown via ECS API. PDB violations are warnings, not errors.
Dry Run and auto-rollback — --dry-run simulates without API calls; --auto-rollback triggers rollback on shutdown failure.
Local deployment mode — deploy_experiment.sh generates the emergency rollback script locally; the full drill flow is executed by the built-in scripts/execute_experiment.py (no external COC dependency).
AZ-scoped targeting — targets all CCE nodes in a specified AZ, simulating real AZ power outage.
Default duration 5 minutes (300s) — sufficient for observing pod rescheduling while limiting blast radius.
Two analysis modes — post-hoc (from execution-log.json) and real-time (from experiment.json during experiment).
All artifacts in one experiment directory — $EXP_DIR is created before discovery; discovery.json, validation.json, experiment.json, execution-log.json, and log-analysis/ all live inside it.
discover_cce.py — invalid kubeconfig detection: hcloud CCE DownloadClusterConfig returns error with exit code 0 on unsupported versions. Fix: validate config contains "apiVersion" or "clusters"; use CreateKubernetesClusterCert as fallback.
discover_cce.py — get_nodes_by_az return value: Returned [] instead of [], [] on kubectl failure, causing ValueError. Fix: return [], [].
validate_targets.py — CPU millicore parsing: float("1930m") caused ValueError. Fix: parse_cpu() / parse_memory() helpers handle m, Ki, Mi, Gi suffixes.
discover_cce.py — wrong ECS instance ID: Extracted from spec.providerID which contains CCE-internal node ID, not ECS server ID. Fix: get_ecs_instance_map(region) queries hcloud ECS ListServersDetails and builds private-IP → ECS-ID map.
deploy_experiment.sh + execute_experiment.py + rollback_experiment.py — unsupported --body parameter: hcloud KooCLI does not support --body for ECS APIs. Fix: use native parameter format --os-stop.servers.1.id=xxx --os-stop.type=SOFT and --os-start.servers.N.id=xxx.