Install
openclaw skills install @huaweicloudskill/huawei-cloud-ecs-shutdown-experimentFull lifecycle ECS shutdown fault injection experiment for Huawei Cloud chaos engineering: prepare → execute → analyze. Phase 1 discovers target ECS instances, validates compatibility, generates experiment configuration. Phase 2 executes BatchStopServers shutdown, polls status, holds for duration, rolls back via BatchStartServers, verifies recovery, generates execution report. Phase 3 collects CES monitoring metrics and LTS application logs, analyzes error patterns, generates analysis report. Key safety: --dry-run, --auto-rollback, mandatory --yes confirmation, independent emergency rollback script. Triggers: ECS故障演练, ECS关机实验, chaos engineering, 故障注入, ECS shutdown experiment, 演练准备, 执行实验, 运行实验, 启动演练, 执行 ECS 关机故障演练, 分析应用日志, 查看应用表现, 应用日志分析, 故障影响分析, experiment prepare, experiment execute, analyze app logs, log analysis.
openclaw skills install @huaweicloudskill/huawei-cloud-ecs-shutdown-experiment⚠️ This skill performs real ECS shutdown operations in Phase 2. Ensure target instances can tolerate downtime. Emergency rollback is always available.
Full lifecycle ECS shutdown fault injection experiment for Huawei Cloud chaos engineering, covering three phases:
| Phase | Name | Action | Key Output |
|---|---|---|---|
| 1 | Prepare | Discover, validate, generate config | experiment.json + rollback_experiment.sh |
| 2 | Execute | Shutdown → hold → rollback → verify | execution-log.json + execution-report.md |
| 3 | Analyze | Collect CES metrics + LTS logs, analyze patterns | log-analysis-result.json + log-analysis-report.md |
Scope: standalone ECS instances in any availability zone (non-CCE).
references/cli-installation-guide.md.HW_ACCESS_KEY, HW_SECRET_KEY, HW_REGION_NAME (optional, default cn-north-4).references/iam-policies.md for the full matrix.hcloud LTS ListLogGroups or LTS console).Environment check (run before any phase):
bash scripts/check_env.sh --cli-region <region>
Does NOT start the experiment. Only discovers targets, validates compatibility, and generates configuration files.
python3 scripts/discover_ecs.py --cli-region <region>
python3 scripts/discover_ecs.py --cli-region <region> --az cn-north-4a --status ACTIVE
python3 scripts/discover_ecs.py --cli-region <region> --name-pattern "web-"
Output: JSON with instance ID, name, status, AZ, flavor, charging mode, and tags.
python3 scripts/validate_targets.py --instances-file instances.json --cli-region <region>
Validation rules (see references/ecs-validation-rules.md for details):
Decision: errors → stop and fix; warnings only → review with user before proceeding.
python3 scripts/generate_experiment.py \
--targets "i-xxx,i-yyy" --cli-region <region> --duration 300 --output-dir ./experiments
Output directory: {timestamp}-ecs-shutdown-{slug}/ containing experiment.json + README.md. See references/experiment-template-guide.md for schema.
bash scripts/deploy_experiment.sh --experiment-dir ./experiments/{dir}/ --cli-region <region>
Generates rollback_experiment.sh (emergency rollback script) in the experiment directory.
Experiment is prepared but NOT started. Review experiment.json and README.md, then proceed to Phase 2 when ready.
Real shutdown operation. Target ECS instances are stopped. Ensure
experiment.jsonfrom Phase 1 is available.
# Real execution (--yes is mandatory; default refuses without it)
python3 scripts/execute_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region> --yes
# Dry Run (simulated, no API calls)
python3 scripts/execute_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region> --dry-run
# Auto-rollback when shutdown fails
python3 scripts/execute_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region> --auto-rollback --yes
6-phase execution flow (see references/execution-workflow.md for details):
| Phase | Action | Description |
|---|---|---|
| 1. Pre-check | Query status | All instances must be ACTIVE |
| 2. Shutdown | BatchStopServers | Inject the fault |
| 3. Monitor | Poll status | Wait until all reach SHUTOFF |
| 4. Wait | sleep(duration) | Hold with stop-condition checks |
| 5. Rollback | BatchStartServers | Restore instances |
| 6. Verify | Poll status | Confirm all back to ACTIVE |
Output: execution-log.json with full execution timeline and state changes.
python3 scripts/generate_report.py --log-file ./experiments/xxx/execution-log.json
Report contains: experiment overview, target instances, phase timeline, state changes, conclusions.
If the experiment must be terminated early:
python3 scripts/rollback_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region>
# Or pass instance IDs directly:
python3 scripts/rollback_experiment.py --ids "i-xxx,i-yyy" --cli-region <region>
Collect CES monitoring metrics and LTS application logs to understand how applications respond to the infrastructure failure. Supports post-hoc mode (after experiment) and real-time mode (during experiment).
Auto-detect mode: execution-log.json found → post-hoc; falls back to experiment.json → real-time.
python3 scripts/analyze_logs.py \
--experiment-dir ./experiments/xxx/ \
--cli-region <region> \
--log-group-id <lts-group-id>
# Dry run (verify workflow without API calls)
python3 scripts/analyze_logs.py --experiment-dir ./experiments/xxx/ --dry-run
The pipeline executes steps 3.3–3.6 automatically. See references/analysis-workflow.md for details.
Queries 6 metrics via hcloud CES ShowMetricData (see references/managed-service-logs.md → CES section):
| Namespace | Metric | Label | Requires Agent? |
|---|---|---|---|
SYS.ECS | cpu_util | CPU Utilization | No |
SYS.ECS | network_incoming_bytes_aggregate_rate | Network Incoming Rate | No |
SYS.ECS | network_outgoing_bytes_aggregate_rate | Network Outgoing Rate | No |
SYS.ECS | disk_read_bytes_rate | Disk Read Rate | No |
SYS.ECS | disk_write_bytes_rate | Disk Write Rate | No |
AGT.ECS | mem_usedPercent | Memory Usage | Yes |
Metrics summarized by phase (pre-shutdown / during-shutdown / post-recovery). Query uses period=300 (5-min aggregation for SYS.ECS).
Queries LTS log group within experiment time window. --log-group-id is required (find via hcloud LTS ListLogGroups). Missing group / empty result = no LTS data, NOT "application unaffected" — check ICAgent config.
Local pattern matching (no network calls):
| Pattern | Severity |
|---|---|
| ERROR/Exception/Traceback | high |
| Connection refused | high |
| Timeout | high |
| HTTP 5xx | high |
| Reconnect/retry | medium |
| Degrade/circuit breaker | medium |
| Recovery/restored | info |
python3 scripts/generate_analysis_report.py --result-file ./experiments/xxx/log-analysis-result.json
Report contains: experiment overview, affected instances, LTS log sources, instance state timeline, CES monitoring metrics, error pattern statistics, error event timeline, application behavior assessment, improvement suggestions.
Before each phase's core operation, confirm with the user:
| Phase | What to Confirm | Mandatory? |
|---|---|---|
| 1 (Generate) | Target instance IDs + names, duration, validation warnings | ✅ Yes |
| 2 (Execute) | Experiment name, target count, duration, dry-run flag | ✅ Yes — --yes required for real execution |
| 3 (Analyze) | Experiment directory, detected mode, time window, LTS log group ID | ✅ Yes |
Wait for explicit confirmation ("确认" / "confirm" / "ok") before proceeding.
<region>= region, defaults tocn-north-4(HW_REGION_NAMEoverrides if set). The placeholder may be omitted at runtime — scripts fall back to the default region.
# Environment check (all phases)
bash scripts/check_env.sh --cli-region <region>
# Phase 1: Prepare
python3 scripts/discover_ecs.py --cli-region <region>
python3 scripts/validate_targets.py --instances-file instances.json --cli-region <region>
python3 scripts/generate_experiment.py --targets "i-xxx,i-yyy" --cli-region <region> --duration 300
bash scripts/deploy_experiment.sh --experiment-dir ./experiments/{dir}/ --cli-region <region>
# Phase 2: Execute
python3 scripts/execute_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region> --yes
python3 scripts/generate_report.py --log-file ./experiments/xxx/execution-log.json
python3 scripts/rollback_experiment.py --experiment-dir ./experiments/xxx/ --cli-region <region> # emergency
# Phase 3: Analyze
hcloud LTS ListLogGroups --cli-region=<region> --cli-output=json # find log group
python3 scripts/analyze_logs.py --experiment-dir ./experiments/xxx/ --cli-region <region> --log-group-id <id>
python3 scripts/generate_analysis_report.py --result-file ./experiments/xxx/log-analysis-result.json
All hcloud CLI commands follow this format:
hcloud <Service> <Action> --cli-region=<region> --cli-output=json [--param=value ...]
--cli-region=<region>: always explicit. KooCLI 7.2.x requires --param=value syntax (equals sign); space-separated --cli-region cn-north-4 is rejected with [USE_ERROR].--cli-output=json: all commands, for machine-parseable output.--body):
--os-stop.servers.1.id=<id> --os-stop.type=SOFT--os-start.servers.1.id=<id>ListServersDetails with --limit=1000 + paging; no N+1 queries.HW_ACCESS_KEY / HW_SECRET_KEY, never hardcoded.| Script | Phase | Purpose |
|---|---|---|
scripts/check_env.sh | All | Environment verification |
scripts/discover_ecs.py | 1 | Discover ECS instances |
scripts/validate_targets.py | 1 | Validate target compatibility |
scripts/generate_experiment.py | 1 | Generate experiment config |
scripts/deploy_experiment.sh | 1 | Deploy experiment (local mode) |
scripts/execute_experiment.py | 2 | Main execution script (6-phase flow) |
scripts/monitor_instances.py | 2 | Instance status monitoring |
scripts/rollback_experiment.py | 2 | Emergency rollback |
scripts/generate_report.py | 2 | Generate execution report |
scripts/collect_logs.py | 3 | LTS log collection + ECS endpoint query |
scripts/collect_ces_metrics.py | 3 | CES monitoring metrics collection |
scripts/analyze_logs.py | 3 | Main analysis pipeline orchestrator |
scripts/generate_analysis_report.py | 3 | Generate Markdown analysis report |
| Document | Phase | Description |
|---|---|---|
references/cli-installation-guide.md | All | hcloud KooCLI installation and configuration |
references/iam-policies.md | All | IAM permissions by phase |
references/ecs-validation-rules.md | 1 | Validation rules with rationale |
references/experiment-template-guide.md | 1 | experiment.json schema |
references/execution-workflow.md | 2 | Detailed execution workflow (per-phase I/O, exceptions) |
references/monitoring-guide.md | 2 | Monitoring guide (state machine, polling, CES alarms) |
references/analysis-workflow.md | 3 | Detailed analysis workflow |
references/managed-service-logs.md | 3 | Managed-service log reference (CES, LTS, audit) |
references/verification-method.md | All | Verification methodology (all phases) |
references/acceptance-criteria.md | All | Acceptance criteria (all phases) |
--dry-run simulates the full flow without API calls (Phase 2 & 3)--auto-rollback starts instances back up when shutdown failsrollback_experiment.py is standalone, does not depend on main script stateexecution-log.json records per-phase timestamps and state changes--yes; otherwise refusesos.environStop: ACTIVE → STOPPING → SHUTOFF
Start: SHUTOFF → STARTING → ACTIVE
Abnormal: any → ERROR
SYS.ECS metrics are aggregated at 5-minute granularity (period=300); shorter periods return emptyreferences/managed-service-logs.md → CES sectionAGT.ECS metrics (memory) require telescope/uniagent Agent installed inside the ECSinstance_timeline is unavailable, so recovery-time analysis is skipped