Install
openclaw skills install @huaweicloudskill/huawei-cloud-ces-aom-capacity-assessmentHuawei Cloud capacity assessment. Assesses capacity for instances across 21 services (NAT/RDS/ECS/ELB/EIP/CSS/DDS/DWS/DCS/DMS/TaurusDB/GeminiDB/DC/ER/EVS/EFS Turbo/APIG/DRS/VPCEP/Bandwidth/CCE) by comparing peak/valley fluctuations of monitoring data between entry instances and target instances, calculating pressure coefficients to predict festival traffic ceilings, and giving scale-out recommendations. Data collection uses hcloud (KooCLI) CES commands (CCE and other cloud-native metrics go through AOM). Supports four modes: single-instance historical, full historical, single-instance T-24, full T-24. Trigger words: 容量评估、扩容建议、节日保障、压力系数、峰谷值倍数、预计节日上限、容量预测、capacity assessment、要不要扩容、ECS/RDS/CCE 够不够用.
openclaw skills install @huaweicloudskill/huawei-cloud-ces-aom-capacity-assessmentAssesses instances in the Excel capacity template: collects monitoring data (hcloud CES; CCE and other cloud-native metrics via hcloud AOM), computes peak/valley multiples, pressure coefficients, and expected festival ceilings, decides whether business peaks (festival events) need scale-out, and writes results back to Excel.
<SKILL_DIR>/scripts/capacity_cli.py # Unified CLI entry point
<SKILL_DIR>/scripts/requirements.txt # Dependency: openpyxl
<SKILL_DIR>/references/metrics.json # 238-metric registry (Chinese name → CES/AOM metric_name + namespace + dimensions + ceiling + unit)
<SKILL_DIR>/references/troubleshooting-dns.md # DNS troubleshooting when hcloud cannot connect
<SKILL_DIR>/references/region-map.md # Chinese region name ↔ region code mapping
<SKILL_DIR>/references/cli-installation-guide.md # hcloud installation/config (required for review)
<SKILL_DIR>/references/iam-policies.md # IAM permissions (CES/AOM read-only, required for review)
<SKILL_DIR>/references/verification-method.md # Verification method (required for review)
<SKILL_DIR>/references/acceptance-criteria.md # Acceptance criteria (required for review)
<SKILL_DIR>/templates/capacity_assessment_template.xlsx # Template (the Excel template users take and fill in)
First-time setup: pip install -r <SKILL_DIR>/scripts/requirements.txt
| Mode | Command | Time range |
|---|---|---|
| Single-instance historical | assess-row --mode historical --start ... --end ... | user-specified |
| Full historical | assess-all --mode historical | user-specified |
| Single-instance T-24 | assess-row --mode t24 | last 24h (automatic) |
| Full T-24 | assess-all --mode t24 | last 24h (automatic) |
Time parameter format: YYYY-MM-DD HH:MM:SS (day windows are split by day automatically, peak=max, valley=min).
When users ask "what is the template / how do I fill it / which metrics are supported", run:
python3 <SKILL_DIR>/scripts/capacity_cli.py info
This command outputs in one shot: Excel template instructions (including the template file location) + the list of supported instance type abbreviations (to prevent users filling Chinese full names that fail recognition) + the full supported-metric list grouped by 云服务/云服务维度/关键指标.
references/region-map.md (Chinese region name ↔ region code mapping, consistent with resolve_region in config.py).python3 <SKILL_DIR>/scripts/capacity_cli.py read-excel <capacity_assessment_template.xlsx>
Returns header columns (with base names), data rows, and required-column validation results (region/实例类型/实例ID/关键指标/入口实例ID).
Single instance: run on the specified row (collects both target and entry instances, computes, and returns cells to write):
python3 <SKILL_DIR>/scripts/capacity_cli.py assess-row <xlsx> --row 2 --mode historical \
--start "2026-07-01 00:00:00" --end "2026-07-14 23:59:59"
# or T-24:
python3 <SKILL_DIR>/scripts/capacity_cli.py assess-row <xlsx> --row 2 --mode t24
Full (time-consuming, background execution + polling):
python3 <SKILL_DIR>/scripts/capacity_cli.py assess-all <xlsx> --mode historical \
--start "..." --end "..." --background
# returns {"background": true, "pid": ..., "progress": "<xlsx>.assess.log", ...}
python3 <SKILL_DIR>/scripts/capacity_cli.py assess-status --progress <xlsx>.assess.log
The background task processes rows one by one and writes everything back to Excel when done (no backup files), finally returning
skipped_rows (metric mismatch / unsupported service) and collection_failures (collection failed) details.
assess-row outputs updates that can be written back via write-excel, or let the assess-all background task write back uniformly.os.replace); on save failure the .tmp is cleaned automatically and the original file is untouched.7/1峰值) are auto-inserted after the 【入口实例ID】 column, sorted by time.python3 <SKILL_DIR>/scripts/capacity_cli.py write-excel <xlsx> --updates updates.json
Automatic cleanup of temp files and history (built into the script, no manual work):
assess-all finishes in the foreground, <xlsx>.assess.log and <xlsx>.assess.log.out are deleted automatically.assess-all --background is auto-deleted after assess-status reads done..bak-* backup; on save failure .tmp is cleaned. The working directory should always contain only the <xlsx> itself — no history/temp residue.python3 <SKILL_DIR>/scripts/capacity_cli.py cleanup <xlsx>
(default --keep 0, deletes all .bak-*/.assess.log/.assess.log.out/.tmp)After every assessment, must report the following to the user:
collection_failures is
该时间窗口内无监控数据 or 返回中未找到该指标, you must explain to the user —
the filled-in key indicator itself is correct and supported by Huawei Cloud monitoring; the failure usually means
the cloud service instance does not report data to the monitoring service (instance stopped / monitoring not enabled / no data points in the window).
Suggest the user self-check: whether the instance is running, whether cloud monitoring reporting is enabled, and whether data really exists in the window.
Do not explain this as "the metric was written wrong", and do not show customers technical fields like metric_name.A correct key indicator ≠ data exists (important): the key indicators in the template are all official metrics supported by Huawei Cloud monitoring — not a mistake. But whether an instance has monitoring data depends on whether the cloud service reports to the monitoring service: when an instance is stopped / not connected / has no data points, the query returns "无监控数据" or "未找到该指标". Such failures are not a wrong metric nor a tool problem; guide the user to check the cloud service's reporting status in the report.
references/metrics.json: 238 metrics across 21 services,
keyed by Chinese metric name (with service prefix, e.g. "ECS CPU使用率"); prefix-less aliases (e.g. "CPU利用率" in the template) are also accepted.
ECS/EVS/Bandwidth/EIP/ELB/DC/DCS/RDS/VPCEP/CCE etc. have been verified as collectible; the rest are organized per Huawei Cloud
monitoring metric lists, dimensions follow the official lists, not yet verified one by one; correct them with real collections if anomalies occur.
3 items are still unregistered (official metric_name not public):
TaurusDB 数据盘使用率 (note: "TaurusDB 磁盘使用率" IS registered, they are different),
GeminiDB Redis 节点带宽利用率 ("节点入/出带宽利用率" is registered, only the direction-less aggregate is not),
VPCEP 终端节点每秒新建连接数.
Rows filling these metrics are marked skipped (metric_mismatch), and the service's supported-metric list is returned for the user to correct.
CCE metric note (8 supported, collected via AOM): CCE monitoring metrics have no official public metric_name in CES;
the skill collects them via hcloud AOM ListSample (entries in metrics.json carry "backend": "aom"), no extra Agent needed.
All 8 AOM metric_names verified on an account with a real CCE cluster (pre-rename names, as used in ListSample input):
cpuUsage、CCE内存利用率=memUsedRate、CCE磁盘使用率=diskUsedRate(PAAS.AGGR)cpuUsage、CCE节点磁盘使用率=diskUsedRate(PAAS.NODE)cpuUsage、CCE POD物理内存使用率=memUsage(PAAS.CONTAINER)
How to fill:Each metric in references/metrics.json has ceiling (resource ceiling) and unit fields, used as the default ceiling.
The "资源上限" column in the template is optional (assess_row prefers the Excel user-entered value, falling back to the registry only when absent):
ECS metric note (already in info output, user-facing): current ECS assessment metrics come from physical-machine (host) level collection,
less accurate than in-instance collection; install the Cloud Eye Agent on that ECS for more accurate metrics.
CES metric collection is invoked by the script as hcloud CES BatchListMetricData --cli-region=<region>,
CCE and other cloud-native metrics (backend=aom) as hcloud AOM ListSample --cli-region=<region>;
<region> is the region code (parsed from the --region parameter of capacity_cli.py, e.g. cn-north-4).
If network/timeout errors occur:
python3 <SKILL_DIR>/scripts/capacity_cli.py smoke --region <region-code>
(smoke probes both CES and AOM backends and reports which one failed)getent hosts ces.<region-code>.myhuaweicloud.com
getent hosts aom.<region-code>.myhuaweicloud.com
100.125.x.x → look up the public IP with a public DNS and write it into /etc/hosts
(see references/troubleshooting-dns.md, includes verified steps and real public-IP tests; AOM is diagnosed the same as CES, just replace the domain with AOM.).Also note: even with DNS OK, some regions may report
The IAM user is forbidden in the currently selected regiondue to missing IAM permissions; such errors go into thecollection_failuresdetails (the row is written with an "无法计算/无法预测" placeholder), reported by the model to the user as-is; they are not script issues to fix.
collect: single-metric collection, --t24 or --start/--end, --region (code).
Debug peak/valley of a specific instance metric directly.calculate: pure computation, input JSON {"mode":"historical","target":{peak,valley,unit}, "entrances":[{peak,valley,unit}],"ceiling":100,"growth":1.5}, outputs all metrics.保障重点实例_容量管理模板