Install
openclaw skills install @huaweiclouddev/huawei-cloud-mrs-host-fault-diagnoseHuawei Cloud MRS cluster fault diagnosis skill. Diagnoses service faults, instance faults, and host faults through progressive root cause localization: quick log scan first, host troubleshooting when host issues are found, detailed investigation when no conclusion is reached. Driven by the built-in LakeWatch API client and the per-component knowledge base under components/. No commands outside the knowledge base are fabricated. Applicable to MRS fault diagnosis and root cause localization scenarios where a service name or node name is provided. Trigger words: "故障诊断", "故障定位", "fault diagnosis", "fault diagnose", "MRS故障", "服务故障", "实例故障", "主机故障", "集群排查", "集群诊断", "启动失败", "停止异常", "KrbServer故障", "DBService故障", "fault troubleshooting"
openclaw skills install @huaweiclouddev/huawei-cloud-mrs-host-fault-diagnoseThis skill diagnoses Huawei Cloud MRS (MapReduce Service) cluster faults. Given a service name and/or node name, it progressively localizes the root cause: quick log scan first, host troubleshooting when host issues are found, detailed investigation when no conclusion is reached.
Architecture: Caller (Agent) -> lakewatch_api_client.py (Python, scripts/) -> LakeWatch API -> MRS cluster (node resource data, logs, MRS Manager proxy); per-component knowledge base (components/<service_name>.md) drives the diagnosis flow; three fault layers (host -> instance -> service) with propagation chain tracing.
Note on language: This SKILL.md and the documents under
references/are written in English per the repository spec. The knowledge base documents underfault_layer/,scenarios/,components/, andpropagation.mdare also in English. Commands and code blocks are English throughout.
Applicable Scenarios:
Typical Use Cases:
Important constraints:
- Read-only: This skill only runs information-gathering commands (view logs, query status, collect resource data). It MUST NOT run any start/stop, modify, or delete operations.
- User confirmation for repair: The skill only provides executable repair suggestions; it MUST NOT directly execute any repair operation. All repair actions require user confirmation.
- Strict execution: Diagnose strictly according to the knowledge base content under this skill directory. Fabricating diagnostic commands outside the knowledge base is prohibited.
pyyaml (YAML parsing), cryptography (Windows AES password encryption only)cryptography dependency)python3 --version (Linux) / python --version (Windows)This skill does NOT require KooCLI (
hcloud). It calls the LakeWatch API throughscripts/lakewatch_api_client.py. For the LakeWatch client setup, see CLI Installation Guide.
--encrypt-password and stored in scripts/lakewatch_api_config.yaml (auth.encrypted_password). Never store the plaintext password.--encrypt-password flow%TEMP%\lakewatch_token\, Linux: /tmp/lakewatch_token/)scripts/lakewatch_api_config.yaml server.host/port)This skill references the per-alarm diagnosis knowledge base from the huawei-cloud-mrs-host-alarm-diagnose skill (sibling directory under skills/bigdata/mrs/). When the fault diagnosis flow encounters a known alarm (12006/12007/25000/25500/27001), it loads the corresponding document from ../huawei-cloud-mrs-host-alarm-diagnose/alarms/<alarm_id>.md.
This skill uses the LakeWatch API client instead of KooCLI. The unified command format is:
# Linux
python3 <skill_dir>/scripts/lakewatch_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2'
# Windows
python <skill_dir>/scripts/lakewatch_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2'
| Element | Rule | Example |
|---|---|---|
python3 / python | Linux uses python3, Windows uses python | python3 lakewatch_api_client.py |
-a, --api | API name to call (defined in lakewatch_api_config.yaml) | -a collect_alarm_node_res_data |
-p, --param | API parameter in key=value form, repeatable | -p 'cluster_id=xxx' |
| Quoting | Every -p value MUST be wrapped in single quotes to prevent shell parsing of [] {} | () | -p 'keywords=["ERROR"]' |
Windows (PowerShell) quote rule: every " inside a value must be replaced with """ (including " inside [] and {}), otherwise the server returns {"message":"Unknown exception","success":false,"code":"500"}:
# Correct on Windows
-p 'keywords=["""ERROR"""]'
-p 'env={"""PID""":"""123"""}'
# Wrong on Windows (will fail)
-p 'keywords=["ERROR"]'
Linux (bash) quote rule: keep " as-is inside the value, wrap the whole value in single quotes:
# Correct on Linux
-p 'keywords=["ERROR","Exception"]'
-p 'env={"PID":"123"}'
For the full API catalog, parameters, and the token/encryption mechanism, see LakeWatch API Client.
Extract fault information from the user input and determine the diagnosis entry:
| User Description | Entry | Step 1 Action |
|---|---|---|
Has service_name, no node_name (e.g. "KrbServer出问题了") | Service fault | Check all instance statuses, find faulty instances |
Has service_name + node_name (e.g. "8-5-225-6上的KrbServer挂了") | Instance fault | Directly check that instance |
Has node_name, no service_name (e.g. "8-5-225-6出问题了") | Host fault | Check host status, then check instances on the host |
Load components/<service_name>.md for component config. Query OMS primary/standby nodes, check process on each node:
python3 lakewatch_api_client.py -a query-management-node-info \
-p 'cluster_id=<cluster_id>'
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=process-basic-info' \
-p 'env={"process_name":"<process_name>"}' \
-p 'node_name=<node_name>'
Decision:
| Result | Next Step |
|---|---|
| All node processes normal | Step 4 detailed investigation |
| Some node processes missing | Step 3 quick log scan (for faulty nodes) |
| API call failed (node unreachable) | Step 4 host troubleshooting |
Load components/<service_name>.md. Directly check process on that node:
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=process-basic-info' \
-p 'env={"process_name":"<process_name>"}' \
-p 'node_name=<node_name>'
Decision:
| Result | Next Step |
|---|---|
| Process normal | Step 4 detailed investigation |
| Process missing | Step 3 quick log scan |
| API call failed (node unreachable) | Step 4 host troubleshooting |
Query OMS primary/standby nodes, query node IP, ping the faulty node from OMS active node:
python3 lakewatch_api_client.py -a query-management-node-info \
-p 'cluster_id=<cluster_id>'
python3 lakewatch_api_client.py -a query-node-ip \
-p 'cluster_id=<cluster_id>' \
-p 'node_name=<node_name>'
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=ping-check' \
-p 'env={"TARGET_IP":"<target_ip>"}' \
-p 'node_name=<oms_active_node>'
Decision:
| Result | Next Step |
|---|---|
| Ping failed | Step 4 host troubleshooting (network/hardware) |
| Ping succeeded | Check all component processes on the host, find faulty instances -> Step 3 quick log scan |
For the faulty node, quickly scan three layers of logs (Controller -> NodeAgent -> component), looking for clear ERROR:
# Controller log
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=<alarm_time>' \
-p 'log_directory=/var/log/Bigdata/controller' \
-p 'log_file_name=exe.log*' \
-p 'keywords=["<service_name>","ERROR","fail","timeout","Exception"]' \
-p 'log_type=local' \
-p 'node_name=<oms_active_node>'
# NodeAgent script log
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=<alarm_time>' \
-p 'log_directory=/var/log/Bigdata/nodeagent/scriptlog' \
-p 'log_file_name=*.log*' \
-p 'keywords=["<service_name>","ERROR","fail","exit"]' \
-p 'log_type=local' \
-p 'node_name=<node_name>'
If service_name is known, also check the component's own log (path from components/<service_name>.md):
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=<alarm_time>' \
-p 'log_directory=<log_directory>' \
-p 'log_file_name=<log_file_name>' \
-p 'keywords=["ERROR","Exception","FATAL","fail","OOM"]' \
-p 'log_type=local' \
-p 'node_name=<node_name>'
Decision:
| Log Result | Next Step |
|---|---|
| Clear ERROR (e.g. OOM/permission/port conflict/config missing) | Output root cause |
| Log shows node unreachable / Agent timeout | Step 4 host troubleshooting |
| Multiple faulty nodes on same host | Step 4 host troubleshooting |
| No clear conclusion | Step 4 detailed investigation |
When the quick log scan yields no conclusion, collect complete data:
Load Propagation Chain to trace the root cause propagation path and impact scope.
## Diagnosis Result
| Item | Content |
|------|---------|
| Diagnosis time | [time] |
| Cluster ID | [cluster_id] |
| Faulty component | [service_name] |
| Faulty node | [node_name] |
### Diagnosis Process
| Step | Result |
|------|--------|
| Instance status | [which nodes normal/abnormal] |
| Quick log scan | [found/not found clear ERROR] |
| Host troubleshooting | [normal/abnormal: ...] |
| Detailed investigation | [process/port/HA/resource results] |
### Propagation Path
[root cause] -> [propagation] -> [symptom] (single-layer root cause if no propagation)
### Root Cause Analysis
**Root cause layer**: [host/instance/service]
**Root cause type**: [specific reason]
### Repair Suggestion
| Priority | Operation | Description | Needs user confirmation |
|----------|-----------|-------------|-------------------------|
| 1 | [operation] | [description] | Yes |
python3 lakewatch_api_client.py -a query-management-node-info \
-p 'cluster_id=<cluster_id>'
python3 lakewatch_api_client.py -a query-node-ip \
-p 'cluster_id=<cluster_id>' \
-p 'node_name=<node_name>'
# Process basic info
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=process-basic-info' \
-p 'env={"process_name":"<process_name>"}' \
-p 'node_name=<node_name>'
# Port check
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=port-check' \
-p 'env={"PORT":"<port>"}' \
-p 'node_name=<node_name>'
# HA resource status
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=ha-resource-status' \
-p 'node_name=<node_name>'
# Disk space / Memory / CPU load
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=disk-space' \
-p 'node_name=<node_name>'
Supported strategy_name values include: system-load, memory-usage, disk-space, disk-io, network-io, file-handle, port-check, high-cpu-processes, high-memory-process, zombie-process, dns-check, network-connectivity-test, process-basic-info, process-file-descriptor, jstack-thread-dump, disk-health-check, disk-smart-info, ha-resource-status, omm-process-tree, and more. See LakeWatch API Client for the full list.
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=<alarm_time>' \
-p 'log_directory=<log_directory>' \
-p 'log_file_name=<log_file_name>' \
-p 'keywords=["ERROR","Exception"]' \
-p 'log_type=local'
When the log time format is non-standard ISO (e.g. [2026-07-07 20:54:25,171]), pass time_pattern:
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=2026/07/07 20:54:00 GMT+08:00' \
-p 'log_directory=/var/log/Bigdata/omm/oms/pms' \
-p 'log_file_name=pms*.log' \
-p 'keywords=["ERROR","Exception"]' \
-p 'log_type=local' \
-p 'time_pattern=^\[([0-9]{4})-([0-9]{2})-([0-9]{2}) ([0-9]{2}):([0-9]{2}):([0-9]{2})||ymdHMS'
# Query cluster services
python3 lakewatch_api_client.py -a access_manager_get \
-p 'cluster_id=<cluster_id>' \
-p 'target_url=api/v2/clusters/<cluster_id>/services'
# Query host processes
python3 lakewatch_api_client.py -a access_manager_get \
-p 'cluster_id=<cluster_id>' \
-p 'target_url=api/v2/clusters/<cluster_id>/hosts/<node_name>/processes'
# Query active alarms
python3 lakewatch_api_client.py -a access_manager_get \
-p 'cluster_id=<cluster_id>' \
-p 'target_url=api/v2/clusters/<cluster_id>/alarms'
target_urlMUST NOT start with/. The proxy requires Agent >= 1.0.5 and reported OMS node info. Only GET is supported currently.
| Parameter | Required/Optional | Description | Default |
|---|---|---|---|
cluster_id | Required | MRS cluster ID | N/A |
service_name | Conditionally required | Faulty component (required for service/instance fault entry) | N/A |
node_name | Conditionally required | Faulty node (required for instance/host fault entry) | N/A |
alarm_time | Optional | Fault occurrence time, format yyyy/MM/dd HH:mm:ss GMT+X:XX | Current time |
strategy_name | Required by collect_alarm_node_res_data | Resource collection strategy | N/A |
log_directory | Required by collect_alarm_log_data | Log directory, must be under /var/log/ | N/A |
log_file_name | Required by collect_alarm_log_data | Log file name, no path separators | N/A |
keywords | Required by collect_alarm_log_data | Log keyword filter, JSON array | N/A |
log_type | Required by collect_alarm_log_data | local or hdfs | N/A |
time_pattern | Optional | Non-standard log time regex, format regex||format | N/A |
target_url | Required by access_manager_get | MRS Manager API path, must NOT start with / | N/A |
The diagnosis report is output in Markdown, containing:
See the template in the Workflow -> Step 6 section.
See Verification Method for the installation, configuration, and function verification steps.
<cluster_id>, <alarm_time>, <node_name>, <target_ip>, <process_name>, etc. with actual user-provided values; never hardcode them.-p values in single quotes; on Windows PowerShell, escape " as """ to avoid code:500 errors.alarm_time must follow yyyy/MM/dd HH:mm:ss GMT+X:XX; for non-standard log time formats, pass time_pattern.| Document | Description |
|---|---|
| CLI Installation Guide | Python dependencies and LakeWatch client setup |
| IAM Policies | LakeWatch/MRS Manager access model and required roles |
| Verification Method | Installation, configuration, and function verification |
| Acceptance Criteria | Pass/fail criteria for skill testing |
| Fault Diagnosis Workflow | Progressive fault diagnosis workflow design |
| LakeWatch API Client | Full API catalog, parameters, token and encryption mechanism |
| Related Commands | Common LakeWatch API commands quick reference |
| huawei-cloud-mrs-host-alarm-diagnose (sibling skill) | Dependency: per-alarm diagnosis knowledge base (../huawei-cloud-mrs-host-alarm-diagnose/alarms/<alarm_id>.md). See Prerequisites section 4 for details. |
| Data Collection | Complete data collection flow (Step 4) |
| Host Fault Diagnosis | Host layer diagnosis |
| Instance Fault Diagnosis | Instance layer diagnosis (includes scenario identification) |
| Service Fault Diagnosis | Service layer diagnosis |
| Propagation Chain | Root cause propagation path tracing |
| Common Scenario | 6-phase common diagnosis framework for all scenarios |
scenarios/<scenario>.md | Scenario-specific checks (install/start/stop/uninstall/reinstall/reinstall_host/scale_out/scale_in) |
components/<service_name>.md | Per-component configuration (process, port, log path, etc.) |
components/_template.md | Template for new component configuration |
--encrypt-password and stored in lakewatch_api_config.yaml. Repair steps are suggestions only.hcloud; it calls the LakeWatch API through lakewatch_api_client.py. Do not mix in hcloud commands.access_manager_get proxy only supports GET requests (PUT is not yet available on the Agent side); collect_alarm_log_data requires log_directory to be under /var/log/; some strategy_name values require extra env parameters.../huawei-cloud-mrs-host-alarm-diagnose/alarms/12006.md, 12007.md). See Prerequisites section 4 for the dependency declaration and handling rules. If the alarm skill is not installed, inform the user and proceed with the generic fault diagnosis flow.