Install
openclaw skills install @huaweicloudskill/huawei-cloud-mrs-host-alarm-diagnoseHuawei Cloud MRS cluster alarm diagnosis skill. Analyzes the root cause of an MRS alarm based on user-provided alarm information (alarm ID, alarm name, alarm details, occurrence time, node IP, related service and logs), then outputs the root cause, repair steps, and verification method. Diagnosis is driven by the built-in LakeWatch API client (lakewatch mode) or the MRS Manager API client (manager mode) and the per-alarm knowledge base under alarms/ (lakewatch) or alarm_manager/ (manager). The API mode is auto-detected by check_api_mode.py. No commands outside the knowledge base are fabricated. Applicable to MRS alarm diagnosis and root cause localization scenarios where an alarm ID is provided. Trigger words: "告警诊断", "告警定位", "alarm diagnosis", "alarm diagnose", "MRS告警", "告警原因", "告警ID", "alarm ID", "root cause"
openclaw skills install @huaweicloudskill/huawei-cloud-mrs-host-alarm-diagnoseThis skill diagnoses Huawei Cloud MRS (MapReduce Service) cluster alarms. Given alarm information (alarm ID, alarm name, occurrence time, cluster ID, node, related service/role), it locates the root cause and outputs repair steps and a verification method.
Architecture: Caller (Agent) → check_api_mode.py (Python, scripts/) determines the API mode → either lakewatch_api_client.py → LakeWatch API → MRS cluster (node resource data, logs, MRS Manager proxy) or manager_api_client.py → MRS Manager REST API (28443). Per-alarm knowledge base (alarms/<alarm_id>.md in lakewatch mode, alarm_manager/<alarm_id>.md in manager mode) drives the diagnosis flow.
Note on language: This SKILL.md, the documents under
references/, and the per-alarm knowledge base underalarms/andalarm_manager/are all written in English per the repository spec. Commands and code blocks are English throughout.
Applicable Scenarios:
Typical Use Cases:
Important constraints:
- Read-only: This skill only runs information-gathering commands (view logs, query status). It MUST NOT run any start/stop, modify, or delete operations.
- User confirmation for repair: The skill only provides executable repair steps; it MUST NOT directly execute any repair operation. All repair actions require user confirmation.
- Strict execution: Diagnose strictly according to the per-alarm knowledge base content. Fabricating diagnostic commands outside the knowledge base is prohibited.
pyyaml (YAML parsing), cryptography (Windows AES password encryption only)cryptography dependency)python3 --version (Linux) / python --version (Windows)This skill does NOT require KooCLI (
hcloud). It calls the LakeWatch API throughscripts/lakewatch_api_client.py(lakewatch mode) or the MRS Manager REST API throughscripts/manager_api_client.py(manager mode). For the client setup, see CLI Installation Guide.
--encrypt-password and stored in scripts/lakewatch_api_config.yaml (auth.encrypted_password). Never store the plaintext password.--encrypt-password flow%TEMP%\lakewatch_token\, Linux: /tmp/lakewatch_token/)Manager mode is enabled when scripts/manager_api_config.yaml exists and auth.encrypted_password is set.
server.host (port default 28443). To obtain it, run grep float_ip /opt/huawei/Bigdata/om-server/OMS/workspace/conf/oms.ini on the OMS node, or ask the cluster administrator.python3 scripts/manager_api_client.py --encrypt-password and stored in scripts/manager_api_config.yaml (auth.encrypted_password). Never store the plaintext password.--encrypt-password flow.aes_key file to be migrated together to decrypt on another machine; Linux SCC ciphertext is not portable across clustersscripts/lakewatch_api_config.yaml server.host/port); the LakeWatch account must have permission to call the MRS Manager proxy and collect node resource/log data on the target clusterThis skill uses the LakeWatch API client instead of KooCLI. The unified command format is:
# Linux
python3 <skill_dir>/scripts/lakewatch_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2'
# Windows
python <skill_dir>/scripts/lakewatch_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2'
| Element | Rule | Example |
|---|---|---|
python3 / python | Linux uses python3, Windows uses python | python3 lakewatch_api_client.py |
-a, --api | API name to call (defined in lakewatch_api_config.yaml) | -a collect_alarm_node_res_data |
-p, --param | API parameter in key=value form, repeatable | -p 'cluster_id=xxx' |
| Quoting | Every -p value MUST be wrapped in single quotes to prevent shell parsing of `[] {} | ()` |
Windows (PowerShell) quote rule: every " inside a value must be replaced with """ (including " inside [] and {}), otherwise the server returns {"message":"Unknown exception","success":false,"code":"500"}:
# Correct on Windows
-p 'keywords=["""ERROR"""]'
-p 'env={"""PID""":"""123"""}'
# Wrong on Windows (will fail)
-p 'keywords=["ERROR"]'
Linux (bash) quote rule: keep " as-is inside the value, wrap the whole value in single quotes:
# Correct on Linux
-p 'keywords=["ERROR","Exception"]'
-p 'env={"PID":"123"}'
For the full API catalog, parameters, and the token/encryption mechanism, see LakeWatch API Client.
When check_api_mode.py reports manager, use manager_api_client.py instead of the LakeWatch client. The unified command format is:
# Linux
python3 <skill_dir>/scripts/manager_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2' --json
# Windows
python <skill_dir>/scripts/manager_api_client.py -a <api_name> -p 'key1=value1' -p 'key2=value2' --json
| Element | Rule | Example |
|---|---|---|
python3 / python | Linux uses python3, Windows uses python | python3 manager_api_client.py |
-a, --api | API name to call (defined in manager_api_apis/) | -a get_instances |
-p, --param | API parameter in key=value form, repeatable | -p 'service_name=DBService' |
--json | JSON formatted output | --json |
--auth | Auth mode: basic (default) or cookie | --auth cookie |
| Quoting | Same quote rules as the LakeWatch client (' wrapping; Windows " → """) | -p 'keywords=["ERROR"]' |
For the full API catalog, authentication modes, metric names, and password encryption mechanism, see MRS Manager API Client.
Extract the following alarm information from the user input:
| Field | Parameter | Required | Description | Example |
|---|---|---|---|---|
| Alarm ID | alarm_id | Optional | Alarm unique ID, e.g. 12089 | 12007 |
| Alarm name | alarm_name | Required | Alarm Chinese name | PMS进程异常 |
| Occurrence time | alarm_time | Required | Alarm occurrence time | 2026/06/11 16:00:32 GMT+08:00 |
| Cluster ID | cluster_id | Optional | MRS cluster ID | fd04c789-39d4-4847-8fc9-4572fec9414f |
| Host name | node_name | Optional | Host where the alarm occurred (from the alarm location info) | 8-5-225-6 |
| Service name | server_name | Optional | Service that raised the alarm (from the alarm location info) | Manager |
| Role name | role_name | Optional | Role that raised the alarm (from the alarm location info) | pms |
| Additional info | additional_info | Optional | Alarm additional information, usually contains key diagnostic clues |
Notes:
. is a node IP; otherwise it is a host name.Run the mode check script to determine whether diagnosis is based on MRS Manager or LakeWatch:
python scripts/check_api_mode.pypython3 scripts/check_api_mode.pyThe script checks whether scripts/manager_api_config.yaml exists and whether encrypted_password is filled in, and returns a JSON result:
{"mode": "manager", "reason": "..."} // manager-based
{"mode": "lakewatch", "reason": "..."} // lakewatch-based
Rules:
manager_api_config.yaml does not exist → default lakewatchencrypted_password is empty → default lakewatchencrypted_password is not empty → managerBased on the Step 2-1 result:
alarms/<alarm_id>.md under this skill directory to get the diagnosis flow. Also read LakeWatch API Client for the Python script usage.alarm_manager/<alarm_id>.md under this skill directory to get the diagnosis flow. Also read MRS Manager API Client for the Python script usage.If no matching alarm document exists, tell the user: 暂不支持此告警的分析。 (This alarm is not supported for analysis.)
Follow the per-alarm knowledge base from Step 2 to execute the diagnosis.
Diagnosis execution notes:
<alarm_time>, <alarm_node>) MUST be substituted with actual values, never hardcodedlakewatch_api_client.py or manager_api_client.py, use python on Windows and python3 on Linuxlakewatch_api_client.py in lakewatch mode, manager_api_client.py in manager mode-p parameter values MUST be wrapped in single quotes; on Windows PowerShell, every " inside a value must be replaced with """Command failure handling: When a command fails, skip the current check item and continue with the other checks.
Output following the template below:
## Diagnosis Result
| Item | Content |
|------|---------|
| Diagnosis time | [time] |
| Cluster ID | [cluster_id] |
| Alarm name | [alarm_name] |
| Alarm ID | [alarm_id] |
| Alarm node | [node info] |
| Alarm occurrence time | [occurrence time] |
### Root Cause
**Preliminary judgment**: [root cause type]
**Analysis basis**:
- [basis 1]
- [basis 2]
### Repair Suggestion
| Priority | Operation | Description | Needs user confirmation |
|----------|-----------|-------------|-------------------------|
| 1 | [operation 1] | [description] | Yes |
| 2 | [operation 2] | [description] | Yes |
# Query the alarm diagnosis skill content by alarm sequence ID
python3 lakewatch_api_client.py -a query_alarm_skill \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_sequence_id=<alarm_serial_no>'
# Collect node resource data for a given strategy (system-load, memory-usage, disk-space, etc.)
python3 lakewatch_api_client.py -a collect_alarm_node_res_data \
-p 'cluster_id=<cluster_id>' \
-p 'strategy_name=system-load' \
-p 'node_name=<node_name>'
Supported strategy_name values include: system-load, memory-usage, disk-space, disk-io, network-io, file-handle, port-check (needs env={"PORT":"<port>"}), high-cpu-processes, high-memory-process, zombie-process, dns-check, network-connectivity-test (needs env={"TARGET_IP":"<ip>"}), process-basic-info (needs env={"process_name":"java"}), process-file-descriptor (needs env={"PID":"<pid>"}), jstack-thread-dump (needs env={"PID":"<pid>"}), disk-health-check, disk-smart-info (needs env={"DISK_DEVICE":"/dev/sda"}), disk-raid-status, ha-resource-status, and more. See LakeWatch API Client for the full list.
# Collect alarm-related log data around the alarm time
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=2026/06/11 16:00:32 GMT+08:00' \
-p 'log_directory=/var/log/hadoop/hdfs' \
-p 'log_file_name=hadoop-hdfs-datanode.log' \
-p 'keywords=["ERROR","Exception"]' \
-p 'log_type=local'
When the log time format is non-standard ISO (e.g. [2026-07-07 20:54:25,171]), pass time_pattern:
python3 lakewatch_api_client.py -a collect_alarm_log_data \
-p 'cluster_id=<cluster_id>' \
-p 'alarm_time=2026/07/07 20:54:00 GMT+08:00' \
-p 'log_directory=/var/log/Bigdata/omm/oms/pms' \
-p 'log_file_name=pms*.log' \
-p 'keywords=["ERROR","Exception"]' \
-p 'log_type=local' \
-p 'time_pattern=^\[([0-9]{4})-([0-9]{2})-([0-9]{2}) ([0-9]{2}):([0-9]{2}):([0-9]{2})||ymdHMS'
# Query audit dump config via the LakeWatch manager-access proxy
python3 lakewatch_api_client.py -a access_manager_get \
-p 'cluster_id=<cluster_id>' \
-p 'target_url=api/v2/audits/config'
# Query audit logs (returns totalCount)
python3 lakewatch_api_client.py -a access_manager_get \
-p 'cluster_id=<cluster_id>' \
-p 'target_url=api/v2/audits?limit=1'
target_urlMUST NOT start with/. The proxy requires Agent >= 1.0.5 and reported OMS node info. Only GET is supported currently.
When in manager mode, use manager_api_client.py for the equivalent queries:
# Query current alarms (with status=1 for active alarms)
python3 manager_api_client.py -a get_alarms -p 'status=1' --json
# Query instance running status of a service
python3 manager_api_client.py -a get_instances \
-p 'cluster_id=<cluster_id>' \
-p 'service_name=DBService' \
--json
# Query host process status
python3 manager_api_client.py -a get_host_process \
-p 'hostname=<node_name>' \
--json
# Query host monitor metrics (dev_ prefix)
python3 manager_api_client.py -a get_host_metrics \
-p 'hostname=<node_name>' \
-p 'metric_names=dev_cpu_surp_avg,dev_load_one_min' \
--json
# Query host resource usage
python3 manager_api_client.py -a get_host_resource \
-p 'hostname=<node_name>' \
--json
# Search logs by keyword (returns task_id, then poll progress)
python3 manager_api_client.py -a start_log_search \
-p 'cluster_id=<cluster_id>' \
-p 'key_word=ERROR' \
-p 'start_time=<alarm_time>' \
-p 'end_time=<current_time>' \
-p 'services=Manager:Manager:Agent' \
-p 'min_log_level=ERROR' \
--json
python3 manager_api_client.py -a get_log_search_progress \
-p 'search_id=<task_id>' \
--json
# Browse a log file directly (file_name must be a full path)
python3 manager_api_client.py -a browse_log \
-p 'hostname=<node_name>' \
-p 'file_name=/var/log/Bigdata/controller/acs/acs.log' \
-p 'start_line=1' \
-p 'end_line=200' \
--json
start_log_searchservicesformat:component:service:role(e.g.HDFS:HDFS:NameNode), multiple separated by;;start_time/end_timeformat:yyyy-MM-ddTHH:mm:ss. See MRS Manager API Client for metric names and full parameter rules.
| Parameter | Required/Optional | Description | Default |
|---|---|---|---|
alarm_id | Optional | Alarm unique ID, used to locate alarms/<alarm_id>.md (lakewatch mode) or alarm_manager/<alarm_id>.md (manager mode) | N/A |
alarm_name | Required | Alarm Chinese name | N/A |
alarm_time | Required | Alarm occurrence time, format yyyy/MM/dd HH:mm:ss GMT+X:XX | N/A |
cluster_id | Optional | MRS cluster ID | N/A |
node_name | Optional | Alarm node host name or IP | N/A |
server_name | Optional | Service that raised the alarm | N/A |
role_name | Optional | Role that raised the alarm | N/A |
additional_info | Optional | Alarm additional information | N/A |
strategy_name | Required by collect_alarm_node_res_data | Resource collection strategy (lakewatch mode) | N/A |
log_directory | Required by collect_alarm_log_data | Log directory, must be under /var/log/ (lakewatch mode) | N/A |
log_file_name | Required by collect_alarm_log_data | Log file name, no path separators (lakewatch mode) | N/A |
keywords | Required by collect_alarm_log_data | Log keyword filter, JSON array (lakewatch mode) | N/A |
log_type | Required by collect_alarm_log_data | local or hdfs (lakewatch mode) | N/A |
time_pattern | Optional | Non-standard log time regex, format `regex | |
service_name | Required by get_instances | Service name to query instances (manager mode) | N/A |
metric_names | Required by get_host_metrics | Comma-separated monitor metric names with dev_ prefix (manager mode) | N/A |
key_word | Required by start_log_search | Log keyword to search (manager mode) | N/A |
current_time | Required by start_log_search | Current time, format yyyy-MM-ddTHH:mm:ss (manager mode) | N/A |
The diagnosis report is output in Markdown, containing:
See the template in the Workflow → Step 4 section.
See Verification Method for the installation, configuration, and function verification steps.
check_api_mode.py, then read alarms/<alarm_id>.md (lakewatch mode) or alarm_manager/<alarm_id>.md (manager mode) before running any command; do not infer diagnostic steps yourself.<cluster_id>, <alarm_time>, <node_name>, <target_ip>, <current_time>, etc. with actual user-provided values; never hardcode them.-p values in single quotes; on Windows PowerShell, escape " as """ to avoid code:500 errors.alarm_time must follow yyyy/MM/dd HH:mm:ss GMT+X:XX (lakewatch mode) or yyyy-MM-ddTHH:mm:ss (manager mode start_log_search); for non-standard log time formats in lakewatch mode, pass time_pattern.| Document | Description |
|---|---|
| CLI Installation Guide | Python dependencies and LakeWatch client setup |
| IAM Policies | LakeWatch/MRS Manager access model and required roles |
| Verification Method | Installation, configuration, and function verification |
| Acceptance Criteria | Pass/fail criteria for skill testing |
| LakeWatch API Client | Full LakeWatch API catalog, parameters, token and encryption mechanism |
| MRS Manager API Client | Full MRS Manager API catalog, authentication modes, metric names and encryption mechanism |
| Related Commands | Common LakeWatch/Manager API commands quick reference |
alarms/<alarm_id>.md | Per-alarm diagnosis knowledge base (lakewatch mode, mirrors the MRS product alarm catalog) |
alarm_manager/<alarm_id>.md | Per-alarm diagnosis knowledge base (manager mode, mirrors the MRS product alarm catalog) |
--encrypt-password and stored in the corresponding config YAML. Repair steps are suggestions only.hcloud; it calls the LakeWatch API through lakewatch_api_client.py or the MRS Manager REST API through manager_api_client.py. Do not mix in hcloud commands.check_api_mode.py to determine the mode. In manager mode use manager_api_client.py and the alarm_manager/ knowledge base; in lakewatch mode use lakewatch_api_client.py and the alarms/ knowledge base. Do not mix clients across modes.access_manager_get proxy only supports GET requests (PUT is not yet available on the Agent side); collect_alarm_log_data requires log_directory to be under /var/log/; some strategy_name values require extra env parameters; in manager mode browse_log requires the full log file path and start_log_search services must follow component:service:role.