Install
openclaw skills install @huaweicloudskill/huawei-cloud-mrs-hdfs-performance-issue-analysisHuawei Cloud MRS HDFS performance issue analysis skill. Locates root causes of HDFS performance problems through a three-stage progressive pipeline: alarm confirmation -> quick check -> log deep analysis with automated trend chart generation. Built-in Python analyzer (scripts/hdfs_perf_analyze.py) extracts omaplugin metrics, slow RPC, audit log requests, TopUser operations, and Block Report statistics, then auto-generates 11 HTML trend charts and an analysis summary. Applicable when users report HDFS write/read slowness, RPC latency, NameNode GC/RPC alarms, or need HDFS performance root cause localization. 触发词:"HDFS性能问题"、"HDFS写入慢"、"HDFS读取慢"、"NameNode RPC冲高"、"RPC响应时间长"、"HDFS性能变慢"、"ALM-14006"、"ALM-14007"、"ALM-14014"、"ALM-14015"、"ALM-14021"、"ALM-14022"、"HDFS性能定位"、"HDFS性能分析"
openclaw skills install @huaweicloudskill/huawei-cloud-mrs-hdfs-performance-issue-analysisYou are an MRS HDFS performance issue analysis expert, responsible for locating root causes of HDFS performance problems on Huawei FusionInsight HD / MRS clusters. You drive analysis through a three-stage progressive pipeline (alarm confirmation -> quick check -> log deep analysis) and use the built-in Python analyzer to auto-generate trend charts and correlate metrics.
Architecture: Caller (Agent) -> scripts/hdfs_perf_analyze.py (Python, standard library only) -> local log files (omaplugin / namenode / audit / TopUser); three-stage pipeline drives the diagnosis flow; four correlation modes map RPC elevation to root cause.
Note on language: This SKILL.md is written in English per the repository spec. Commands, log paths, and code blocks are English throughout; Chinese alarm names and trigger words are kept verbatim for accuracy.
Applicable Scenarios:
Typical Use Cases:
Performance Issue Categories (4 root cause modes):
| Mode | Name | Judgment Method | Root Cause |
|---|---|---|---|
| MODE01 | Large write volume | RPC elevated + Blocks Total rises synchronously | Large write operations (create/addBlock) drive RPC up |
| MODE02 | Business operations surge | RPC elevated + total operations rise synchronously (ops chart confirms the specific op type) | Business-side operation volume increases RPC |
| MODE03 | Balance / stop instance / large delete | RPC elevated + pending deletion / under-replicated / excess blocks rise | Balance, stop-instance, or large delete operations drive RPC up |
| MODE04 | NameNode node performance insufficient | RPC elevated + disk IO rises or Block Report processing time increases | NameNode node itself has insufficient performance (disk IO, CPU, load, power-saving mode) |
Related Alarms:
| Alarm ID | Alarm Name | Related Issue |
|---|---|---|
| ALM-14006 | HDFS 文件数超过阈值 | NameNode file object count exceeds memory plan |
| ALM-14007 | NameNode 堆内存使用率超过阈值 | NameNode memory insufficient |
| ALM-14014 | NameNode 进程 GC 时间超过阈值 | NameNode GC problem |
| ALM-14015 | DataNode 进程 GC 时间超过阈值 | DataNode GC problem |
| ALM-14021 | NameNode RPC 处理平均时间超过阈值 | NameNode RPC processing capacity insufficient |
| ALM-14022 | NameNode RPC 队列平均时间超过阈值 | NameNode RPC queue backlog |
Log Paths:
| Log Type | Path |
|---|---|
| NameNode runtime log | /var/log/Bigdata/hdfs/nn/hadoop-omm-namenode-<hostname>.log |
| NameNode audit log | /var/log/Bigdata/audit/hdfs/nn/hdfs-audit-namenode.log |
| NameNode Agent log (omaplugin) | /var/log/Bigdata/nodeagent/monitorlog/omaplugin.log |
| NameNode GC log | #{BigdataLogHome}/hdfs/nn/namenode-omm-gc.log |
| TopUser operation log | /var/log/Bigdata/audit/hdfs/nn/5min-TopUserOpCounts.log |
python3 --version (Linux) / python --version (Windows)scripts/hdfs_perf_analyze.py) performs local static log analysis only, no cluster connection required<log_dir>/trends/hdfs dfs -ls /system/balancer.id, grep on cluster log paths, etc.) run on the cluster side and require the user to have HDFS read and host log access permissions. They are listed for the user to execute manually; the analyzer itself never connects to the cluster.The analysis follows a strict three-stage progressive pipeline. Do not skip stages.
Ask the user whether any of the alarms listed in Section 1 Related Alarms are present, and wait for the user's confirmation.
| User Answer | Next Step |
|---|---|
| Has alarm (e.g. ALM-14021 / ALM-14022) | Go to Step 2 Scenario A |
| Has alarm (e.g. ALM-14006 / ALM-14007 / ALM-14014 / ALM-14015) | Go to Step 2 Scenario A, focus on NameNode memory / file object / GC checks |
| No alarm | Go to Step 2 Scenario B |
In this stage, the user runs read-only commands on the cluster. No logs are required.
Provide the corresponding possible causes and quick-check commands based on the alarm type.
For ALM-14021 / ALM-14022 (RPC capacity insufficient), provide these quick-check commands directly:
# 1. Check for Balance task (balancer.id creation time = balance start time)
hdfs dfs -ls /system/balancer.id
# 2. Check slow RPC, take the top 100 by latency, ascending
grep -i "slow rpc" /var/log/Bigdata/hdfs/nn/hadoop-omm-namenode-*.log | awk '{match($0, /took ([0-9]+)ms/, arr); print arr[1]" "$0}' | sort -rn | head -100 | sort -n | cut -d' ' -f2-
# 3. In the audit log, find the directory of the operation corresponding to the slow RPC
grep "<time>" /var/log/Bigdata/audit/hdfs/nn/hdfs-audit-namenode.log | grep -E "<op_type>"
# 4. Check NameNode GC
grep -i "fullgc\|Full GC" /var/log/Bigdata/hdfs/nn/namenode-omm-*-gc.log*
# 5. Check rack imbalance
grep -i "TOO_MANY_NODES_ON_RACK" /var/log/Bigdata/hdfs/nn/namenode-omm-server-*.log
# 6. Check NameNode node CPU and disk IO
top
iostat -x 1 5
# CPU over 60% or iostat %util persistently over 90% indicates a problem
# 7. Check whether CPU is in power-saving mode
# Step 1: if the count is 0, it is performance mode; otherwise go to step 2
ls /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor | wc -l
# Step 2: empty result means performance mode; non-empty means power-saving mode
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor | grep -v ondemand | grep -v performance | grep -v smartass
For ALM-14006 / ALM-14007 / ALM-14014 / ALM-14015 (memory / file object / GC), check whether NameNode file object count exceeds the memory plan. See Section 5 Root Cause Knowledge Base Chapter 1 and Chapter 2.
Provide 9 generic possible causes and the following quick-check commands:
# 1. Check slow rpc
grep -i "slow rpc" /var/log/Bigdata/hdfs/nn/hadoop-omm-namenode-node-hostname.* | sort -nk 21
# 2. Check for Balance task
hdfs dfs -ls /system/balancer.id
# 3. Check large directory scan operations
grep -E "listStatus|contentSummary|quotaUsage" /var/log/Bigdata/audit/hdfs/nn/hdfs-audit-namenode.log
# 4. Check NameNode node CPU, load, disk IO
top && uptime && iostat -x 1 10
# 5. Check DataNode disk space (NameNode native page)
# 6. Check NameNode GC
grep -i "fullgc\|Full GC" /var/log/Bigdata/hdfs/nn/namenode-omm-*-gc.log*
# 7. Check rack imbalance
grep -i "TOO_MANY_NODES_ON_RACK" /var/log/Bigdata/hdfs/nn/namenode-omm-server-*.log
# 8. Check whether CPU is in power-saving mode
ls /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor | wc -l
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor | grep -v ondemand | grep -v performance | grep -v smartass
Decision:
| Quick-Check Result | Next Step |
|---|---|
| Problem resolved | Workflow ends, output root cause and solution |
| Problem persists | Go to Step 3 |
When the quick check does not resolve the issue, the user provides logs for the abnormal time window.
| Log Type | Path |
|---|---|
| NameNode runtime log | /var/log/Bigdata/hdfs/nn/hadoop-omm-namenode-<hostname>.log |
| NameNode audit log | /var/log/Bigdata/audit/hdfs/nn/hdfs-audit-namenode.log |
| NameNode Agent log (omaplugin) | /var/log/Bigdata/nodeagent/monitorlog/omaplugin.log |
| TopUser operation log | /var/log/Bigdata/audit/hdfs/nn/5min-TopUserOpCounts.log |
Note: 5min-TopUserOpCounts.log is the canonical filename (the older /var/log/Bigdata/audit/hdfs/nn/top-user-operation-info path is deprecated).
Before running the analyzer, ask the user for the abnormal time window (e.g. 2026-06-11 09:00:00 ~ 2026-06-11 12:00:00). Do not run analysis without a confirmed time range.
Run the Python analyzer per Section 4 Core Commands. The analyzer extracts omaplugin metrics, slow RPC, audit log requests, TopUser operations, and Block Report statistics, then generates 11 HTML trend charts and an analysis summary in <log_dir>/trends/.
After the trend charts are generated, compare the RPC elevation window against the other metrics. For each mode, explicitly state whether it is PRESENT or ABSENT and provide evidence (e.g. "RPC peak at 09:15, Blocks Total rose from 142,855,636 to 142,910,200 in the same window").
| Mode | Judgment Method | Root Cause | Solution |
|---|---|---|---|
| MODE01 | RPC elevated + nn_blockstotal_trend.html Blocks Total rises synchronously | Large write operations (create/addBlock) | Find the write source, schedule writes off-peak; increase dfs.namenode.handler.count |
| MODE02 | RPC elevated + nn_topuser_all_trend.html total operations rise synchronously; use nn_topuser_ops_trend.html to confirm the specific op type (listStatus / contentSummary / getfileinfo / delete) | Business-side operation volume surge | Negotiate with business side to lower op frequency; for large ops (listStatus / contentSummary) switch to smaller-grained stats; increase dfs.namenode.handler.count |
| MODE03 | RPC elevated + any of nn_pendingdeletionblocks_trend.html / nn_underreplicatedblocks_trend.html / nn_excessblocks_trend.html rises | Balance / stop instance / large delete | Large delete: batch off-peak, lower dfs.namenode.replication.work.multiplier.per.iteration; Balance: stop Balance, large cluster uses -f for partial node migration; stop instance: schedule off-peak, pre-assess RPC impact |
| MODE04 | RPC elevated + nn_io_trend.html disk IO rises OR nn_blockreport_trend.html Block Report processing time increases | NameNode node insufficient performance (disk IO high / CPU high / load high / power-saving mode / log compression archive IO surge) | Disk IO high: migrate log dir to dedicated disk, use higher-performance disk, schedule log rolling off-peak; CPU high: find and migrate high-CPU process, keep NameNode CPU under 50%; load high: find load source, migrate non-HDFS processes; power-saving: switch to performance mode, disable power-saving in BIOS; log archive: schedule log rolling/compression off-peak |
If multiple modes are PRESENT, rank them by impact and give the comprehensive root cause.
The analysis report is output in Markdown containing:
# Linux
python3 <skill_dir>/scripts/hdfs_perf_analyze.py <log_dir> <time_start> <time_end>
# Windows
python <skill_dir>\scripts\hdfs_perf_analyze.py <log_dir> <time_start> <time_end>
Example:
python3 scripts/hdfs_perf_analyze.py "/tmp" "2026-06-11 09:00:00" "2026-06-11 12:00:00"
# Extract a specific omaplugin metric (e.g. RPC processing time)
grep -E "key=nn_rpcprocessingtimeavgtime_client_rt" <log_dir>/omaplugin/*.log
# Extract slow RPC and sort by latency
grep -i "slow rpc" <log_dir>/hadoop-omm-namenode/*.log | awk '{match($0, /took ([0-9]+)ms/, arr); print arr[1]" "$0}' | sort -rn | head -100
# Extract processReport entries (blocks + processing time)
grep "processReport" <log_dir>/hadoop-omm-namenode/*.log | grep -oE "blocks: [0-9]+|processing time: [0-9]+ msecs"
# Extract 5min TopUser op counts
grep -E "300s Operation Collecter" <log_dir>/5min-TopUserOpCounts.log
| Parameter | Required/Optional | Description | Default |
|---|---|---|---|
log_dir | Required | Directory containing the log files (supports zip auto-extraction for omaplugin / namenode / audit logs) | N/A |
time_start | Required | Analysis start time, format YYYY-MM-DD HH:MM:SS | N/A |
time_end | Required | Analysis end time, format YYYY-MM-DD HH:MM:SS | N/A |
log_dir)| File | Description | Supported Format |
|---|---|---|
omaplugin*.log | NameNode Agent monitoring log | .log / .zip |
hadoop-omm-namenode*.log | NameNode runtime log | .log / .zip |
hdfs-audit-namenode*.log | NameNode audit log | .log / .zip |
5min-TopUserOpCounts.log | TopUser operation statistics log | .log |
The analyzer generates 11 output files in <log_dir>/trends/:
| File | Description | Threshold Line |
|---|---|---|
nn_rpc_processing_trend.html | RPC processing average time trend | 100ms |
nn_rpc_queue_trend.html | RPC queue average time trend | 400ms |
nn_pendingdeletionblocks_trend.html | Pending deletion blocks count trend | - |
nn_underreplicatedblocks_trend.html | Under-replicated blocks count trend | - |
nn_excessblocks_trend.html | Excess blocks count trend | - |
nn_blockstotal_trend.html | Total blocks count trend | - |
nn_io_trend.html | NameNode disk IO read/write rate trend (combined) | - |
nn_topuser_ops_trend.html | TopUser per-operation-type request count trend | - |
nn_topuser_all_trend.html | Total operations trend (sum of all RPC types, excluding all, hover shows value and time) | - |
nn_blockreport_trend.html | NameNode Block Report trend (per-minute blocks sum + processing time sum, dual Y-axis) | - |
analysis_summary.txt | Analysis summary (contains root cause analysis) | - |
Trend chart styling rules:
#0d1117, chart area #161b22)M / K abbreviations (e.g. 142,855,636 not 142.9M)After the analyzer finishes, the report must list the trend chart file location (e.g. "Trend charts generated in <log_dir>/trends/") and each file's description so the user can open them in a browser.
| File Count | File System Object Count | Recommended JVM Params |
|---|---|---|
| 5,000,000 | 10,000,000 | -Xms6G -Xmx6G -XX:NewSize=512M -XX:MaxNewSize=512M |
| 10,000,000 | 20,000,000 | -Xms12G -Xmx12G -XX:NewSize=1G -XX:MaxNewSize=1G |
| 25,000,000 | 50,000,000 | -Xms32G -Xmx32G -XX:NewSize=3G -XX:MaxNewSize=3G |
| 50,000,000 | 100,000,000 | -Xms64G -Xmx64G -XX:NewSize=6G -XX:MaxNewSize=6G |
| 100,000,000 | 200,000,000 | -Xms96G -Xmx96G -XX:NewSize=9G -XX:MaxNewSize=9G |
| 150,000,000 | 300,000,000 | -Xms164G -Xmx164G -XX:NewSize=12G -XX:MaxNewSize=12G |
Note: when modifying GC_OPTS, only modify the first four parameters.
Repair:
| Per-DataNode Block Count | Recommended JVM Params |
|---|---|
| 2,000,000 | -Xms6G -Xmx6G -XX:NewSize=512M -XX:MaxNewSize=512M |
| 5,000,000 | -Xms16G -Xmx16G -XX:NewSize=1G -XX:MaxNewSize=2G |
Spec limit: a single DataNode instance supports up to 5,000,000 blocks.
Repair:
Possible causes:
Solutions per cause:
| Cause | Solution |
|---|---|
| Balance task running | Stop Balance; large cluster uses -f for partial node migration |
| Large directory scan | Lower op frequency; smaller-grained stats; disable fine-grained monitoring |
| NameNode CPU high | Migrate non-HDFS processes; keep CPU under 60% |
| NameNode disk IO insufficient | Use higher-performance disk; mount NameNode instance on a dedicated disk |
| DataNode disk usage > 90% | Scale out nodes; delete unused data; temporarily set dfs.namenode.redundancy.considerLoad=false |
| Rack node imbalance | Keep rack node count consistent |
| File objects exceed memory plan | Adjust NameNode memory config or delete unused files |
Possible causes:
fs.du.interval misconfigured (8 version before 8203 default 60000, adjust to 600000)Slow log types:
| Slow Type | Meaning |
|---|---|
| Slow BlockReceiver write packet to mirror | Network write block latency |
| Slow BlockReceiver write data to disk cost | Block write to OS cache or disk latency |
| Slow flushOrSync | Block write to OS cache or disk latency |
| Slow manageWriterOsCache | Block write to OS cache or disk latency |
Repair:
fs.du.interval.manageWriterOsCache slowness, set dfs.datanode.drop.cache.behind.writes=false and dfs.datanode.drop.cache.behind.reads=false.| File Size | File Object Count |
|---|---|
| < 128MB | 1 (file) + 1 (block) = 2 |
| > 128MB (e.g. 128G) | 1 (file) + 1024 (blocks) = 1025 |
Primary/standby NameNode max file object count: 300 million (corresponds to 150 million small files).
| Item | Spec |
|---|---|
| Max block replica count per DataNode instance | 5,000,000 |
| Max block replica count per DataNode per disk | 500,000 |
| Minimum disk count per DataNode | 10 |
HDFS Block * 3 (default replication factor).
HDFS Block * 3 / DataNode node count.
| Parameter | Default | Tuning Value | Description |
|---|---|---|---|
dfs.namenode.handler.count | 64 | 192 | NameNode handler thread count |
ipc.server.read.threadpool.size | 15 | - | NameNode request thread pool size |
dfs.namenode.redundancy.considerLoad | true | false | Temporarily bypass busy nodes |
fs.du.interval | 60000 | 600000 | 8 version before 8203 needs adjustment |
dfs.datanode.drop.cache.behind.writes | false | - | Whether to drop write cache |
dfs.datanode.drop.cache.behind.reads | false | - | Whether to drop read cache |
Prepare a test log directory with the four log files listed in Section 5.1.
Run the analyzer:
python3 scripts/hdfs_perf_analyze.py "<test_log_dir>" "2026-06-11 09:00:00" "2026-06-11 12:00:00"
Verify the 11 output files exist in <test_log_dir>/trends/:
analysis_summary.txtOpen one HTML file in a browser, confirm the dark-theme SVG renders, hover over a data point, and confirm the tooltip shows metric name / value / time.
For metrics with thresholds (RPC processing 100ms, RPC queue 400ms), confirm the red dashed threshold line is rendered.
#0d1117); open them in a browser for hover tooltips.142,855,636, not 142.9M).| Reference | Description | Related Section |
|---|---|---|
| FusionInsight HD Performance Problem Location Guide | Source document for the three-stage pipeline, four correlation modes, and the capacity specs | 1. Overview, 5. Root Cause Knowledge Base |
| MRS HDFS Product Documentation | NameNode / DataNode architecture, memory planning, block report mechanism | 1. Overview, 7. Root Cause Knowledge Base |
| HDFS Operation Guide | Balance operation, rack config, JVM tuning | 4. Core Commands, 9. Common Tuning Parameters |
When encountering an HDFS performance issue, check in the following order:
Method 1: From the HDFS home page, click NameNode(Active).
Method 2: Open the URL directly:
https://<Manager IP>:20026/HDFS/NameNode/<instanceId>/dfshealth.html
instanceId can be found on the Manager page by hovering over the instance name.
NameNode startup time reference:
| File Count | Object Count | Approx. Startup Time |
|---|---|---|
| 10,000,000 | 20,000,000 | ~3 minutes |
| 60,000,000 | 150,000,000 | ~8 minutes |
| 120,000,000 | 300,000,000 | ~30 minutes |
# Set Balance max bandwidth (optional)
hdfs dfsadmin -setBalancerBandwidth 209715200 # 200MB/s
# Balance all DataNode nodes
sh /opt/client/HDFS/hadoop/sbin/start-balancer.sh -threshold 10
# Balance a subset of DataNode nodes
sh /opt/client/HDFS/hadoop/sbin/start-balancer.sh -include -f /tmp/includeHost.txt -threshold 10
# View Balance task progress
cat /opt/client/HDFS/hadoop/logs/hadoop-root-balancer-<hostname>.out
# Stop Balance task
sh /opt/client/HDFS/hadoop/sbin/stop-balancer.sh
This is expected. HDFS reserves 10% disk space for Yarn by default, so when disk usage reaches 90%, HDFS considers the space full.
omaplugin*.zip, hadoop-omm-namenode*.zip, and hdfs-audit-namenode*.zip into subdirectories before parsing.<log_dir>/trends/; the analyzer creates this directory if it does not exist.top-user-operation-info path is deprecated.