Install
openclaw skills install @crazyss/linux-kernel-crash-debugDebug Linux kernel crashes using evidence-first vmcore analysis, the crash utility, and memory/concurrency debugging tools. Use when users mention kernel crash, kernel panic, vmcore analysis, kernel dump debugging, crash utility, kernel oops debugging, pstore or ramoops, soft/hard lockup, hung task, OOM, locating root causes of kernel issues, regression bisection, mutex ownership, ARM64 lock-pointer recovery, KASAN, KFENCE, KCSAN, Lockdep, drgn, Kprobes, Kmemleak, memory corruption, out-of-bounds access, use-after-free, race, deadlock, or memory leak detection.
openclaw skills install @crazyss/linux-kernel-crash-debugThis skill guides you through analyzing Linux kernel crash dumps using the crash utility.
claude skill install linux-kernel-crash-debug.skill
# Method 1: Install via ClawHub
clawhub install linux-kernel-crash-debug
# Method 2: Manual installation
mkdir -p ~/.openclaw/workspace/skills/linux-kernel-crash-debug
cp SKILL.md ~/.openclaw/workspace/skills/linux-kernel-crash-debug/
# Analyze a dump file
crash vmlinux vmcore
# Debug a running system
crash vmlinux
# Raw RAM dump
crash vmlinux ddr.bin --ram_start=0x80000000
0. Preserve checksums, vmcore-dmesg, build IDs, modules, config, and command line
1. crash> sys # Validate release/build and panic context
2. crash> log # Find the FIRST anomaly, not only the last panic
3. crash> bt / bt -a # Compare panic task with all active CPUs
4. crash> mod # Confirm faulting module symbols are available
5. crash> struct / kmem # Test a specific object-lifetime hypothesis
6. Search upstream and verify good/bad kernels before claiming a regression
Read references/evidence-first-workflow.md before deep analysis. It defines
the evidence-quality gates, failure-type routing, hypothesis ledger, tool
escalation rules, and root-cause report format. Never equate the panic task,
fault site, corruption site, and root cause without supporting evidence.
Default to offline, read-only analysis. Treat sudo, module load/unload, boot
or service configuration, writes under debugfs or /proc/sys, live tracing,
and SysRq actions as live-host mutations.
kdumpctl test. Explain the prerequisites and hand the final trigger to an
authorized human following an approved drill with console access, backups,
workload evacuation, and a verified rollback/recovery plan.If you are an AI/Agent using this skill, do not invoke crash interactively as it will block your subshell.
./scripts/agent-crash.sh which maps precisely to the workflows below but safely truncates outputs:
./scripts/agent-crash.sh -k vmlinux -c vmcore triage - Runs sys, a high-signal log index, panic/all-CPU backtraces, and module inventory../scripts/agent-crash.sh -k vmlinux -c vmcore flow-oom - Top 15 memory checks../scripts/agent-crash.sh -k vmlinux -c vmcore flow-deadlock - Pulls UN task stacks../scripts/agent-crash.sh -k vmlinux -c vmcore dis-regs <func> <pid> - Assembly regression../scripts/agent-crash.sh -k vmlinux -c vmcore check-poison <addr> - Pattern match memory poisons../scripts/agent-crash.sh -k vmlinux -c vmcore run "rd ffff880123456780".references/agentic-heuristics.md for extended expert methodologies.references/evidence-first-workflow.md: report symbol/dump quality,
identify the earliest anomaly, keep competing hypotheses, and attach a
confidence level plus a disproof test to the conclusion.| Item | Requirement |
|---|---|
| vmlinux | Must have debug symbols (CONFIG_DEBUG_INFO=y) |
| vmcore | kdump/netdump/diskdump/ELF format |
| Version | vmlinux must exactly match the vmcore kernel version |
# Install crash utility
sudo dnf install crash
# Install kernel debuginfo (match your kernel version)
sudo dnf install kernel-debuginfo-$(uname -r)
# Install additional analysis tools
sudo dnf install gdb readelf objdump makedumpfile
# Optional: Install kernel-devel for source code reference
sudo dnf install kernel-devel-$(uname -r)
sudo dnf install crash gdb binutils makedumpfile kexec-tools
# Enable the matching debuginfo repository first if needed
sudo dnf debuginfo-install kernel-$(uname -r)
sudo apt install crash kdump-tools kexec-tools gdb binutils makedumpfile
# Debian and Ubuntu use different debug-symbol repositories/package suffixes;
# query the exact running-kernel package before installing it.
apt-cache search "linux-image-$(uname -r).*dbg\|linux-image-$(uname -r).*dbgsym"
sudo zypper install crash kexec-tools makedumpfile
# Optional SUSE kdump UI and matching kernel debuginfo
sudo zypper install yast2-kdump
zypper se -s 'kernel*debug*'
# Enable debug symbols in kernel config
make menuconfig # Enable CONFIG_DEBUG_INFO, CONFIG_DEBUG_INFO_REDUCED=n
# Or set directly
scripts/config --enable CONFIG_DEBUG_INFO
scripts/config --enable CONFIG_DEBUG_INFO_DWARF_TOOLCHAIN_DEFAULT
# Check crash version
crash --version
# Verify debuginfo matches kernel
crash /usr/lib/debug/lib/modules/$(uname -r)/vmlinux /proc/kcore
| Command | Purpose | Example |
|---|---|---|
sys | System info/panic reason | sys, sys -i |
log | Kernel message buffer | log, log | tail |
bt | Stack backtrace | bt, bt -a, bt -f |
struct | View structures | struct task_struct <addr> |
p/px/pd | Print variables | p jiffies, px current |
kmem | Memory analysis | kmem -i, kmem -S <cache> |
| Command | Purpose | Example |
|---|---|---|
ps | Process list | ps, ps -m | grep UN |
set | Switch context | set <pid>, set -p |
foreach | Batch task operations | foreach bt, foreach UN bt |
task | task_struct contents | task <pid> |
files | Open files | files <pid> |
| Command | Purpose | Example |
|---|---|---|
rd | Read memory | rd <addr>, rd -p <phys> |
search | Search memory | search -k deadbeef |
vtop | Address translation | vtop <addr> |
list | Traverse linked lists | list task_struct.tasks -h <addr> |
The most important debugging command:
crash> bt # Current task stack
crash> bt -a # All CPU active tasks
crash> bt -f # Expand stack frame raw data
crash> bt -F # Symbolic stack frame data
crash> bt -l # Show source file and line number
crash> bt -e # Search for exception frames
crash> bt -v # Check stack overflow
crash> bt -R <sym> # Only show stacks referencing symbol
crash> bt <pid> # Specific process
Crash session has a "current context" affecting bt, files, vm commands:
crash> set # View current context
crash> set <pid> # Switch to specified PID
crash> set <task_addr> # Switch to task address
crash> set -p # Restore to panic task
# Output control
crash> set scroll off # Disable pagination
crash> sf # Alias for scroll off
# Output redirection
crash> foreach bt > bt.all
# GDB passthrough
crash> gdb bt # Single gdb invocation
crash> set gdb on # Enter gdb mode
(gdb) info registers
(gdb) set gdb off
# Read commands from file
crash> < commands.txt
| Aspect | x86_64 | ARM64 |
|---|---|---|
| crash command | crash vmlinux vmcore | crash vmlinux vmcore; add -m only as a recovery path |
| KASLR | Usually auto-handled from VMCOREINFO | Usually auto-handled; derive -m kaslr=<offset> only for raw/damaged metadata |
| Virtual address bits | fixed for the analyzed build | VMCOREINFO first; -m vabits_actual=<N> is a fallback |
| Physical base | phys_base from VMCOREINFO | VMCOREINFO first; -m phys_offset=<addr> is a fallback |
| VA-PA offset | __START_KERNEL_map | VMCOREINFO first; -m kimage_voffset=<val> is a fallback |
| Frame pointer | RBP (often optimized away) | FP (x29) explicit |
| Calling convention | RDI/RSI/RDX/RCX/R8/R9 | X0-X7 |
For complete ARM64 address parameter derivation, see
references/arm64-crash-params.md. For kdump end-to-end setup, seereferences/kdump-setup-guide.md.
crash_arm64 \
-m vabits_actual=39 \
-m phys_offset=0x80000000 \
-m kimage_voffset=0xffffffc000000000 \
-m kaslr=0x0 \
vmlinux vmcore
First try
crash vmlinux vmcore. Use explicit values only when VMCOREINFO is absent/damaged or the input is raw RAM.kaslr=0means KASLR was disabled; never assume that value or reuse another boot's parameters.
crash> sys # Confirm panic
crash> log | tail -50 # View logs
crash> bt # Call stack
crash> bt -f # Expand frames for parameters
crash> struct <type> <addr> # Inspect data structures
crash> bt -a # All CPU call stacks
crash> ps -m | grep UN # Uninterruptible processes
crash> foreach UN bt # View waiting reasons
crash> struct mutex <addr> # Inspect lock state
crash> kmem -i # Memory statistics
crash> kmem -S <cache> # Inspect slab
crash> vm <pid> # Process memory mapping
crash> search -k <pattern> # Search memory
crash> bt -v # Check stack overflow
crash> bt -r # Raw stack data
Sources: mutex lock pointer and rwsem lock derivation, Kernel Panic Lab.
When a task sleeps in mutex_lock() or a rwsem slow path, trace the first ARM64 argument (x0) at the call site:
# Path A: a global lock is constructed directly
adrp x0, 0xffffffc00ac1e000
add x0, x0, #0x7f0 # lock = page + offset
bl mutex_lock
# Path B: the caller passes a callee-saved register
mov x0, x19 # lock pointer is the saved x19 value
bl mutex_lock
# Find the exact "add x29, sp, #N" and "stp/str ..., x19, [sp,#M]",
# derive SP from the frame pointer, then read the x19 stack slot with rd.
crash> struct mutex <lock_addr> -x
# mutex.owner packs flags into its low 3 bits on common kernels:
# owner_task = owner.counter & ~0x7
crash> struct task_struct <owner_task>
crash> bt <owner_pid>
stp x20, x19, [sp,#32] saves x20 at sp+32 and x19 at sp+40. Never reuse example offsets blindly: derive them from the vmcore's matching vmlinux. Verify the mutex layout and owner flag definitions against the analyzed kernel.
For exact FP/SP arithmetic,
stpslot ordering, owner masking, and failure checks, readreferences/arm64-lock-analysis.md. For a complete rwsem example, read Case 11 inreferences/case-studies.md.
x86_64 equivalent: Use RBP chain with
bt -f. Note that with-fomit-frame-pointer, this technique may fail; in that case usebt -For look for explicit stack frames.
Three independent paths to diagnose memory leaks:
The first path is read-only. Enabling page_owner or writing kmemleak controls
changes a live kernel and must follow the Live-System Safety Contract above.
Preserve the current output before clearing detector state.
# === Layer 1: /proc 三件套 (read from running system or captured info) ===
# MemAvailable 持续下降 + SUnreclaim 持续增加 → slab 内存泄露
cat /proc/meminfo
cat /proc/slabinfo
cat /proc/buddyinfo
# === Layer 2: SLAB-specific (slub_debug) ===
# In bootargs: slub_debug=u,kmalloc-512
# Then read:
cat /sys/kernel/debug/slab/kmalloc-512/alloc_traces
cat /sys/kernel/debug/slab/kmalloc-512/free_traces
# === Layer 3: >8K allocations (page_owner) ===
# SUnreclaim rises but slabinfo flat → kmalloc > 8K uses alloc_pages directly
# Enable CONFIG_PAGE_OWNER + boot with page_owner=on
# If page_owner is not already enabled, use an approved maintenance runbook;
# do not enable it as part of automated triage.
# Periodic dumps, then diff:
./page_owner_sort --cull name,ator,stacktrace page_owner_begin.txt > begin.txt
./page_owner_sort --cull name,ator,stacktrace page_owner_end.txt > end.txt
# Compare begin.txt vs end.txt - rising stacks are leaks
# === Alternative: kmemleak ===
# CONFIG_DEBUG_KMEMLEAK + kmemleak=on bootarg
echo scan > /sys/kernel/debug/kmemleak
cat /sys/kernel/debug/kmemleak
crash> bt -f # Get pointers
crash> struct file.f_dentry <addr>
crash> struct dentry.d_inode <addr>
crash> struct inode.i_pipe <addr>
crash> kmem -S inode_cache | grep counter | grep -v "= 1"
crash> list task_struct.tasks -s task_struct.pid -h <start>
crash> list -h <addr> -s dentry.d_name.name
For detailed information, refer to the following reference files:
| File | Content |
|---|---|
references/advanced-commands.md | Advanced commands: list, rd, search, vtop, kmem, foreach |
references/vmcore-format.md | vmcore file format, ELF structure, VMCOREINFO |
references/case-studies.md | Debugging cases: kernel BUG, deadlock, OOM, NULL pointer, stack overflow |
references/debug-tools-guide.md | Advanced debugging tools: KASAN, Kprobes, Kmemleak, UBSAN (require kernel rebuild) |
references/kdump-setup-guide.md | NEW End-to-end kdump configuration (x86_64 + ARM64, crashkernel syntax, sysrq triggers) |
references/arm64-crash-params.md | NEW ARM64-specific crash address parameters (vabits_actual, phys_offset, kimage_voffset, kaslr) |
references/arm64-lock-analysis.md | ARM64 assembly/stack recovery of mutex and rwsem pointers, plus mutex owner decoding |
references/evidence-first-workflow.md | NEW Evidence gates, timeline reconstruction, failure routing, tool escalation, regression verification, report template |
references/sources.md | NEW Complete bibliography of reference materials used to enhance this skill |
Usage:
crash> help <command> # Built-in help
# Or ask Claude to view reference files
crash: vmlinux and vmcore do not match!
# -> Ensure vmlinux version exactly matches vmcore
crash: cannot find booted kernel
# -> Specify vmlinux path explicitly
crash: cannot resolve symbol
# -> Check if vmlinux has debug symbols
⚠️ Dangerous Operations
The following commands can cause system damage or data loss:
| Command | Risk | Recommendation |
|---|---|---|
wr | Writes to live kernel memory | NEVER use on production systems - can crash or corrupt running kernel |
| GDB passthrough | Unrestricted memory access | Use with caution, may modify memory or registers |
| Kprobes/ftrace/debugfs writes | Changes live instrumentation and may expose runtime data | Require explicit authorization, bounded capture, and cleanup |
| Boot/service configuration | Persists across reboot or changes crash recovery | Back up current state and provide rollback before applying |
SysRq crash / kdumpctl test | Deliberately panics the host | Human-operated approved drill only; agents must not execute |
🔒 Sensitive Data Handling
shred or secure delete when disposing of vmcore files🛡️ Best Practices
makedumpfile -d to filter sensitive pages before analysisbt, files, vm commands are affected by current contextwr command modifies running kernel, extremely dangerousThis is an open-source project. Contributions are welcome!
See CONTRIBUTING.md for guidelines.