Complete methodology for local LLM performance optimization. Core principle: maximize context while fully covering GPU memory — find the sweet spot where GPU runs at full speed. Step-by-step 4-phase 10-step control variable testing process. Works for ALL llama.cpp / llama-server models on ANY hardware. Cases: Qwen3.5-MoE, Qwen3.6-35B, Qwen3.6-27B (2026-04-28).

Install

openclaw skills install @hoperealize/llama-params-optimizer