Install
openclaw skills install @sunshinejnjn/music-with-comfyuiUse to generate music and audio through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like "make me a song", "make up a song", "compose", "write a song", "compose a track", "arrange a song", "编一个歌", "写歌", "作曲", "编个曲", "做首歌", or any similar instruction to write/compose/arrange a song or melody (in any language). Important disambiguation: in Chinese, "写歌" / "写一个歌" / "来个歌" here means generate an actual audio song through ComfyUI (AceStep 1.5 → MP3), NOT merely writing lyrics text. If you only want lyrics written down (words, no audio), that is plain text writing and NOT this skill. Not for simple audio file conversion, volume adjustment, or transcribing/normalizing an existing audio file (that is plain audio processing, not generation). The model is AceStep Audio 1.5 — it turns a music-description (style/genre tags) plus optional lyrics into a generated MP3. You may specify any of theme, music length, style, lyrics — one or more; if the user gives an image, the style and/or lyrics can be derived from it.
openclaw skills install @sunshinejnjn/music-with-comfyuiCall a user-defined ComfyUI server (local or remote) to generate music / audio with AceStep Audio 1.5 (text + optional lyrics → MP3). The user may specify any of theme, music length, style, and lyrics (one or more). If no style or lyrics are given explicitly, the agent derives them from the user's request — and from any image the user provides (see Image → Music).
The workflow expects several model files on the server (see Models and MODEL_URL.txt):
acestep_v1.5_xl_turbo_bf16.safetensorsqwen_0.6b_ace15.safetensors + qwen_4b_ace15.safetensorsace_1.5_vae.safetensorsconfig.json at the skill root holds every default; any value can be overridden by an env var (see the Data & Privacy table there).
The agent reads the image; the workflow reads only text. AceStep 1.5's
graph has no image input — the sole text node is TextEncodeAceStepAudio1.5
(tags + lyrics). When the user sends an image and does not also specify a
full style prompt and/or lyrics, let the agent derive the tags and lyrics,
then hand only that text to the script:
--tags (and, if the song has words, draft lyrics in the right --language).If the user does give explicit style and/or lyrics, honour those first and use the image only as a mood reference.
When to use this section:
This skill fires whenever the user asks to write, compose, arrange, or make a song / melody / music track, in any language — even if no genre or length is given. Common trigger phrases include, but are not limited to:
The signal is the intent to create new music (write / compose / arrange / make / generate a song/melody/track). Match it loosely across languages; if the intent is clearly "create a new song", use this skill.
--tags)This skill is for generation, not post-processing of existing audio. Do not use it (and do not look here) for:
COMFYUI_URL env var > comfyui_url in config.json; bundled default http://127.0.0.1:8188). A non-local host prints a data-flow warning on every run; if no server answers, the script exits with configuration instructions instead of a raw timeout.output_dir (config.json or COMFYUI_OUTPUT_DIR). Don't submit sensitive content if the endpoint or disk is a concern.| Mode | Workflow | File |
|---|---|---|
| Music (AceStep 1.5) | AceStep Audio 1.5 (text/lyrics → audio) | workflows/acestep_audio_api.json |
The workflow is an adapted version of a real ComfyUI graph. Original source + exactly what was changed → references/attribution.md.
⚠️ The music description is MANDATORY and must be passed via
--tags(or--prompt) — never as a positional word.
python3 music_with_comfyui.py music "lo-fi track"→music: error: the following arguments are required: --tagsAlways write the description through
--tags "..."on the first attempt. A positional word wastes a whole turn before the command ever reaches ComfyUI.
Full commands and the timeout table → references/cli-reference.md
--tags (required) — the music description: genre, tempo, instrumentation, mood, vocal style. This is the primary prompt. ComfyUI-escaped (line breaks as \t/\n); keep it clean.--lyrics (optional) — structured Sonic-lyrics text. Omit (or pass empty) for an instrumental; AceStep 1.5 generates instrumental audio.--duration (optional) — seconds (default: config default_duration, 120). The audio latent's seconds is set to match, so keep them consistent.--seed (optional) — not sent by default. If omitted, the script uses no seed (random per run). Pass --seed <n> only to lock/reproduce a specific output.--bpm, --keyscale, --language, --timesignature, --cfg-scale, --temperature, --top-p, --top-k, --min-p) — optional overrides of the defaults in config.json music.default_params; each omits cleanly if unset.--tags-image (removed) — AceStep 1.5 takes no image. If the user gives a reference image, the agent reads it and derives tags/lyrics itself (see Image → Music).--format / --output-dir — output format (jpg/png for the image skill; here audio format is whatever the server's SaveAudioAdvanced writes).Guidance for crafting AceStep prompts lives in references/prompts.md. Key points:
--tags as a comma/space-separated description of genre, tempo, instrumentation, mood, and vocal style. Keep it clean and natural.\\t (tab) / \\n (newline). If you write the tags through --tags "..." the script passes the string through, so keep line breaks to a minimum or escape them.[Verse 1], [Chorus], ...). Keep to the language set via --language.The script auto-detects the same categories as the image skill and reports them with install/source guidance (see references/error-handling.md):
git clone itself).UnloadAllModels) — bypassed automatically without interrupting generation.Send the generated file — do not just describe it. Deliver in the user's original session, never raw paths/URLs. Per-channel prefixes and the staging/cleanup routine → references/delivery.md
output_dir/music/ (prefix amusic_<timestamp>, e.g. amusic_20261006_153000.mp3).MEDIA: line for WhatsApp, a native attachment elsewhere). Do not send the raw file path to the group unless asked.ffmpeg -y -i <src> -c:a aac -b:a 128k -ar 44100 -ac 1 -movflags +faststart <out>.m4a
then send the .m4a via MEDIA: as a plain line. (mp3 also works if 44.1 kHz; 44.1 kHz is the key.).zip (e.g. zip -j out/<name>.zip <src>.mp3) and send the zip as a document attachment (MEDIA:<path>.zip, or message(action="send", channel=..., asDocument=true, media=<path>.zip)).A generated song is not just an audio clip: hand it over as a complete package, and do it all in one message.
--tags, and the lyrics — keep it short and evocative. State it clearly as the song's name.[Verse] / [Chorus] intact, no stray control characters or unescaped line breaks).Deliver all three together — title, lyrics, and audio — not in separate hand-offs:
🎵 **<Song Title>**
<style note — one line>
[Verse 1]
...
[Chorus]
...
MEDIA:/…/amusic_<timestamp>.mp3
Never omit the title or the lyrics: a finished song should arrive with its name and, when it has words, its lyrics, in the same message as the audio. If the user supplied a title, use theirs; if the song is instrumental, skip the lyrics but still give it a title.
The measured GPU sampling time is model-dependent (short, looped clips are fast; long durations cost seconds-to-tens-of-seconds of sampling). When a run stretches well beyond the configured timeout_seconds, the delay is almost always agent-side overhead before the run, not inference. Kill the waste:
SKILL.md, config.json, or the script for a known config / repeat / same-job. Defaults are baked in.timeout_seconds in config; anything past that is overhead. Never pile on tool calls before the audio exists.