Install
openclaw skills install @permew/wan-3-0-prime-reference-to-videoWan 3.0 Prime Reference to Video generates video clips from reference images, reference videos and reference audio on RunComfy. Wan 3.0 Prime Reference to Video binds up to 10 reference images, 5 reference videos and 5 reference audio clips to a prompt that names them as Image 1, Video 1 and Audio 1, so a character, product or location stays consistent across a 2 to 30 second shot at 480p, 720p or 1080p with a synchronized audio track. Wan 3.0 Prime Reference to Video runs on the fast Wan 3.0 Prime tier (wan3.0-video-prime) and is billed per counted second, where reference videos add duration and reference images and audio do not. This skill documents the full Wan 3.0 Prime Reference to Video input schema, pricing, prompting patterns and the routing rules for Wan 3.0 Prime text-to-video, Wan 3.0 Prime image-to-video, Wan 2.7 and Seedance 2.0 Pro. Calls runcomfy run wan-ai/wan-3.0-prime/reference-to-video through the local RunComfy CLI. Triggers on "wan 3 prime reference to video", "wan 3.0 prime", "wan3 prime", "reference to video", "ref2v", "keep the same character across shots", "video from reference images", or any explicit ask to generate video from references with Wan 3.0 Prime.
openclaw skills install @permew/wan-3-0-prime-reference-to-videoruncomfy.com · Wan 3.0 Prime Reference to Video · CLI docs
Wan-AI Wan 3.0 Prime Reference to Video builds a clip from a prompt plus image, video and audio references, on the fast Prime tier (wan3.0-video-prime), hosted on the RunComfy Model API.
openclaw skills install @permew/wan-3-0-prime-reference-to-video
Wan 3.0 Prime Reference to Video is the reference-conditioned endpoint of Wan-AI's Wan 3.0 Prime family. Instead of describing a subject in prose, you attach the subject as reference media and then address it inside the prompt by number: Image 1, Video 1, Audio 1, following the order of the arrays you passed. That numbered binding is what keeps a face, a costume, a product's geometry or a location steady through the shot, and it is why Wan 3.0 Prime Reference to Video exists as a separate endpoint from plain text-to-video.
The Prime tier targets the same Wan 3.0 visual quality with faster inference. The reference workflow matches standard Wan 3.0 reference-to-video.
Ideal for: character consistency, branded product scenes, multimodal storytelling.
| You want | Use |
|---|---|
| Same character / product / set across a shot, driven by references | Wan 3.0 Prime Reference to Video ✓ |
| Many references at once (10 images + 5 videos + 5 audio) | Wan 3.0 Prime Reference to Video ✓ |
| A reference-guided clip longer than 15s (up to 30s) | Wan 3.0 Prime Reference to Video ✓ |
| Prompt only, no reference media | Wan 3.0 Prime text-to-video |
| Animate one still, optionally toward a last frame | Wan 3.0 Prime image-to-video |
| Lip-sync to a voiceover track you already have | Wan 2.7 (audio_url) |
| Cinematic multi-modal short-form with in-pass speech | Seedance 2.0 Pro |
| Open-weights reference-to-video alternative | MiniMax H3 Open reference-to-video |
If the user said "Wan 3 Prime", "Wan 3.0 Prime", "reference to video" or "ref2v" explicitly, route to Wan 3.0 Prime Reference to Video regardless.
npm i -g @runcomfy/cli (or npx -y @runcomfy/cli --version)runcomfy login opens a browser device-code flow.RUNCOMFY_TOKEN=<token> instead of runcomfy login.wan-ai/wan-3.0-prime/reference-to-video| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
prompt | string | yes | — | Up to 20,000 chars. Scene, subject, motion, camera, lighting, style. Name references as Image 1, Video 1, Audio 1. |
reference_images | array | conditional | example image | Up to 10. Subject / object / scene consistency. |
reference_videos | array | conditional | [] | Up to 5, MP4 or MOV, 1–15s each, 15s total. Motion or scene guidance. |
reference_audios | array | conditional | [] | Up to 5, 15s total. Guides sound or timing. |
resolution | enum | no | 720p | 480p, 720p, 1080p. |
aspect_ratio | enum | no | 16:9 | adaptive, 16:9, 9:16, 1:1, 4:3, 3:4. |
duration | int | no | 5 | 2–30 whole seconds. |
prompt_extend | bool | no | true | Model rewrites the prompt for richer detail. Off = literal + faster. |
enable_audio | bool | no | true | Output carries a synchronized audio track. Off = silent clip. |
seed | int | no | random | 0–2147483647. Reuse for reproducible variants. |
At least one of reference_images, reference_videos, reference_audios is required. Wan 3.0 Prime Reference to Video rejects a prompt-only call; if the user has no reference media, route to Wan 3.0 Prime text-to-video instead.
Wan 3.0 Prime Reference to Video bills per counted second = output duration plus the combined duration of every reference video attached. Reference images and reference audio are not billed as duration, and toggling enable_audio does not change the rate.
| Resolution | Rate per counted second |
|---|---|
| 480p | $0.0624 |
| 720p | $0.124 |
| 1080p | $0.249 |
Worked examples: a 5s 720p clip with image references only = 5 counted seconds ≈ $0.62. The same clip with a 10s reference video attached = 15 counted seconds ≈ $1.86. A 30s 1080p clip with no reference video ≈ $7.47.
Two practical consequences: trim reference videos to the shortest clip that carries the motion, and draft at 480p (about 4× cheaper per second than 1080p) before the final render. The figure shown before submit is an estimate — reference durations are measured after the run, so the final charge settles then.
Default (image reference, 5s, 720p, 16:9, audio on):
runcomfy run wan-ai/wan-3.0-prime/reference-to-video \
--input '{
"prompt": "Image 1 walks slowly through a sunlit botanical garden, pauses beside a glass pavilion, then turns toward the camera with a relaxed smile; soft dappled light, gentle handheld motion, cinematic.",
"reference_images": ["https://.../subject.webp"]
}' \
--output-dir <absolute/path>
Cheap draft pass (480p, short, literal prompt):
runcomfy run wan-ai/wan-3.0-prime/reference-to-video \
--input '{
"prompt": "Image 1 rotates slowly on a marble pedestal, a highlight sweeps across the glass, soft studio bokeh behind.",
"reference_images": ["https://.../perfume-bottle.jpg"],
"resolution": "480p",
"duration": 3,
"prompt_extend": false
}' \
--output-dir <absolute/path>
Multi-modal (images + motion reference + audio reference), vertical, 1080p:
runcomfy run wan-ai/wan-3.0-prime/reference-to-video \
--input '{
"prompt": "Image 1 wearing the jacket from Image 2 crosses the rain-slick street from Video 1; camera dollies forward, neon reflections shimmer. Match the pacing of Audio 1.",
"reference_images": ["https://.../actor.jpg", "https://.../jacket.jpg"],
"reference_videos": ["https://.../street-plate.mp4"],
"reference_audios": ["https://.../rhythm-ref.mp3"],
"aspect_ratio": "9:16",
"duration": 8,
"resolution": "1080p",
"seed": 12345
}' \
--output-dir <absolute/path>
The CLI submits the request, polls it, fetches the result, and downloads *.runcomfy.net / *.runcomfy.com URLs into --output-dir. Ctrl-C cancels the remote request before exit.
Name references by number. Image 1, Video 1, Audio 1 follow the array order you passed. This is the core mechanic: "Image 1 stands beside the counter" beats a paragraph describing the person's face, and it beats "the man in the reference" once more than one reference is attached.
Split stable identity from evolving action. Face, costume, product geometry, brand mark, set → references. Motion, camera, mood, lighting, weather → prompt. Describing stable identity in prose burns characters and drifts.
Front-load shot grammar. "Slow forward push", "camera dollies forward", "slow subtle push-in", "handheld", "seen from above" all land as directives. Then state one primary action, not four competing ones.
prompt_extend is on by default. Short prompts get auto-enriched, which usually helps. Turn it off when the prompt is already precise, when brand copy must stay verbatim, or when you want a shorter turnaround.
Ladder the duration. Lock the motion at 2–5s, then raise toward 30s once the shot reads right. Duration and resolution are the two cost multipliers.
aspect_ratio: "adaptive" lets the output follow the reference framing instead of forcing 16:9 — useful when references are already vertical or square.
Anti-patterns:
A rugged Atlantic coastline at sunset seen from above; slow forward push
as waves roll onto dark rocks, warm clouds drift across the sky, soft
golden light, cinematic, smooth motion.
A rain-slicked European city street at night, neon signs reflecting in the
wet cobblestones; the camera dollies forward as a tram glides past,
reflections shimmer, moody cinematic lighting.
A luxury perfume bottle on a marble pedestal; it rotates slowly as a
highlight sweeps across the glass, soft studio bokeh behind, clean
premium product look, subtle motion.
| Use case | Why this model |
|---|---|
| Character continuity across shots | Up to 10 image references, addressed by number |
| Branded product scenes | Product geometry held by reference, motion driven by prompt |
| Multimodal storytelling | Image + video + audio references in one call |
| Longer reference-guided clips | 2–30s, past the 15s ceiling of most siblings |
| Cost-tiered iteration | 480p drafts, 1080p finals, same prompt and seed |
| code | meaning |
|---|---|
| 0 | success |
| 64 | bad CLI args |
| 65 | bad input JSON / schema mismatch (no reference supplied, duration outside 2–30) |
| 69 | upstream 5xx |
| 75 | retryable: timeout / 429 |
| 77 | not signed in or token rejected |
Full reference: docs.runcomfy.com/cli/troubleshooting.
runcomfy run wan-ai/wan-3.0-prime/reference-to-video --input '<json>' --output-dir <dir>..runcomfy.net / .runcomfy.com URL in the result is downloaded into --output-dir.Ctrl-C cancels the in-flight request before billing.runcomfy login writes the API token to ~/.config/runcomfy/token.json with mode 0600 (owner-only). Set RUNCOMFY_TOKEN to bypass the file entirely in CI / containers. The skill reads no other environment variable, no shell history, and no system files.--input. The CLI does not shell-expand it; the body goes to the Model API over HTTPS. No shell-injection surface from prompt content.model-api.runcomfy.net (request submission) and *.runcomfy.net / *.runcomfy.com (download allowlist for generated output). No telemetry, no callbacks, no remote scripts piped into a shell.