Install
openclaw skills install @wangminrui2022/h3-video-editing-promptWrite or revise MiniMax H3 Ref2VA prompts for editing a source video from reference images or audio, especially masked or false-colour person replacement, face replacement, costume or object replacement, cross-shot identity persistence, overlay cleanup, and lip-sync preservation. Use when a request
openclaw skills install @wangminrui2022/h3-video-editing-promptWrite the target result, not a post-production command. In detailed_description, prefer carries the appearance of, is rebuilt as, or continues from; do not tell H3 to “swap”, “key”, or “remove” pixels.
This skill writes prompts. It must not imply that wording alone can guarantee a spatial mask, an untouched time range, or frame-accurate lip sync.
For uncommon target types, read references/universal-replacement-template.md. For copy-ready examples and pipeline failure handling, read references/editing-recipes.md.
Unless the user requests analysis or another format:
text block.subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:
The target video is an edited version of <Video 1>.outputs/ folder only when the current environment provides a workspace and the user wants a file.<Video N> identifies the source video or its whole-run structure, not the person inside it.<Subject N> identifies reusable visible content such as the source performer or reference identity.<Subject N>; do not add a standalone <Picture N> line unless the picture is a keyframe or composition anchor.<Audio N> only when an audio signal is actually available to the target workflow.<Theme N>, <Topic N>, Topic definition, Task, or Retain analysis.Task types are video editing, reference generation, audio reuse, audio reference, keyframe completion, and video continuation. Combine only the relationships that truly occur.
The mask identifies the target; it does not make the rest of the frame immutable.
not carried over.<Video 1> when they are visible or can be inferred from exposed joints.Opaque and translucent masks are different:
A negative, cyan/blue, green, or other false-colour overlay may corrupt the complete appearance while leaving the face, body contours, clothing folds, mouth shapes, and motion readable. Treat this as appearance-corrupted but geometry-readable, not as an opaque blank:
Define <Subject 1> as the same persistent marked character, not as an anonymous region:
<Subject 1> is the same marked target character wherever any part of that character appears in <Video 1>, including full, partial, cropped, edge-of-frame, foreground, and post-cut appearances.
For every measured cut, write a new [Shot N] At MM:SS.mmm paragraph and reassert that every visible portion of <Subject 1> carries the same <Subject 2> identity. A cut changes composition and visible area; it never resets the replacement assignment.
Prompt wording cannot guarantee cross-cut identity. If the output still restores the source character, leaks false colour, or flickers at a cut, generate each target-bearing shot separately and concatenate at the original hard cuts.
Use this as the default input contract when the user supplies three assets:
<Video 1>: the reference-condition video. Its coloured overlay identifies the only target region; its visible frames provide motion, pose, expression, gaze, mouth motion, timing, camera, lighting, contact, and occlusion.<Subject 1>: the masked carrier abstracted from <Video 1>, used for performance and geometry rather than retained appearance.<Subject 2>: the target person or object from <Picture 1>, used for the complete appearance inside the overlay.<Audio 1>: the separately supplied audio, only when its actual role and alignment have been verified.Do not assume every condition video matches the inspected sample. Measure each clip. For a full-run translucent overlay, no temporal segmentation is needed and all readable motion and facial evidence remain usable.
If a person-shaped mask is filled with an object of very different topology, preserve the carrier's centre, scale, orientation, path, and timing rather than claiming that the object reproduces human joints or facial motion.
<Video 1>.<Audio 1>, audio reuse, and fully_copy even when it is a separate file.audio reference; it is not an absolute lip-sync timeline.Use source frames and the synchronized final audio jointly. The frames own the visible articulation; the audio checks the opening, closure, sustained sound, pause, breath, and emphasis timing. The audio may be embedded in the source video or supplied separately, but it must share the same timeline. Do not make audio replace the source mouth trajectory.
Compact wording:
Her visible articulation follows <Video 1>: the upper- and lower-lip contours, mouth corners, jaw opening, visible teeth and tongue, full closures, pauses, and coupled head motion keep the source frames' timing. <Audio 1> is reused unchanged and validates the same openings, closures, sustained sounds, pauses, breaths, and emphasis. No new dialogue is added.
Mark <Audio 1> as fully_copy only when it is actually reused. Do not invent <d> content. Known dialogue may be preserved for semantics, but it must not define a hand-written mouth timeline.
Preserve the visible lip shapes and timing from <Video 1>, but do not claim audio-driven synchronization and do not define <Audio 1> in the H3 prompt. Reattach the untouched original track after generation. If precise sync is required, use a dedicated audio-driven lip-sync or face-reenactment stage after H3.
Her upper- and lower-lip contours, mouth corners, jaw opening, visible teeth and tongue, full closures, pauses, and coupled head motion keep the exact visible sequence and timing from <Video 1>. No audio-based synchronization is claimed in this generation pass.
The audio can provide approximate articulation timing, but the source video cannot supply erased mouth geometry. Say so in analysis; do not promise frame accuracy. For exact sync, plan a dedicated downstream lip-sync pass using the original audio.
Stop rewriting the prompt as if phrasing alone will fix it. Obtain the rendered clip, keep the original audio, and classify the error:
Do not hand-write a phoneme schedule and do not claim that another prompt revision will produce exact sync.
When the face is readable, cover only observable ownership:
Avoid emotion labels and inferred intent. Do not describe a person as singing, speaking, livestreaming, rehearsing, or thinking unless the source or user establishes it. Do not add new action, dialogue, or events.
This node receives frames without their audio track. Therefore:
audio reuse, <Audio 1>, and fully_copy from the H3 prompt;overall_soundscape: N/A and non_diegetic_music: N/A unless the user explicitly requests new audio and a later stage supports it;audio reuse is used for the final copied signal; audio reference is used only for characteristics such as timbre or style.<Subject 1> is a persistent target across shots rather than an anonymous region.[Shot N] paragraph, and partial/cropped post-cut appearances explicitly retain <Subject 2>.