Install
openclaw skills install skills-sh:minimax-ai/minimax-h3/minimalist-product-ad-generatorMinimalist Product Ad Generator Use this Skill to guide non-professional users through a flow-style workflow that creates a minimalist product advertising video for e-commerce promotion and product launches. The target users are e-commerce sellers, small brand owners, indie…
openclaw skills install skills-sh:minimax-ai/minimax-h3/minimalist-product-ad-generatorUse this Skill to guide non-professional users through a flow-style workflow that creates a minimalist product advertising video for e-commerce promotion and product launches. The target users are e-commerce sellers, small brand owners, indie creators, and individual sellers; the user should provide at least one product image or related asset.
The core principle is: confirm assets and brief first, build product facts and a product narrative spine, lock product visuals with independent anchor photos, control video with a precise beat storyboard, then finish the film with native audio or music-based editing. This Skill no longer defaults to a 4-panel anchor sheet, because video models may reproduce the panel layout. By default, it uses three separate anchor photos plus one precise beat text storyboard as the video control system.
Every time this Skill is triggered, run a complete but lightweight start gate before any analysis, copywriting, anchor generation, or video generation. Confirm only the information required for production. Do not ask about music here; music is handled later in the music step.
The start gate must confirm all of the following in one pass:
Product images and related materials
Product variant status
Target duration
Aspect ratio
Apple-style template
In-frame copy
Video model strategy
Start gate is mandatory. If the user has already uploaded product material, only skip the upload request; do not skip the rest of the start gate. The agent must still confirm product variant status, target duration, aspect ratio, Apple-style template, and in-frame copy mode in one pass before analysis, anchor generation, storyboard, or video generation. The video model is not a start-gate question: use MiniMax-H3 by default and display it in the completed start-gate result. If the user says “you decide / just do it,” the agent may choose recommended defaults for the remaining fields, but must still display the adopted choices as the completed start-gate result before continuing.
Independent anchor photos are the default visual control system
In-frame copy must appear as integrated video motion, not a subtitle fallback
Product body color is a hard fidelity constraint
Every product needs a product-specific narrative
Progress must be visible
SF Pro Display Semibold in prompts.If no product material is available, ask the user to upload it. If material is already available, use it and analyze:
If the image quality is not usable, stop and give concrete reshoot advice.
Output a short product fact summary:
Turn the start gate answers into an executable production brief. Do not ask again for parameters that have already been confirmed. The brief drives the narrative spine, copy, anchor sheet, and storyboard.
The production brief must include:
Multi-variant product rules:
Before copywriting and anchor generation, choose a lightweight product narrative spine. Do not use a broad corporate brand-film structure; this Skill focuses on short Apple-style films for physical products.
If the user has not selected a direction, offer 2-3 concise directions, recommend one, and continue after confirmation. If the user says “you decide,” use the recommended one.
Recommended spines:
Product Launch (default recommendation)
Feature Touch
Color Family
Output one sentence describing the selected narrative spine, for example:
This film uses the Color Family spine: the purple main variant establishes the hero view, other colors enter as supporting layers, and the full product set closes with the English copy.
Before copywriting and anchor generation, define the motion language of the film. This step does not generate assets; it defines motion intensity, transition logic, and rhythm peaks for the later storyboard.
Rules:
Transitions are driven by real product or visual elements
One main action per beat
Set strong and quiet moments
Keep safe space clear
Avoid fake tech decoration, mirrored white stages, and empty openings
Output one short “motion language” statement, for example:
This film uses product-edge highlights and the purple main variant sliding motion to drive transitions. The main peaks are variant entrance and full-copy landing, with a stable final hold for the product and copy.
Before generating the text anchor and storyboard, obtain a final English copy line for the frame.
Rules:
After the user confirms the copy, proceed to anchor generation. If the user says “you decide,” choose the recommended copy and continue.
Generate three independent anchor photos in the same aspect ratio and resolution as the user-selected video setting. They must be three separate image outputs, not one combined 3-panel or 4-panel sheet. Background follows the selected style, and all three photos must share unified light, shadow, grade, and background language while remaining standalone product photos.
Anchor photo roles:
Multi-variant handling:
Final copy anchor typography rules:
Show the three independent anchor photos and ask the user to approve, edit, or regenerate a specific photo before proceeding to the precise beat storyboard.
Before generating any video clip, create and read back a precise beat text storyboard table. It is not artwork; it is the execution table for the video model.
The storyboard table must be organized as “User Choice Statement → Beat Content → Principles.”
Write the confirmed style, aspect ratio, duration, narrative spine, main variant, product main color, variant strategy, and copy. Example:
White-tech style, 16:9, 10 seconds, Color Family spine, purple as the main variant, other colors as supporting layers, copy: Color Meets Sound.
Recommended columns:
| Time range | Shot | Shot purpose | Visual lead | Visual / camera move | Variant state | Copy | Text color | Text effect | Transition / continuity | Rhythm intent |
|---|
Table rules:
Rhythm intent vocabulary: setup, establish, prepare, impact, brake, settle.
After the user confirms the storyboard table, proceed to video generation.
Generate video directly from the confirmed three independent anchor photos + precise beat text storyboard table. Video style, background, lighting, mood, and pacing follow the selected style. Aspect ratio and size must strictly follow the user-selected setting; if the user chose 16:9, generate 16:9 and do not auto-change to another ratio.
Video reference rules:
The video prompt must include:
Video generation defaults to MiniMax-H3 native audio first, and the Apple-style tech BGM direction below should be written into the video prompt. If H3's native music is unpleasant, too loud, too weak, out of sync, or the user asks for new music, use music-2.6 to generate a longer standalone instrumental BGM and replace the video audio track. Do not default to ElevenLabs; use ElevenLabs only when the user explicitly requires a strict target duration, and follow the fee-confirmation rule.
Default music direction:
Around 100 BPM, fast tech feeling, block chords, pluck and airy noise bed, kick + sub-bass + sine sweep, wooden percussion, sudden cut-off, within 0.5s everything stops, leaving only the pluck tail to decay.
Extended interpretation:
Before final delivery, edit or assemble based on music analysis. Do not simply lay music under the video.
Music analysis:
Audio assembly rules:
Pre-delivery verification:
Delivery includes:
music-2.6, do not force duration, and cut the best segment from the longer track.Use this Skill for requests like:
Do not use this Skill for general video editing, documentary explainers, KOC talking-head ads, or complex UI/screen-text demos unless the user explicitly wants the Apple-style product ad workflow.
d21241f0a4b3