Install
openclaw skills install @sdk-team/alibabacloud-video-prompt-architectGenerates structured, high-quality prompts for AI video and image generation models. Transforms natural language descriptions into optimized prompts adapted for 18 models including Happy Horse, Seedance, Kling, Pika, Midjourney, Recraft, FLUX, and more. Use when creating video prompts, image prompts, product images, posters, or adapting prompts across different AI generation models. Triggers: "生成视频提示词", "视频prompt", "文生视频", "图生视频", "文生图", "商品图", "海报生成", "AI生成提示词", "prompt architect", "media prompt"
openclaw skills install @sdk-team/alibabacloud-video-prompt-architectAutomatically decompose a user's natural language request and generate high-quality structured prompts adapted for different AI models.
User natural language input → Intent analysis → Structured decomposition → Mode adaptation → Self-review reflection → Output high-quality Prompt
Components: Intent Parsing Engine + Structured Prompt Generator + Multi-Mode Adapter + Negative Prompt Generator + Quality Self-Checker
Accept the user's natural language description, which can be a brief sentence, e.g.:
Based on user input, automatically determine the most suitable generation mode and select the target model according to user specification or default strategy:
| Trigger Keywords | Generation Mode | Default Model |
|---|---|---|
| Video, animation, motion, dynamic, clip | Text-to-Video | Happy Horse |
| Reference image, multi-character, image fusion video | Reference-to-Video | Happy Horse r2v |
| First frame, image-to-video, animate image | Image-to-Video | Happy Horse i2v |
| Image, photo, illustration, wallpaper, concept art | Text-to-Image | Nano Banana |
| Product, merchandise, e-commerce, showcase | Product Image Generation | seedream |
| Poster, promotion, advertisement, banner | Poster Generation | Midjourney |
| Model | Use Case | Language | Prompt Style |
|---|---|---|---|
| Happy Horse | Business custom videos, platform integration, lightweight creativity | Chinese (mandatory) | Structured Chinese, camera terms directly usable |
| Seedance | Short videos, narrative segments, motion shots | Chinese (mandatory) | Structured Chinese: subject + action + scene + camera + style |
| Kling | Cinematic videos, ads, drama segments | Chinese (mandatory) | Structured Chinese, emphasizing camera movement and visual texture |
| Wanx | Chinese-native video, general content, marketing videos | Chinese (mandatory) | Chinese structure, reduce abstract words |
| Veo | High-quality video, commercials, cinematic visuals | English | Complete description, emphasizing atmosphere, rhythm, scene details |
| Sora | Complex narratives, multi-character, physical consistency | English | Semantically coherent, clear subject relationships and action logic |
| Hailuo | Short videos, social content, quick production | Chinese (mandatory) | Concise and direct, emphasizing action, emotion, and style |
| Runway | Creative ads, brand content, stylized videos | English | Natural language + style directives, emphasizing brand consistency and visual style |
| Pika | Social media shorts, creative effects videos, viral content | English | Concise and dynamic, emphasizing creative transitions and special effects (melt/inflate/explode/crumble) |
| Model | Use Case | Language | Prompt Style |
|---|---|---|---|
| Nano Banana | Quick image generation, creative exploration, lightweight visuals | English | Concise and efficient: subject + style + composition |
| GPT Image | General images, design sketches, marketing graphics | English | Natural language, semantically clear |
| Grok Image | Creative images, social media visuals, personalized graphics | English | Natural description, emphasizing effects and themes |
| seedream | Commercial images, posters, high-quality pictures | Chinese (mandatory) | Structured Chinese: subject + scene + lighting + texture + composition |
| Qwen Image | Chinese design needs, general images, marketing visuals | Chinese (mandatory) | Chinese-organized, emphasizing purpose and style |
| Midjourney | Posters, concept art, stylized illustrations | English | Style + composition + material + lighting keyword combinations |
| FLUX | API integration, developer workflows, self-hosted | English | Structured English: subject + scene + style + quality, precise control |
| Ideogram | Posters/ads/covers, text rendering, typographic images | English | Natural language + text content directives, emphasizing layout and readability |
| Recraft | Vector graphics, brand design, icons/logos, print materials | English | Design-oriented: subject + style + color palette + output format, supports SVG/EPS vector output |
If the user does not specify a model, use the default model based on task type. Users can switch at any time. For detailed model adaptation rules, see
references/generation-modes.md
Decompose user input into the following 8 structured components, each generated independently:
Apply differentiated adaptation to the structured prompt based on the target model:
Video Model Adaptation Strategies:
[Image N] reference image syntax. Creative Analysis table in Chinese.Image Model Adaptation Strategies:
--ar, --style, :: weight syntaxFor detailed model adaptation rules, see
references/generation-modes.mdFor prompt templates, seereferences/prompt-templates.md
Perform self-checks on generated prompts before output to ensure quality standards:
| Check Item | Rule | Action on Failure |
|---|---|---|
| Intent Completeness | Are all key elements from the user's original description covered? | Add missing elements |
| Component Consistency | Are there semantic conflicts between the 8 components? (e.g., "minimalist style" vs "element-dense") | Resolve conflicts, prioritize user's core intent |
| Length Compliance | Does it meet the target model's prompt length range? Count ONLY the text inside the Positive Prompt code block — exclude markdown markers (```), section headers, labels like "Positive Prompt:", and Negative Prompt content. For English: count space-separated tokens. For Chinese: count characters (excluding punctuation marks and spaces). | If too long, trim by priority; if too short, add details. Report accurate count, not rough estimate. |
| Language Compliance | Does it use the language required by the target model? | Convert to correct language |
| Negative Prompt Validation | Do negative prompts contradict positive descriptions? | Remove contradictory items |
| Physical Plausibility | Are action/scene descriptions physically reasonable? (CRITICAL for video models). Check for: humans hovering/floating in mid-air without mechanical aid, defying gravity, impossible body positions, perpetual motion, objects passing through solid matter, humans remaining perfectly still for extended durations (e.g., 30s frozen), physically impossible camera movements, time-compressed natural processes. | MUST correct to plausible description AND output ⚠️ Note in Notes section to inform user of the correction and reasoning |
| Model Feature Adaptation | Does it follow the target model's field order and special syntax? | Rearrange per model specs |
Items that fail self-check must be corrected before output. If trade-offs remain after correction, place "⚠️ Note" in the
### 📝 Notessection BEFORE### ✅ Final Prompt— never after Final Prompt. ⚠️ Note is MANDATORY for any Physical Plausibility correction — never silently modify physically implausible descriptions. The Note must explain: (1) what was implausible, (2) how it was corrected, (3) any remaining limitations.
When the target is a video generation model, output recommended parameters in a ### 📊 Video Parameters table BEFORE ### ✅ Final Prompt. Never place parameters after Final Prompt.
| Parameter | Description | Example Values |
|---|---|---|
| resolution | Resolution | 720P / 1080P / 4K |
| ratio | Aspect Ratio | 16:9 / 9:16 / 1:1 / 4:3 / 3:4 |
| duration | Video Duration | 3-15 seconds (depending on model support) |
When user requirements involve multiple scenes, multiple shots, or storylines, automatically switch to multi-shot sequencing mode:
## 🎬 Multi-Shot Narrative — [Theme]
### Shot Overview
| Shot # | Duration | Scale | Core Content |
|--------|----------|-------|--------|
| Shot 1 | 3s | Wide | [Summary] |
| Shot 2 | 4s | Medium | [Summary] |
| Shot 3 | 3s | Close-up | [Summary] |
### 📝 Narrative Plan
**Transition Suggestions:**
- Shot 1 → Shot 2: [Transition method, e.g., fade/cut/camera-linked]
- Shot 2 → Shot 3: [Transition method]
**Narrative Consistency:**
- Subject Appearance: [Key features to maintain across shots]
- Lighting Continuity: [Ensure adjacent shots have consistent lighting]
- Color Palette: [Unified color scheme]
---
### ✅ Shot Prompts (copy-ready)
**Shot 1 — [Scene Name]:**
[Complete Prompt]
**Shot 2 — [Scene Name]:**
[Complete Prompt]
**Shot 3 — [Scene Name]:**
[Complete Prompt]
### ✅ Final Prompt MUST be the last section in output — nothing follows it### ✅ Final Prompt## 🎬 [Mode Name] | [Target Model]
> 📋 Model: [Model Name] | Language: [Chinese/English] | Suggested Length: [Range]
---
### 💡 Creative Analysis
| Component | Content |
|-----------|---------|
| **Subject** | [Specific description] |
| **Scene** | [Specific description] |
| **Camera** | [Specific description] |
| **Lighting** | [Specific description] |
| **Composition** | [Specific description] |
| **Action/Emotion** | [Specific description] |
| **Quality** | [Specific description] |
---
### 📝 Notes (optional, only when needed)
> Model-specific tips, reference image descriptions, or usage notes go here.
> This section MUST appear BEFORE Final Prompt, never after.
---
### 📊 Video Parameters (video models only)
| Parameter | Value |
|-----------|-------|
| Resolution | [720P / 1080P / 4K] |
| Ratio | [16:9 / 9:16 / 1:1] |
| Duration | [3-15s] |
> This section MUST appear BEFORE Final Prompt, never after.
---
### ✅ Final Prompt
> The following is the complete Prompt ready for direct copy-paste:
> **NOTHING is allowed after this section — no suggestions, notes, explanations, or any additional text.**
**Positive Prompt:**
[Combine all components into a coherent Prompt, ready to paste directly into the model]
**Negative Prompt:**
[Negative prompt content, only output for models that support negative prompts]
When the target model is HappyHorse, use the following enhanced format:
## 🐴 HappyHorse — [Mode Name]
> 📋 Model: `happyhorse-t2v` | Resolution: 1080P | Ratio: 16:9 | Duration: 5s
---
### 💡 Creative Analysis
| Component | Content |
|-----------|---------|
| **Subject** | [Specific description] |
| **Scene** | [Specific description] |
| **Camera** | [Specific description] |
| **Lighting** | [Specific description] |
| **Composition** | [Specific description] |
| **Action/Emotion** | [Specific description] |
| **Quality** | [Specific description] |
---
### 📝 Notes (r2v mode: describe reference images here)
**Reference Image Notes (r2v mode only):**
- [Image 1]: [Describe image content and purpose]
- [Image 2]: [Describe image content and purpose]
---
### ✅ Final Prompt
> The following is the complete Prompt ready for direct copy-paste:
> **NOTHING is allowed after this section — no suggestions, notes, explanations, or any additional text.**
**Positive Prompt:**
[Compose a coherent Chinese Prompt, leveraging HappyHorse's strong understanding of Chinese camera language. This MUST be in Chinese.]
**Negative Prompt (use "avoid..." guidance in positive Prompt):**
[Negative guidance embedded in positive Prompt using "避免..." phrasing, as HappyHorse does not support independent negative prompts]
references/generation-modes.md. Word/character count applies ONLY to the text inside the Positive Prompt code block: for English count space-separated words, for Chinese count characters (excluding punctuation and spaces). Do NOT include markdown syntax, section headers, labels, or Negative Prompt content in the count. Always verify the exact count before output — never report an approximate estimate.references/generation-modes.md[Image 1], [Image 2] in prompts to reference images--ar 16:9, --style raw, element::2 weight syntax(element:1.5) weight syntax### ✅ Final Prompt (or ### ✅ Shot Prompts (copy-ready) in multi-shot mode), NO additional content is allowed. This includes but is not limited to: suggestions, tips, notes, explanations, comparisons, follow-up questions, alternative versions, video parameters, model-specific recommendations, file save confirmations, execution logs, status messages, processing summaries, or any operational output. The prompt code blocks MUST be the absolute last content in the output. All auxiliary information (Notes, Video Parameters, Transition Suggestions, etc.) must appear BEFORE the Final Prompt section. If the agent needs to log information internally, it must do so silently without displaying anything to the user after Final Prompt.If the user expresses preferences during conversation, or demonstrates consistent style tendencies across multiple uses, remember and automatically apply in subsequent generations:
| Configuration | Description | Example |
|---|---|---|
| Default Video Model | User's preferred video generation model | Seedance |
| Default Image Model | User's preferred image generation model | Midjourney |
| Preferred Style | User's commonly used visual style | cinematic / Chinese traditional / minimalist / cyberpunk |
| Common Aspect Ratio | User's most common aspect ratio | 16:9 / 9:16 / 1:1 |
| Language Preference | Prompt language preference | Chinese-first / English-first |
| Style Keywords | User's frequently used core style words | cinematic quality, warm tone, shallow DOF |
| Strategy | Description |
|---|---|
| Default Version | Always use the latest stable version of each model |
| User-Specified Version | Support user-specified versions (e.g., "Midjourney", "HappyHorse") |
| Version Feature Differences | When different versions have different prompt rules, generate per specified version rules |
| New Version Release | Update references/generation-modes.md to add new version adaptation rules |
| Version Parameters | Midjourney uses --v to specify version; other models specify via model name |
| Resource | Path | Description |
|---|---|---|
| Prompt Template Library | references/prompt-templates.md | Detailed templates and keyword libraries for each mode |
| Generation Mode Guide | references/generation-modes.md | Detailed adaptation rules for each generation mode |
| Usage Examples | references/examples.md | Complete input/output examples |
| Acceptance Criteria | references/acceptance-criteria.md | Skill quality acceptance criteria |