Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Use this skill whenever the user wants to describe, caption, summarize, or interpret the content of images, videos, or files. Supports single/multiple inputs, URLs, local paths, base64 (images only), and optional structured JSON output via response_format + system prompt.

Install

openclaw skills install @tridefender/glmv-caption-tunnel