Multimodal Base

Supports image understanding, OCR, speech-to-text, and text-to-speech synthesis with multi-voice and multimodal unified processing using OpenAI and Edge TTS.

Install

openclaw skills install @yuyonghao-123/yuyonghao-multimodal-base