OCR for photos and images using MinerU. Extract text from photographs, screenshots, camera captures, and image files with high accuracy. Features: image OCR for .png, .jpg, .jpeg, .webp files. VLM mode for complex visual content. Handles photos of documents, whiteboards, signs, receipts, and more. Multiple output formats (Markdown, HTML, JSON, LaTeX). Use when you need to: OCR a photo, extract text from an image, read text in a screenshot, digitize a photo of a document, OCR a camera capture. Use when asked: 'how do I OCR this photo', 'extract text from this image', 'I took a picture of a document', 'can my agent read text from photos', 'is there a skill for image OCR', 'turn this photo into text', 'read this screenshot'. Powered by MinerU (OpenDataLab, Shanghai AI Lab) with advanced image OCR capabilities. Supports English, Chinese, and multilingual text in images. Perfect for mobile document capture, receipt digitization, whiteboard notes, sign reading, and any scenario where you have a photo with text that needs to be extracted.

Install

openclaw skills install @mzlzyca/photo-ocr