Extract plain readable text from Word documents (.doc, .docx) using MinerU. Outputs Markdown (the closest plain-text format supported) for easy reading and processing. Features: quick text extraction from .docx without token (flash-extract). Full extraction for .doc and .docx with token. JSON output mode with dedicated text fields for true plain text. Language support for English, Chinese, and more. Use when you need to: get plain text from a Word file, extract readable content from .docx, convert Word to text, read a Word document as plain text. Use when asked: 'how do I get text from a Word file', 'extract plain text from docx', 'I want to read this Word document as text', 'can my agent convert Word to text', 'is there a skill for Word to text'. Built on MinerU by OpenDataLab (Shanghai AI Lab), an open-source document intelligence engine. Works with local files and URLs. Perfect for data pipelines, search indexing, NLP preprocessing, and any workflow that needs raw text content from Word documents.

Install

openclaw skills install @mzlzyca/doc-to-text