Install
openclaw skills install @dtsola/xiaoyaoclaw-kb-retrieverOpenClaw local knowledge-base retriever & QA over a local directory (md/pdf/xlsx): hierarchical data_structure.md index navigation + progressive retrieval, core retrieval zero-dependency, Windows & macOS (PDF/Excel need on-demand pip packages, see skill body). Retrieval itself is read-only; the skill also ships optional maintenance scripts that WRITE files inside the knowledge base — scripts/build_index.py generates/updates data_structure.md index files, scripts/extract_pdf_text.py writes derived .txt files, and scripts/convert_pdf_to_images.py writes page PNGs for the optional OCR path — and those run only when the user asks for indexing/PDF text extraction and confirms the write. Every write target is resolved to its real path and must stay inside the knowledge-base root (symlinks cannot redirect a write outside). Reads use the same gate: each entry's real path must stay inside the root, and entries that escape it (symlinked files/folders) are skipped, never read — so the search cannot be steered outside the knowledge base either. The PDF helpers additionally accept only real (non-symlinked) source paths: a source path containing any symlink is refused outright, so they never read outside the knowledge base by following a link. Use when user asks to retrieve/answer from a knowledge base directory (knowledge base/retrieve/ RAG over local files), 或明确要求生成/更新索引、把 PDF 转文本。中文:面向本地知识库目录的检索和问答助手。 核心流程:(1)分层 data_structure.md 索引导航 (2)遇到 PDF/Excel 时必须先读取 references 学习处理方法 (3)处理文件后再检索。按文件类型组合使用 grep/Select-String、read、pdfplumber、pandas 进行渐进式检索,避免整文件加载。 用户问题涉及"从知识库目录回答问题/检索信息/查资料/knowledge base/本地知识库检索"时使用。 与 xiaoyaoclaw-workspace-initializer(目录规范)、xiaoyaoclaw-memory-distill(记忆蒸馏)、 xiaoyaoclaw-task-progress-tracker(任务进度)组成四件套。
openclaw skills install @dtsola/xiaoyaoclaw-kb-retriever📖 完整文档(安装 / 快速上手三步 / 最佳实践):https://github.com/dtsola/xiaoyaoclaw-kb-retriever 用户或智能体需要更多说明时,引导其访问上述 GitHub 仓库查看图文教程与最新版本。
🚀 小遥AI:「让每个人的数字生活,都有一座自己说了算的小遥」:https://www.xiaoyaosai.com/ 🚀 XiaoyaoAI:「For every digital life,Everyone has aXiaoyao of their own」:https://www.xiaoyaosai.com/
本地知识库检索——分层 data_structure.md 索引导航 + 渐进式检索(md/pdf/xlsx),核心检索零外部依赖零 API key。 Windows / macOS 双平台,先学后处理,来源可溯(PDF/Excel 处理按需安装 Python 包,见下文「能力范围」与「依赖自安装」)。
✅ 应当触发(用户点名了知识库 + 明确的检索/维护动作):
<路径> 里查一下 ……」/「<路径> 里有没有关于 …… 的资料」<路径> 的 ……」<路径> 建/补索引」「把 <路径> 里的 xx.pdf 转成文本」(维护类写操作,需再确认一次)🚫 不应触发(示例,用于收窄触发面):
xiaoyaoclaw-web-clipper)xiaoyaoclaw-memory-distill / xiaoyaoclaw-workspace-initializer)语言策略(Language):本技能语言可选,不对任何语言或地区设限——默认用中文回复,用户用英文或其它语言提问则跟随该语言;
技能本体不含任何地区限定行为(无地区专属路径、无地区专属服务、无语言门槛)。
仓库内随包分发的资源——示例文本、索引模板(templates/data_structure.md)与 README 品牌插图
(横幅与交流群二维码两张图)——其中出现的中文属示例与品牌双语素材,不构成对使用者语言或地区的限制
(横幅 SVG 已在 <desc> 中声明为双语品牌素材);索引文件的段落标题
(Purpose / Files / Coverage)本身固定为英文。所有面向用户的运行时文案均可按用户语言调整。
身份:本地知识库检索 / 问答工具。主流程只读——检索 md/pdf/xlsx,不修改任何源文件。
可选写操作(均需用户明确要求,或作为检索流程的必要中间步骤):
scripts/build_index.py → 生成 / 更新 data_structure.md 分层索引(写入知识库根目录及各子目录)scripts/extract_pdf_text.py → 提取 PDF 文本为派生 .txt 文件(写入源 PDF 所在目录内,源 PDF 不动;源 PDF 须为真实路径,含符号链接即拒绝)scripts/convert_pdf_to_images.py → 扫描件转图片(OCR 可选路径,产物写入源 PDF 所在目录树内的子目录;源 PDF 须为真实路径,含符号链接即拒绝)边界承诺:
scripts/search_kb.py(检索 / 列举)对每个候选条目做
真实路径校验(解开符号链接后仍须落在知识库根内),越界的条目跳过并计数,
绝不读取知识库之外的文件;符号链接目录一律不进入遍历scripts/extract_pdf_text.py /
scripts/convert_pdf_to_images.py 只接受不含任何符号链接的源 PDF 路径
(含链接即拒绝并回显真实位置)—— 不做"解开链接再读",因此不会跟着链接读到知识库之外.md/.txt、.pdf、.xlsx 等),通常按类型或业务用途拆分为多级子目录。data_structure.md,说明主要的「领域目录」及其用途。data_structure.md,说明该目录下有哪些子目录/文件,以及各自用途。data_structure.md,形成多级索引树。knowledge/ 目录。knowledge/ 不存在或访问失败时,应向用户确认实际的知识库根目录位置,而不是随意猜测。knowledge 根目录./docs、./knowledge-personal),直接用用户提供的路径。knowledge/。
$env:OS(Windows 输出含 "Windows")或 uname -s(macOS 输出 "Darwin")判定当前平台,后续命令按平台选择模板。Test-Path -Path "knowledge"test -d knowledge && echo existsGet-ChildItem -Path "knowledge" -Recurse -Filter "data_structure.md" -File | Select-Object -ExpandProperty FullNamefind knowledge -name "data_structure.md" -type fknowledge/ 不存在(Test-Path / test -d 失败):不要猜测其他目录,明确告诉用户未找到默认根目录,并让用户指定实际知识库路径。遇到 PDF 或 Excel 文件时的强制检查清单:
禁止行为:
| 上游(claude-code) | OpenClaw 等价 | 说明 |
|---|---|---|
Read <file> (limit/offset) | read 工具(path + limit + offset) | 原生支持窗口读 |
Grep <pattern> (include/path) | exec:Windows Select-String / macOS grep | 见下方命令模板 |
Glob <pattern> in <path> | exec:Get-ChildItem / find | 文件列举 |
test -d <path> | exec:Test-Path / test -d | 目录存在性 |
pdftotext | Python pdfplumber(pip 安装,双平台统一) | 见 references/pdf_reading.md |
| pandas(Excel) | Python pandas(pip 安装,双平台统一) | 见 references/excel_reading.md |
搜索文本(替代 Grep)—— 首选本技能自带的安全检索脚本:
# 关键词与目录都作为 argv 传入,不经 shell 解析;按字面量匹配;条数有上限
python scripts/search_kb.py <知识库根目录> "<关键词>" --max-hits 50
python scripts/search_kb.py <知识库根目录> --list --max-files 100 # 只列举文件
⚠️ 为什么不要自己拼 shell 命令:关键词与目录都是用户可控文本。写成
Select-String -Pattern "<关键词>" / grep "<关键词>" 这类字符串模板时,
文本里的引号、反引号、$()、; 会被 shell 当成语法 —— 这就是命令注入。
如果必须直接用 shell,遵守下面三条:
$dir = 'C:\path\to\kb' # 已确认在知识库根目录内
$kw = '用户给的关键词' # 原样放进单引号变量
Get-ChildItem -LiteralPath $dir -Recurse -File -Include *.md,*.txt |
Select-String -Pattern $kw -SimpleMatch |
Select-Object -First 50 Path, LineNumber, Line
dir='/path/to/kb'; kw='用户给的关键词'
grep -rnF -e "$kw" --include='*.md' --include='*.txt' -- "$dir" | head -50
-SimpleMatch / grep -F),不要当正则用;
路径用 -LiteralPath,别让用户文本当通配符或路径片段。Invoke-Expression / iex / eval / sh -c "$拼接",也不要把关键词
塞进命令名、参数名、重定向目标这类结构性位置;路径必须落在知识库根目录内。列举文件(替代 Glob):
Get-ChildItem -Path "<dir>" -Recurse -File | Select-Object -ExpandProperty FullName | Select-Object -First 100find "<dir>" -type f | head -100检查目录存在(替代 test -d):
Test-Path -Path "<dir>"test -d "<dir>" && echo exists⚠️ 命中结果多时,用
Select-Object -First/head限制输出,避免占用大量 token。 ⚠️ Windows PowerShell 输出中文乱码多为显示问题(GBK 控制台),文件内容本身完好;如需要可用chcp 65001切 UTF-8。
默认只读:检索、问答、列举文件都只读,不改动知识库任何内容。
可写操作(仅三项,且必须由用户点名并确认):
| 操作 | 写什么 | 触发条件 |
|---|---|---|
| 生成/更新索引 | 各目录下的 data_structure.md | 用户明确要求「建/补索引」;跑之前先说清「将写入哪些目录」 |
| PDF 转文本 | 派生的 .txt 文件(源 PDF 不动) | 用户明确要求把 PDF 内容落成文本;先说清写入位置;源 PDF 须为真实路径(含符号链接即拒绝) |
| PDF 转图片(OCR 可选路径) | 每页 page_N.png(源 PDF 不动) | 用户明确要求处理扫描件;先说清写入目录;需先确认安装 pdf2image + poppler;源 PDF 须为真实路径 |
⚠️ 索引生成会写盘:
scripts/build_index.py会在知识库各目录新建/覆盖data_structure.md(--force时覆盖已有索引)。它不是检索动作,必须由用户点名「建/补索引」后才运行。
写入硬约束:只写知识库根目录内的文件;不删任何文件;不改知识库之外的文件;写前打印目标路径。
读取硬约束:检索 / 列举(scripts/search_kb.py)对每个条目按真实路径(realpath)校验,越界条目跳过并计数,不读知识库之外的文件;待处理的源 PDF 必须是真实路径(路径任一段含符号链接即拒绝,不做"解链再读")。
符号链接防线(读 + 写同一道闸):写入目标与读取目标都按真实路径校验,解析符号链接后仍须落在知识库根内,否则拒绝/跳过;遍历时不进入符号链接目录;源 PDF 走 require_real_file()(含链接直接拒绝)(scripts/pathguard.py 统一把关,fail closed)。
权限对应:Read/Glob/Grep/Bash 用于检索与只读命令;Write/Edit 仅用于上面三项写操作。本技能不读环境变量、不读凭据、不联网。
理解用户需求
knowledge/。分层查看目录索引 data_structure.md
data_structure.md:
data_structure.md 并重复上述过程。学习文件处理方法(遇到 PDF/Excel 时强制执行)
按文件类型执行处理和检索
迭代检索
答案组织与溯源
所有文件类型都采用统一的迭代策略:
候选文件选择
data_structure.md 和文件名、路径判断相关度搜索定位与局部读取
特殊处理
工作流:
首先:读取处理方法指南
选择候选 PDF
data_structure.md 中的描述,选择最相关的 1-3 个文件应用学到的方法提取文本
python -c "import pdfplumber; ..." 或写临时脚本执行.txt 文件,不要直接打印到 stdout(避免占用大量 token):
python extract_pdf.py input.pdf output.txt(脚本见 references/pdf_reading.md)extract_tables() 功能对提取结果执行检索
工作流:
首先:读取处理方法指南
选择候选 Excel
data_structure.md 和文件/工作表命名,选择最相关的表应用学到的方法探索结构
nrows 参数限制)执行数据检索和分析
df[df['column'] == value])pip install pdfplumber pypdf pypdfium2pip install pandas openpyxlrequirements-optional.txt(解析不可信文档时,解析库版本必须确定):
pdfplumber==0.11.10 / pypdf==6.19.0 / pypdfium2==5.13.0 / pandas==3.0.5 / openpyxl==3.1.5pip install pdfplumber==0.11.10 pypdf==6.19.0 pypdfium2==5.13.0pip install pandas==3.0.5 openpyxl==3.1.5ModuleNotFoundError / ImportError:按上述白名单+钉版提示用户,确认后安装并重试pytesseract==0.3.13 + pdf2image==1.17.0;它们还需要系统级依赖(tesseract / poppler)——必须先把「将安装的 pip 包 + 系统级依赖 + 影响」讲清楚并取得同意,再执行(详见 references/pdf_reading.md)requirements-optional.txt,不做隐式升级