Back to skill

Security audit

Memory Palace

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its memory-management purpose, but it needs review because it stores personal data, can send memories to configured LLM endpoints, runs risky optional install logic, and has unsafe memory ID file handling.

Review before installing. Use it only if you are comfortable with long-term local storage of memories, avoid saving secrets or regulated personal data, disable or tightly configure LLM-enhanced operations unless you trust the provider endpoint, and avoid the optional postinstall Python setup until dependencies and model sources are pinned and the path traversal bug is fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
src/llm/subagent-client.ts:388
Finding

Stored Memory and API Credentials Can Be Transmitted to an Unrestricted Provider Endpoint

Content
View full analysis
{ const controller = new AbortController(); const timeoutId = setTimeout(() => controller.abort(), timeout * 1000); try { const response = await fetch(`${baseUrl}/chat/completions`, { method: 'POST', ...[truncated 3294 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
src/storage.ts:242
Finding

Path Traversal Through Unsanitized Memory Identifiers

Content
View full analysis
{ return fileLock.withLock(id, async () => { try { const filePath = this.getFilePath(id); const content = await fs.readFile(filePath, 'utf-8'); ``` ```ts // src/storage.ts:304-307 async deleteFile(id: string): Promise { return fileLock.withLock(id, async () => { const filePath = this.getFilePath(id); await fs.unlink(filePath).catch(err => { ``` ```ts // src/storage.ts:395-405 async permanentDelete(id: string): Promise { // Delete from both main directory and trash const filePath = this.getFilePath(id); const trashPath = this.getTrashPath(id); await Promise.all([ fs.unlink(filePath).catch(err => { if ((err as NodeJS.ErrnoException).code !== 'ENOENT') throw err; }), fs.unlink(trashPath).catch(err => { if ((err as NodeJS.ErrnoException).code !== 'ENOENT') throw err; ``` ```ts // src/manager.ts:188-190 async get(params: string | GetParams): Promise { // Handle backward compatibility - accept string or object const id = typeof params === 'string' ? params : params.id; const memory = await this.storage.load(id); ``` ### Technical Analysis Memory IDs received through public manager and CLI methods are concatenated with `.md` and passed to `path.join`. The implementation does not valid ...[truncated 2004 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/check-vector-deps.cjs:132
Finding

Unpinned Python Dependencies Installed by an npm Postinstall Script

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (116)

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · README.md (reported line 187)May include surrounding context.

md
| `store(params)` | Store a new memory |
| `get(id)` | Get memory by ID |
| `update(params)` | Update a memory |
| `delete(id, permanent?)` | Delete memory (soft by default) |
| `recall(query, options?)` | Search memories |
| `list(options?)` | List with filters |
| `stats()` | Get statistics |

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的核心能力是“持久化记忆管理”,即记录和检索用户长期记忆内容。实际代码的核心目的却是评估两种搜索方法(向量搜索 vs 文本搜索)的效果差异。虽然代码表面上涉及‘语义搜索’,但这是在固定测试数据上做实验评测,而不是实现一个可供助手使用的记忆系统。代码没有数据库、文件持久化、记忆写入/更新、历史对话读取、时间相关推理等实现;相反,它会加载 Hugging Face 上的 SentenceTransformer 模型、运行测试集、计算 precision/recall/MRR 并输出总结。这属于 materially different primary purpose,应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

声明描述的是面向用户记忆的持久化管理能力,应包含记忆存储、检索、语义搜索、时间推理或对话记忆回忆等核心行为。但该代码片段没有实现任何记忆写入、读取、索引、语义搜索或时间推理逻辑。相反,它接收已有测试结果对象,计算改进幅度、总体评分、显著性简化判断,生成对比报告和可视化文本,并在命令行中打印示例报告。这属于测试分析/报表工具,而非记忆管理技能,因此主用途与实际行为存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个面向对话场景的持久化记忆技能:根据用户提供的信息进行记忆存储、回忆和时间推理。所给代码并没有实现此类对话触发的记忆管理逻辑,而是在 scripts/ab-test/run-jarvis-test.ts 中执行测试流程:加载 testMemories 和 testQueries,向 /data/.subagent/.jarvis 写入测试数据,调用 store/recall 做检索验证,统计成功率、命中率、按难度分布的结果,并将最终测试报告写入文件。虽然代码涉及“存储和检索记忆”,与声明主题相关,但其主要目的明显是测试/评估底层记忆系统,而不是提供声明中的用户记忆管理能力。因此属于实质性描述不符。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是面向真实用户记忆的持久化管理能力,即保存、回忆用户信息和对话上下文,并支持语义搜索/时间推理。而该代码块并未实现面向用户的记忆管理逻辑,也没有体现语义搜索或时间推理本身;它只是一个测试 harness,通过调用抽象的 MemoryPalaceClient 接口来评估外部记忆系统的存储与检索效果。其主要目的明显是测试、评估和度量,而非提供记忆管理技能本身。因此属于主要用途与实际行为不一致的明显描述-行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述的是一个可管理和检索持久化记忆的技能,应当能够记录用户信息、保存项目状态,并基于语义或时间进行回忆。给定代码却只是静态测试查询集合和辅助筛选函数,主要服务于 A/B 测试或评估流程。它没有任何持久化存储、对话记忆写入、真实搜索、索引、推理或回忆历史内容的实现。虽然测试样本内容围绕“记忆查询”场景,与声明领域相关,但该代码块的实际功能是测试支撑,而非声明中的核心能力,因此属于明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明的核心能力是记忆管理与持久化存储,时间推理只是其中的辅助能力之一。但实际代码仅是 test-time-reasoning.ts 测试脚本,调用 createTimeReasoning 来解析时间查询和条件时间,并输出测试结果。代码没有任何与记忆写入、更新、删除、检索、语义搜索、用户资料保存、项目状态持久化相关的行为,也没有访问存储或历史会话数据。虽然声明中提到“支持时间推理”,这部分与代码一致,但该代码块的主要用途与声明的主要用途存在实质偏差,因此应判定为描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是面向用户场景的记忆持久化与回忆能力;而这段代码并未实现‘记住新信息’或通用的记忆管理流程,主要是在测试环境中对已有记忆系统做检索验证和失败案例分析。它会访问特定测试数据库路径、调用 stats()、执行预设查询并打印评估结果,这更像内部调试/评估工具,而不是所声明的用户功能本身。虽然代码确实使用了记忆检索相关组件,和“支持语义搜索”部分有一定关联,但其主要目的与声明的技能用途存在实质偏差,因此应判定为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

声明描述的是一个面向用户记忆持久化与回忆的技能,核心应包括保存用户信息/偏好、后续检索这些记忆,并可能结合时间信息进行推理。给出的代码却只是测试脚本:一方面可在本地用 sentence-transformers 对预置样本文档做向量相似度计算,另一方面可向本地 HTTP 向量服务发送健康检查、索引和搜索请求。虽然它与“语义搜索”有部分相关性,但其主要目的明显是验证向量检索服务是否工作正常,而不是执行记忆管理本身。代码没有持久化真实用户记忆、没有会话状态管理、没有时间推理、没有基于触发词进行记忆写入或回忆,因此与声明的主要用途存在实质性不符。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

声明描述的是一个供对话中使用的“持久化记忆管理”技能,核心用途应是保存、检索和回忆真实用户/项目记忆。该代码块却是独立的验证脚本,其主要目的不是在实际交互中管理记忆,而是测试安装是否正确、量化检索命中率、验证删除/恢复/统计功能,并最终清理测试数据。虽然代码确实调用了 store/recall/delete/restore/stats 等记忆相关接口,与“记忆系统”领域有关,但其主要行为是 QA/benchmark,而不是声明中的生产用途。此外,声明特别提到“支持语义搜索和时间推理”,本代码只做关键词命中式验证和常规 recall,没有展示时间推理实现。因此描述与实际代码存在实质性偏差。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明的核心用途是记忆管理与回忆历史信息,但提供的代码仅是一个时间模式解析器,负责把若干中文时间短语转换为结构化时间结果。虽然声明里提到“支持时间推理”,这与代码的一部分能力相符,但代码没有任何持久化、记忆检索、语义搜索、用户信息管理或对话历史管理逻辑。因此其主要行为与声明的主要用途明显不一致,应判定为描述与实现不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
87% confidence
Finding

声明描述的是较通用的长期记忆/个人信息记忆能力,重点在记住用户信息、项目状态和历史对话,并提到语义搜索与时间推理。实际代码则明显聚焦于‘experience’对象的管理:recordExperience 只创建 type='experience' 的记忆;getExperiences/getRelevantExperiences 只检索 experiences 位置下的经验;verifyExperience 会维护 verified、verifiedCount、effectivenessScore、usageCount 等经验验证指标;extractFromMemories 还会调用 LLM 从已有记忆中提炼经验。这说明其主要用途不是通用持久化记忆管理,而是经验沉淀与复用。描述中虽提到“完成任务后有可复用经验”与“支持语义搜索”,与部分行为相符,但没有准确反映代码的核心能力——经验验证/评分/提炼,也未体现其范围局限于 experience 类型,而非广义个人资料、偏好或全部对话记忆。因此存在实质性描述与行为不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose centers on durable memory management and retrieval of remembered user/project information. The supplied code instead defines ConceptExpanderLLM, which takes a search query and returns expanded keywords and domains. It uses predefined mappings, an LLM call with a JSON prompt, and a fallback term extractor. No persistence layer, memory CRUD, user-profile storage, conversation-history retrieval, or temporal reasoning appears in this code. While concept expansion could be a supporting component for semantic search, this chunk’s primary function is query expansion rather than memory management, so the description does not accurately represent the actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

声明强调的是“记忆管理”能力,即保存、维护、检索和回忆用户或项目相关信息;而代码仅对传入的 memory 列表进行分析,总结经验教训。代码中的核心行为包括:格式化 memories、构造提示词、调用 LLM 提取 experience/context/lessons/bestPractices/relatedTopics,以及在 LLM 不可用时基于关键词做简单经验抽取。这属于对记忆内容的后处理分析,而不是记忆系统本身的持久化、检索或时间推理功能。虽然它处理‘memories’,与记忆域相关,但其主要目的与声明存在实质偏差。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明描述聚焦于“记忆管理”能力:保存用户信息、项目状态、经验复用,以及语义搜索和时间推理。但给出的代码块实际是 SubagentClient,其核心职责是包装 LLM 调用:异步加载 ~/.openclaw/openclaw.json 等配置,解析 provider/model,调用外部聊天补全 API,执行超时控制、重试、fallback、JSON 提取和测试模拟。虽然测试模拟中出现了时间解析、总结、经验提取等与记忆系统可能相关的辅助任务,但这只是通用 LLM 响应模拟,不等于实现了记忆持久化或搜索。最关键的是,代码还具备声明中未体现的资源访问与能力:读取本地文件、使用 API 凭证、对外发起网络请求。因此该代码块与声明用途存在明显且实质性的不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个完整的长期记忆管理能力,包括保存信息、后续检索、语义搜索和时间推理。但这段代码只负责对“记忆内容”做摘要加工:构造提示词调用 LLM,返回 summary、keyPoints、importance、suggestedTags、category,并提供简单的本地回退摘要逻辑。它没有任何数据库/文件写入、索引、搜索、时间解析、上下文召回或状态管理代码。因此代码的主要用途与声明的核心能力存在实质不一致,应判定为 mismatch。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

声明的核心用途是长期记忆管理与检索,而代码唯一展示的功能是解析时间表达式(如“明天”“下周三”“3天后”)为 YYYY-MM-DD 日期。虽然声明中提到“支持时间推理”,这与代码部分相关,但这只是整体记忆系统中的一个子能力,无法覆盖其声明的主要职责。代码没有任何持久化、数据库/文件访问、记忆索引、检索、用户资料管理或对话历史管理逻辑,因此描述与实际行为存在显著不一致。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
88% confidence
Finding

该代码与“持久化记忆管理”的核心描述部分一致,因为它确实负责将记忆持久化到文件中并管理其生命周期。但声明里的“支持语义搜索和时间推理”在提供的代码块中并未体现:没有检索、嵌入、排序、查询接口,也没有任何基于时间的推断算法,只有 createdAt/updatedAt/deletedAt/decay 等元数据的读写。代码的实际主功能更窄,属于底层文件存储与回收站管理,而不是完整的记忆管理能力。因此描述对该代码块存在夸大,属于部分不匹配。

Content

No source excerpt is available for this finding.

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · scripts/check-vector-deps.cjs (reported line 161)May include surrounding context.

js
, 'green');
    log('');
    log('   The BGE-small-zh-v1.5 model (~100MB) will be downloaded', 'blue');
    log('   automatically when you first start the vector service.', 'blue');
    log('');
    log('   To start the vector service:', 'green');
    log('   $ python scripts/vector-service.py &', 'green');
    log('');
    log('   Or set environment variable permanently:', 'blue');
    log('   $ echo "export HF_ENDPOINT=https://hf-mirror.com" >> ~/.bashrc', 'blue');

  } catch (error) {
    log('');
    log('❌ Installation failed.', 'red');
    log('   Please try manual installation:', 'yellow');
    log('   $ pip install sentence-transformers numpy', 'blue');
  }
}

main().catch(console.error);

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Memory IDs flow directly into path.join() to form filenames without validation or canonicalization. An attacker who can influence an ID can use path traversal sequences or absolute paths to read, overwrite, move, or delete files outside the intended storage directory, which is especially risky in a persistence component handling long-lived user data.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
79% confidence
Finding

The skill declares installation and likely runtime capabilities that can access the environment and network, but it does not declare any explicit tool scope such as permissions or allowed-tools. That weakens containment and makes it harder for a host agent to enforce least privilege, especially because the skill also references external package installation and optional model downloads.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The invocation description and the rest of the skill documentation are written to operate in Chinese, including Chinese trigger phrases such as "记住" and "别忘了", without indicating that other languages are supported or that the restriction is intentional. This can violate language/locale policy when a skill implicitly mandates one language without user opt-in.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill's core description frames long-term retention of user personal information, preferences, and prior conversation context as normal operation, without stating consent requirements or boundaries. That is dangerous because it normalizes durable storage of potentially sensitive user data by default, increasing privacy, compliance, and misuse risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill encourages persistent storage of personal information and preferences but provides no warning about retention, sensitivity, access controls, or deletion expectations. In a memory skill, that omission materially increases privacy risk because users may disclose sensitive facts without understanding they are being stored long-term.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The usage guidance promotes storing personal details and conversation history in persistent memory without describing privacy boundaries, data minimization, or retention limits. In context, a memory-management skill makes this especially risky because users are likely to provide intimate or identifying information for future reuse.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/check-vector-deps.cjs:30

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/background/vector-search.ts:132

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/background/vector-search.ts:127

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/llm/subagent-client.ts:92

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/llm/subagent-client.ts:388