Back to skill

Security audit

Bud Semantic Memory

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent memory-search purpose, but it asks users to store an unused Gemini API key and persists sensitive memory data with weak local permission controls.

Review before installing. This skill is not clearly malicious, but it handles long-term memory content and creates persistent copies. Do not create the documented Gemini credential file for this version, restrict permissions on ~/.openclaw/semantic-memory/ and ~/.openclaw/workspace/memory/, and consider pinning chromadb in a controlled environment before use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:74
Finding

Unnecessary plaintext Gemini API credential setup

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 74-79
Vulnerability Type: Unnecessary plaintext secret storage
Risk Level: Medium

Vulnerable Code

markdown
## Requirements

- **ChromaDB** — Local vector database (`pip install chromadb`)
- **Gemini API key** — For generating embeddings (optional, falls back to text search)
  - Get key at: https://makersuite.google.com/app/apikey
  - Save to: `~/.openclaw/credentials/gemini.json` as `{"api_key": "YOUR_KEY"}`

Technical Analysis

The setup instructions direct users to place a Gemini API key in a plaintext JSON file without requiring owner-only file permissions or recommending a secret-management facility.

This credential is not necessary for the reviewed implementation. semantic_memory.py uses ChromaDB's DefaultEmbeddingFunction and does not read ~/.openclaw/credentials/gemini.json. The instruction therefore requests sensitive credential storage beyond the minimum access necessary for the Skill's declared implementation.

The documentation also conflicts with the code: it states that Gemini provides embeddings, while the implementation uses ChromaDB's built-in embedding function.

Attack Path

  1. A user follows the documented setup procedure.
  2. The user creates ~/.openclaw/credentials/gemini.json containing a valid API key.
  3. The file receives permissions determined by the user's shell, editor, and umask because no restrictive mode is required.
  4. If those permissions allow access, another local account, unrelated process, backup service, or diagnostic tool reads the key.
  5. The exposed key is used against the associated Gemini account, subject to its configured permissions and quotas.

Impact Assessment

Successful exploitation can disclose the user's Gemini API credential. An attacker could consume the associated API quota, incur charges where billing is enabled, or access API capabilities granted to that key. T ...[truncated 162 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the Gemini credential instructions because the current implementation does not use that credential.
  • Update the documentation to state that embeddings are generated by ChromaDB's built-in embedding function.
  • If Gemini support is implemented later, obtain the credential from an environment variable, operating-system credential store, or dedicated secret manager.
  • If file-based storage is unavoidable, create the credential directory with mode 0700 and the credential file with mode 0600.
  • Never log the API key or include it in command-line arguments.
  • Document credential rotation and revocation procedures.

T09 · Insecure Skill Coding Practices

Warning
Location
semantic_memory.py:28
Finding

Insufficient filesystem permission hardening for sensitive memory data

Content
View full analysis

Vulnerability Details

File Location: semantic_memory.py, lines 28-32, 87-93, and 157-164
Vulnerability Type: Insecure local storage permissions
Risk Level: Medium

Vulnerable Code

python
DATA_DIR = TOOL_DIR / "data"
LOG_FILE = TOOL_DIR / "memory.log"
COLLECTION_NAME = "memories"

def ensure_dirs():
    os.makedirs(DATA_DIR, exist_ok=True)
    os.makedirs(MEMORY_DIR, exist_ok=True)
    os.chmod(DATA_DIR, 0o755)

Full memory documents are then copied into the persistent database:

python
collection.add(
    ids=[doc_id],
    embeddings=[embedding],
    documents=[content],
    metadatas=[{"file": mem_file.name, "path": str(mem_file)}]
)

New memory files are also opened without explicitly enforcing an owner-only mode:

python
def add_memory(text, source="manual"):
    """Add a new memory"""
    today = datetime.now().strftime("%Y-%m-%d")
    mem_file = MEMORY_DIR / f"{today}.md"

    with open(mem_file, 'a') as f:
        f.write(f"\n## {datetime.now().strftime('%H:%M')} [{source}]\n")
        f.write(f"{text}\n")

Technical Analysis

The Skill processes potentially sensitive long-term Agent memory. It stores complete memory-file contents in ChromaDB and appends user-supplied memory text to Markdown files.

Despite the sensitivity of this information, ensure_dirs() explicitly assigns mode 0755 to the persistent ChromaDB data directory. This permits directory listing and traversal by other local users. Files created beneath that directory, the log file, and Markdown memory files otherwise receive permissions based on the process umask; the Skill does not verify or enforce owner-only access.

A 0755 directory does not independently make every database file readable. Disclosure occurs when files or nested directories also receive group- or world-readable permissions. Because the code leaves those modes to library behavior and the ambient umask, ...[truncated 1195 chars]

Remediation
View remediation

Remediation Suggestions

  • Create TOOL_DIR, DATA_DIR, and MEMORY_DIR with mode 0700.
  • Replace os.chmod(DATA_DIR, 0o755) with an owner-only mode such as 0700.
  • Create memory and log files with mode 0600, using os.open with explicit modes where necessary.
  • Apply a restrictive process umask, such as 0o077, during sensitive file creation.
  • Audit and correct permissions on existing database, log, and memory files during initialization.
  • Ensure ChromaDB-created nested directories and files are also restricted.
  • Avoid storing full documents in the vector database if embeddings and minimal metadata are sufficient.
  • Clearly document that indexed memory is duplicated into persistent local storage.

T08 · Insecure Dependencies

Warning
Location
SKILL.md:8
Finding

Unpinned ChromaDB dependency permits uncontrolled supply-chain updates

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 8-10
Vulnerability Type: Unpinned third-party dependency
Risk Level: Medium

Vulnerable Code

yaml
"openclaw": {
  "requires": { "bins": ["python3"] },
  "install": ["chromadb"]
}

The documentation similarly recommends an unpinned installation:

markdown
- **ChromaDB** — Local vector database (`pip install chromadb`)

Technical Analysis

The Skill installs chromadb without a version constraint, lock file, or integrity hash. Consequently, installation resolves whichever release the configured package index serves at that time.

Third-party Python packages can execute code during installation and whenever imported. The script imports ChromaDB and its embedding implementation at startup. A compromised package release, compromised package index, or maliciously substituted package from an unsafe index configuration could therefore execute code with the privileges of the user running the Skill.

No evidence was found that the named chromadb package is itself malicious. The vulnerability is the absence of dependency pinning and integrity controls.

Attack Path

  1. A user or automated Skill installer executes the documented dependency installation.
  2. Package resolution requests chromadb without a fixed reviewed version or expected hash.
  3. A compromised future release, compromised index, or package substituted through an unsafe package-index configuration is selected.
  4. Attacker-controlled package code executes during installation or when semantic_memory.py imports ChromaDB.
  5. The code runs with the same filesystem, network, and process privileges as the user running the Skill.

Impact Assessment

A compromised dependency could read or modify the user's files, access Agent memories and locally available credentials, make network requests, or execute arbitrary commands with the invoking user's privileges. The ...[truncated 169 chars]

Remediation
View remediation

Remediation Suggestions

  • Pin ChromaDB to a specifically reviewed version rather than installing the latest available release.
  • Maintain a lock file containing exact transitive dependency versions.
  • Require package hashes through a hash-locked requirements file.
  • Install only from a trusted, explicitly configured package index.
  • Review dependency updates before changing pinned versions.
  • Run the Skill in an isolated virtual environment under a non-privileged account.
  • Avoid installing Python dependencies with administrator or root privileges.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (10)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill is presented primarily as a semantic search tool, but the documentation shows it can also add new memories and maintain on-disk logs. That mismatch can mislead users and agents into granting or invoking the skill with broader data modification expectations than the description suggests, increasing the risk of unintended persistence, tampering, or privacy-impacting behavior.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation introduces Gemini API-based embedding generation, which means memory contents may be sent to an external service despite the skill being described as local semantic search using ChromaDB. This is dangerous because users may reasonably assume their memories remain entirely local, leading to unconsented disclosure of potentially sensitive memory data to a third party.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill does not clearly warn that indexing with Gemini may transmit memory contents to the Gemini API for embedding generation. Because memory files can contain secrets, personal data, or operational context, silent or under-disclosed exfiltration to a third-party API is a significant privacy and confidentiality risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
76% confidence
Finding

Memory content is processed by the embedding function for indexing without clear disclosure to the user that their text will be transformed and stored in a vector database. Even though the processing appears local rather than remote, embeddings can still encode sensitive semantic information and create additional persistent copies beyond the original markdown files.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a vector-based semantic search tool that indexes existing memory files and enables meaning-based retrieval. This code goes beyond indexing/search by appending new memory entries to markdown files, which is a content-creation and file-modification capability not reflected in the stated description.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The add_memory path persists arbitrary text to local disk under the user’s home directory without any explicit warning, consent prompt, retention notice, or file permission hardening. In a memory tool, local persistence is expected, but the lack of disclosure can still lead users to store sensitive data unintentionally, increasing confidentiality and privacy risk on multi-user or backed-up systems.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Claiming local vector storage while omitting that embeddings may be produced by an external API creates a misleading privacy model. Even if only embeddings or source text are sent during indexing, the discrepancy can cause users to expose sensitive memory data under false assumptions about locality and isolation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The 'Why This Matters' section uses Chinese text ('找') in examples even though the rest of the document is in English, with no indication that the skill is intended for a Chinese-speaking or locale-specific audience. This can be interpreted as an unexplained language preference rather than an opt-in or justified locale constraint.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The inline comment states 'Get embedding (or skip if no API key)', which contradicts the documented implementation using ChromaDB's built-in embedding function that requires no API key. This is active intent/documentation divergence rather than mere omission, because it describes a different operational condition than the code is designed for.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.