Back to skill

Security audit

Agent Memory Local

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly does local memory search, but it can send local memory excerpts to an external reranking API by default when an API key is present.

Install only if you are comfortable with the memory index storing plaintext snippets locally and with optional SiliconFlow reranking. Set MEMORY_RERANK=0 for local-only use, avoid setting a generic API_KEY in the runtime environment, and do not index secrets or confidential notes unless you have reviewed the data-sharing behavior.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/retrieve.py:292
Finding

Default-Enabled Transmission of Private Memory Content to an External Reranking Service

Content
View full analysis

Vulnerability Details

File Location: scripts/retrieve.py:292-307, 333-356
Vulnerability Type: Sensitive information exposure through implicit external reranking
Risk Level: High

Vulnerable Code

python
def rerank_enabled() -> bool:
    raw = os.environ.get('MEMORY_RERANK', '').strip().lower()
    if raw in {'0', 'false', 'off', 'no'}:
        return False
    if raw in {'1', 'true', 'on', 'yes'}:
        return True
    return RERANK_ENABLED_DEFAULT


def load_siliconflow_key() -> str | None:
    for env_name in ('SILICONFLOW_API_KEY', 'API_KEY'):
        val = os.environ.get(env_name)
        if val:
            return val
    return None
python
api_key = load_siliconflow_key()
if not api_key:
    return {'enabled': True, 'applied': False, 'reason': 'no_api_key'}

documents = [f"{h.get('title', '')}\n{h.get('text', '')}".strip() for h in subset]
payload = json.dumps({'model': RERANK_MODEL, 'query': query, 'documents': documents}).encode('utf-8')
req = urllib.request.Request(
    RERANK_URL,
    data=payload,
    headers={'Authorization': f'Bearer {api_key}', 'Content-Type': 'application/json'},
    method='POST',
)
try:
    with urllib.request.urlopen(req, timeout=RERANK_TIMEOUT) as resp:
        data = json.loads(resp.read().decode('utf-8', errors='ignore'))

Technical Analysis

The Skill indexes and retrieves content from MEMORY.md, memory/learnings.md, and dated memory files. These files may contain private decisions, preferences, incident records, operational details, credentials, or other sensitive long-term agent data.

Remote reranking is enabled by default through RERANK_ENABLED_DEFAULT = True. If the process environment contains either SILICONFLOW_API_KEY or the generic API_KEY variable, the Skill sends the user's query and up to ten selected memory documents to https://api.siliconflow.cn/v1/rerank.

The transmitted documen ...[truncated 3306 chars]

Remediation
View remediation

Remediation Suggestions

  1. Disable remote reranking by default. Set RERANK_ENABLED_DEFAULT = False and require an explicit setting such as MEMORY_RERANK=1.

  2. Remove the generic credential fallback. Accept only SILICONFLOW_API_KEY, preventing unrelated API_KEY values from activating the integration or being sent to SiliconFlow.

  3. Require informed consent. Before the first external request, clearly disclose that the query and selected memory excerpts will leave the local machine. In interactive contexts, request confirmation; in automated contexts, require explicit configuration.

  4. Provide a strict offline mode. Add an option that categorically prevents network requests, regardless of environment variables. Local-first installations should preferably use this mode by default.

  5. Minimize transmitted content. Send only the smallest necessary excerpt, remove unnecessary titles or metadata, and impose strict document and character limits.

  6. Redact sensitive data. Filter likely API keys, access tokens, passwords, private keys, connection strings, email addresses, and other sensitive patterns before constructing the request. Allow users to configure additional redaction rules.

  7. Prevent repeated disclosure in smart-query mode. Perform candidate selection locally and invoke remote reranking at most once for the final candidate set, rather than once for every rewritten query.

  8. Document data handling accurately. State the provider, endpoint, transmitted fields, activation conditions, retention implications, and how to disable all networking.

  9. Add regression tests. Verify that no network call occurs when reranking is not explicitly enabled, that generic API_KEY does not activate SiliconFlow, and that configured redaction removes secrets from request payloads.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill is presented as 'local-first' and positioned as an alternative to remote memory platforms, yet it includes optional reranking via an external API. That mismatch can cause users to unknowingly send queries and memory-derived excerpts to a third party, creating confidentiality and compliance risks precisely because the context involves long-term memory files that may contain sensitive operational history.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 212)May include surrounding context.

md
- `retrieve.py` — direct retrieval engine

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/agent_memory_local.py (reported line 28)May include surrounding context.

python
def run(script: str, *args: str, workspace: str = '') -> int:
    cmd = [*python_cmd(), str(BASE / script), *args]
    env = os.environ.copy()
    if workspace:
        env['AGENT_MEMORY_WORKSPACE'] = workspace
    return subprocess.call(cmd, env=env)

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/explain.py (reported line 29)May include surrounding context.

python
def run(script: str, *args: str, workspace: str = '') -> int:
    cmd = [*python_cmd(), str(BASE / script), *args]
    env = os.environ.copy()
    if workspace:
        env['AGENT_MEMORY_WORKSPACE'] = workspace
    return subprocess.call(cmd, env=env)

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
90% confidence
Finding

Using os.environ.copy() to launch a child process unnecessarily exposes all environment variables to that subprocess, which is a common secret-spillage pattern. In an agent skill context, local helper scripts may parse untrusted workspace data or produce logs/errors visible to users, increasing the chance that inherited credentials are leaked indirectly.

Content

Scanner excerpt · scripts/memory_query.py (reported line 48)May include surrounding context.

python
def run_query(q: str, k: int) -> dict:
    out = subprocess.check_output([*python_cmd(), str(RETRIEVE), q, str(k)], text=True, env=os.environ.copy())
    return json.loads(out)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is described as local-first memory retrieval, but it can transmit the user's query and retrieved memory snippets to a third-party rerank API when an API key is present, and reranking is enabled by default. This creates a confidentiality and transparency issue because sensitive workspace memory may leave the local environment contrary to user expectations.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises operational capabilities that imply file access, shell execution, environment access, and possible networking, but it does not declare any tool scope or permission boundaries. In an agent environment, this increases the risk of over-privileged execution and makes it harder for users or policy engines to understand and constrain what the skill may do.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation mentions optional external reranking but does not warn users that retrieval inputs may include local memory content sent to a remote service. Because these memory files may contain sensitive notes, incidents, preferences, or internal decisions, the lack of a prominent disclosure can lead to inadvertent data exfiltration.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JSON file contains multiple natural-language query strings and expected outputs exclusively in Chinese, indicating the skill behavior or regression coverage is tied to a specific language. Under the policy rules, forcing a specific language without user opt-in or documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/agent_memory_local.py (reported line 31)May include surrounding context.

python
env = os.environ.copy()
    if workspace:
        env['AGENT_MEMORY_WORKSPACE'] = workspace
    return subprocess.call(cmd, env=env)


def main() -> int:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/benchmark.py (reported line 33)May include surrounding context.

python
query = case['query']
    top_k = int(case.get('top_k', 5))
    cmd = [*python_cmd(), str(CLI), '--workspace', str(WORKSPACE), mode, query, '-k', str(top_k)]
    payload = json.loads(subprocess.check_output(cmd, text=True))
    results = payload.get('results') or []
    expect_file = [x.lower() for x in case.get('expect_any_file_contains', [])]
    expect_text = [x.lower() for x in case.get('expect_any_text_contains', [])]

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The index writer persists each chunk's full plaintext (text) alongside its vector into the local index. Because this skill is specifically designed to ingest long-term memory files that may contain sensitive notes, decisions, dates, or personal/workspace context, the index becomes a secondary cleartext copy of all memory content and increases exposure if the index directory is read, synced, committed, or exfiltrated.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/explain.py (reported line 29)May include surrounding context.

python
def run_json(script: Path, query: str, top_k: int) -> dict:
    out = subprocess.check_output([*python_cmd(), str(script), query, str(top_k)], text=True, env=os.environ.copy())
    return json.loads(out)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/memory_query.py (reported line 48)May include surrounding context.

python
def run_query(q: str, k: int) -> dict:
    out = subprocess.check_output([*python_cmd(), str(RETRIEVE), q, str(k)], text=True, env=os.environ.copy())
    return json.loads(out)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code forwards the entire parent environment to the child process with os.environ.copy(), which may include secrets such as API keys, tokens, proxy credentials, or service configuration not needed for retrieval. If retrieve.py is compromised, logs its environment, crashes verbosely, or is influenced by workspace content, those secrets can be exposed or misused.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring states local-first retrieval while also advertising optional remote reranking enabled by default when a key exists. This mismatch is security-relevant because operators may trust the component with sensitive local memory under the false assumption that processing remains entirely local.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

The external URL itself is not malicious, but in this skill context it represents a real outbound data path from a local-memory retriever to a third-party service. That becomes dangerous because the component handles potentially sensitive memory content and the network use is not aligned with the advertised local-first trust model.

Content

Scanner excerpt · scripts/retrieve.py (reported line 34)May include surrounding context.

python
TOKEN_RE = re.compile(r"[A-Za-z0-9_\-\u4e00-\u9fff]{2,}")
DATE_RE = re.compile(r"(20\d{2})-(\d{2})-(\d{2})")

RERANK_URL = 'https://api.siliconflow.cn/v1/rerank'
RERANK_MODEL = 'BAAI/bge-reranker-v2-m3'
RERANK_ENABLED_DEFAULT = True
RERANK_MIN_RESULTS = 2

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The presence of a hardcoded external rerank endpoint and model introduces network-based processing into a tool whose purpose is local memory retrieval. Even if intended as a quality enhancement, it expands the attack surface and can expose sensitive memory contents to an external service.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/retrieve.py (reported line 259)May include surrounding context.

python
return state

    try:
        proc = subprocess.run([*python_cmd(), str(BUILD_SCRIPT)], cwd=str(WORKSPACE), capture_output=True, text=True, timeout=AUTO_REBUILD_TIMEOUT, check=True)
        rebuilt_meta = load_meta()
        state['status'] = 'rebuilt'
        state['rebuilt'] = True

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code reads third-party API credentials from environment variables and uses them for a remote service in a supposedly local-first retriever. This increases the chance that external transmission occurs silently in environments where such keys are already configured, defeating least surprise and potentially leaking sensitive memory data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The rerank request sends both the user query and document text derived from local memory files to an external API without any explicit warning, consent, or sanitization at the call site. Because memory data often contains internal decisions, dates, root causes, or preferences, this can result in unintended disclosure of sensitive workspace information.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/smoke_test.py (reported line 28)May include surrounding context.

python
def run_json(*args: str) -> dict:
    out = subprocess.check_output([*python_cmd(), str(CLI), *args], text=True)
    return json.loads(out)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The smoke test hard-codes Chinese natural-language queries in multiple command invocations, which implies the skill is validated against a specific language only. There is no nearby comment or user-facing note indicating that the tool is intentionally Chinese-only or that other languages are unsupported, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/smoke_test.py (reported line 55)May include surrounding context.

python
assert_true(len(smart.get('attempted_queries', [])) >= 1, 'smart-query missing attempted queries')
    assert_true(len(smart.get('results', [])) >= 1, 'smart-query returned no results')

    explain_out = subprocess.check_output([
        *python_cmd(), str(CLI), '--workspace', str(WORKSPACE), 'explain', '飞书昨天为什么断联了', '--smart', '-k', '2'
    ], text=True)
    explain_payload = json.loads(explain_out)

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The phrase "Chinese-friendly" expresses a language/locale preference in the skill's architecture description, but the file does not indicate that users can choose their preferred language or that this locale focus is optional. Under the policy criteria, language constraints should be opt-in or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.