Back to skill

Security audit

elite-human-memory

Security checks for vulnerabilities and agentic risk

Overview

This memory skill is not malicious, but it should be reviewed because it enables broad automatic long-term storage and reuse of user context without enough guardrails.

Install only if you are comfortable giving the host agent a long-term memory store. Configure narrow storage paths and scopes, require explicit user consent for writes beyond direct 'remember this' requests, avoid external vector stores for sensitive data unless governed, and provide review/delete controls before enabling scheduled maintenance or automatic promotion.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
INTEGRATION.md:9
Finding

Persistent Prompt Injection Through Unsanitized Agent Memory

Content
View full analysis

Vulnerability Details

File Location: INTEGRATION.md, lines 9–18; related persistence policy in SKILL.md, lines 123–127
Vulnerability Type: Persistent agent memory poisoning
Risk Level: Medium

Vulnerable Code Snippets

INTEGRATION.md, lines 9–18:

python
# On agent startup or context load
if needs_memory_context(query):
    memories = load_semantic_memories(query)  # or keyword fallback
    memories += apply_metadata_filters(memories, scope="project")
    inject_into_prompt(memories)

# When user says "remember this" or strong signals appear
if should_record_memory(user_input):
    entry = create_memory_entry(user_input, metadata={...})
    write_to_daily_file(entry)
    evaluate_for_promotion(entry)

SKILL.md, lines 123–127:

markdown
**Auto-write memory when:**
- User gives explicit “remember this” instructions
- Clear decisions or repeated preferences appear
- New long-running context is established

Technical Analysis

The documented integration stores content derived directly from user_input and subsequently injects retrieved memories into the agent prompt. It does not require sanitization, conversion of raw text into constrained factual fields, instruction detection, provenance-based trust enforcement, or explicit separation between untrusted memory data and executable agent instructions.

Consequently, attacker-controlled text can cross the trust boundary from conversation data into persistent agent state. If the stored text contains directives, semantic or keyword retrieval can place those directives into a later prompt. The model may then interpret the recalled content as instructions instead of inert historical data.

Metadata filtering by project scope does not neutralize embedded instructions. Promotion into semantic memory can increase the persistence and retrieval frequency of a poisoned entry.

Attack Path

  1. An attacker submits instruction-like content that satisfies a recording trigger, such a ...[truncated 1305 chars]
Remediation
View remediation

Remediation Suggestions

  1. Treat all retrieved memories as untrusted data and establish a higher-priority instruction that the agent must never follow commands contained in memory records.
  2. Do not persist raw user input by default. Extract facts into a typed schema with constrained fields, preserving raw text only when necessary and clearly labeling it as an untrusted quotation.
  3. Detect and reject or quarantine instruction-like content, including requests to ignore policies, reveal hidden context, invoke tools, modify security settings, or override future instructions.
  4. Require explicit user confirmation before storing instruction-like text or promoting any entry into long-term semantic memory.
  5. Preserve provenance, author, scope, confidence, creation time, and trust level, and enforce these attributes during retrieval.
  6. Insert recalled material into a clearly delimited data section rather than concatenating it with trusted instructions.
  7. Restrict memory retrieval by user, tenant, project, and session boundaries to prevent poisoned content from affecting unrelated contexts.
  8. Give users mechanisms to inspect, edit, quarantine, and delete stored entries and their associated vector embeddings.
  9. Add tests proving that directives embedded in stored memories cannot override system or developer instructions, trigger tools, or expose prompt-visible data.
  10. Revalidate promoted memories rather than assuming that repetition or semantic relevance establishes trustworthiness.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README promotes persistent filesystem storage, configurable paths, vector storage, and scheduled maintenance, but does not warn that the skill may retain user/agent data or modify local storage over time. For a cross-platform memory system, missing retention and modification warnings increase the risk of sensitive data being stored indefinitely, reviewed later, or written into unexpected locations by adopters.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README instructs integrators to hook unspecified 'behavioral triggers' and 'Auto-read/Auto-write' conditions into an agent loop without defining safety boundaries, approval requirements, or scope limits. In a memory skill, this can lead to over-collection, unintended persistence of sensitive context, or autonomous reads/writes at unsafe times because implementers may guess the intended behavior differently.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill describes persistent episodic and semantic storage but does not present a clear upfront warning that conversational data may be retained on disk or in an external vector store. Users and integrators may therefore deploy it without adequate notice, consent, or data-handling expectations, especially given the cross-platform and optional external storage design.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The auto-read trigger includes the subjective condition that the current context 'feels incomplete or contradictory,' which is too ambiguous for a memory system that may access persisted user data. This can cause the agent to retrieve stored personal context without a clear user request, increasing privacy risk and the chance of inappropriate disclosure or overcollection.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The auto-write criteria ('clear decisions,' 'repeated preferences,' 'new long-running context') are broad and underspecified, so the agent may persist conversational data that the user did not meaningfully consent to store. In a portable multi-agent memory skill, this creates a real risk of retaining sensitive or unnecessary personal/project information across sessions and environments.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.