Back to skill

Security audit

Agent Memory Reflector

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it claims, but it persistently stores full agent prompts and responses in plaintext with weak scoping and no retention or privacy controls.

Review before installing or using this in shared, backed-up, or source-controlled workspaces. Do not log secrets, credentials, personal data, proprietary code, or sensitive retrieved context unless you are comfortable with it being stored locally in plaintext. Use separate working directories per agent and periodically delete .agent_memory if you proceed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
tool.py:22
Finding

Plaintext Conversation Storage and Cross-Agent Data Mixing

Content
View full analysis

Vulnerability Details

File Location: tool.py:22-52 and tool.py:58-92
Vulnerability Type: Plaintext storage of potentially sensitive data with insufficient agent-level isolation
Risk Level: Medium

Vulnerable Code

python
class MemoryReflector:
    def __init__(self, memory_dir: str = ".agent_memory", window_size: int = 50):
        self.memory_dir = memory_dir
        self.window_size = window_size
        self.memory_log = os.path.join(memory_dir, "memory.jsonl")
        self.reflection_log = os.path.join(memory_dir, "reflections.jsonl")
        
        os.makedirs(memory_dir, exist_ok=True)

    def log_interaction(self, agent_id: str, prompt: str, response: str, metadata: Dict = None):
        """Log an agent's input/output for future reflection."""
        entry = {
            "timestamp": datetime.utcnow().isoformat(),
            "agent_id": agent_id,
            "prompt": prompt,
            "response": response,
            "metadata": metadata or {},
            "entry_hash": self._hash_entry(prompt, response)
        }
        with open(self.memory_log, "a") as f:
            f.write(json.dumps(entry) + "\n")

    def _load_recent_memories(self) -> List[Dict]:
        """Load the most recent interactions from memory."""
        if not os.path.exists(self.memory_log):
            return []

        with open(self.memory_log, "r") as f:
            lines = f.readlines()

        entries = [json.loads(line) for line in lines]
        return sorted(entries, key=lambda x: x["timestamp"], reverse=True)[:self.window_size]

Reflection reports can propagate prompt content into a second plaintext file:

python
if h in seen_hashes:
    patterns["repeated_queries"].append(mem["prompt"][:120])
seen_hashes.add(h)

if any(phrase in mem["response"].lower() for phrase in ["i don't know", "unsure", "might be", "could be"]):
    patterns["high_uncertainty_
...[truncated 3215 chars]
Remediation
View remediation

Remediation Suggestions

  1. Create the memory directory with an explicitly restrictive mode such as 0700, and verify or correct the mode when the directory already exists.
  2. Create memory and reflection files with mode 0600 by using os.open() with explicit flags and permissions rather than relying solely on the process umask.
  3. Separate storage by agent identifier. Validate and normalize the identifier before using it in a path to prevent path traversal.
  4. Pass the requested agent identifier into _load_recent_memories() and filter every loaded record so reflection only processes records belonging to that agent.
  5. Avoid storing complete prompts and responses by default. Support configurable field allowlists, secret redaction, truncation, or opt-in content retention.
  6. Do not copy complete prompts into reflection reports. Store aggregate counts or irreversible references unless raw content is explicitly required.
  7. Define configurable retention limits and provide secure deletion or purge functionality.
  8. Where sensitive records must be retained, encrypt them at rest with keys managed separately from the data files.
  9. Document that the memory directory contains potentially sensitive conversation data and must not be placed in shared or source-controlled locations.
  10. Add tests confirming restrictive permissions, per-agent isolation, redaction behavior, and the absence of one agent's records from another agent's reflection report.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (2)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The tool writes complete prompts and responses to a persistent JSONL log on disk, which can capture secrets, credentials, personal data, or proprietary reasoning without any minimization, encryption, or retention control. In an agent setting, prompts and outputs commonly contain sensitive context, so local persistence creates a realistic confidentiality and compliance risk if the host is shared, backed up, or later accessed by other processes or users.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This component intentionally retains natural-language interactions in .agent_memory/memory.jsonl, creating a durable exposure surface for sensitive operational data. Because the skill's purpose is memory and reflection, it systematically accumulates past prompts and responses, which increases blast radius over time and makes accidental disclosure, forensic recovery, or secondary misuse more likely.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.