T02 · Agent Memory Poisoning
- Location
references/learning-loop.md:58- Finding
Untrusted Inputs Can Poison Persistent Agent Memory and Operational Rules
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a business automation guide, but it asks agents to store broad business and customer history and perform public or financial actions with weak controls, so it should be reviewed before installation.
Install only after adding explicit guardrails: make agents draft-only by default, require human approval for public posts, outreach batches, deployments, invoices, and payments, restrict connected accounts by department, redact secrets and personal data before storage, partition HR/finance/customer memory, define retention and deletion rules, and require review before captured content changes long-term memory, SOPs, templates, or strategy.
references/learning-loop.md:58Untrusted Inputs Can Poison Persistent Agent Memory and Operational Rules
assets/department-prompts.md:359Department Agents Are Granted Excessive Authority for Consequential External Actions
SKILL.md:79Broad Auto-Capture Creates an Uncontrolled Repository of Sensitive Business and Personal Data
The document explicitly promotes comprehensive auto-capture of decisions, actions, outcomes, and context without any data minimization, consent, or privacy safeguards. In a learning/memory skill, this creates a strong risk of collecting sensitive personal, customer, operational, or confidential data and reusing it in later contexts.
The instruction to capture context, timing, and participants encourages collection of surrounding conversational and identity metadata that may reveal private relationships, internal operations, or sensitive situational details. In a memory-driven agent system, this metadata can later leak through prompts, summaries, or strategic updates even if the original data was not meant for reuse.
┌─────────────────────────────────────────────────────────┐
│ 1. AUTO-CAPTURE │
│ • Log all decisions, actions, outcomes │
│ • Capture context, timing, participants │
│ • Tag by department, project, category │
└────────────────────┬────────────────────────────────────┘
↓
The auto-capture section explicitly includes department agent output and customer interactions, both of which can contain confidential business data, user inputs, and embedded secrets. Because the captured data is intended for later processing and feedback to other components, the skill increases the chance of sensitive information being surfaced, propagated, or leaked across the system.
Automatically capturing customer interactions is particularly risky because those interactions commonly contain personal data, support case details, billing information, or other sensitive content. The skill provides no warning, filtering, or handling rules, so it normalizes indiscriminate ingestion of customer data into memory systems.
The skill recommends placing a strategic business document into permanent agent instructions, which can cause long-term storage and repeated exposure of sensitive business context such as goals, bottlenecks, positioning, and internal priorities. That persistent context may be surfaced unintentionally in later outputs, inherited by downstream workflows, or retained beyond intended business need.
Using permanent instructions for strategic context is a form of persistent context injection: information from one stage is embedded into future model behavior across sessions. While intended to improve consistency, it can unintentionally bias later decisions, expose sensitive context in unrelated conversations, and make it harder to revoke or compartmentalize confidential business information.
### Step 2: Create Agent with Strategy
Feed strategic document into AI agent's permanent instructions. This ensures:
- Every decision is filtered through the strategy
- Agent can push back on misaligned requests
- Context is maintained across sessions
The skill explicitly instructs operators to log all decisions, actions, and outcomes into a knowledge system without any data-minimization, consent, retention, or access-control guidance. In a business OS context, this can capture customer data, employee data, credentials, commercial strategy, or regulated information and make it persistently retrievable by other agents or future tasks.
The instruction to log all decisions, actions, and outcomes promotes indiscriminate retention of natural-language task history. This increases the chance that sensitive user inputs, confidential documents, personal data, or operational secrets are stored verbatim and later leaked through prompts, retrieval, or agent-to-agent sharing.
The best-practice guidance says 'Capture everything' and treats Auto-Capture as non-negotiable, which encourages blanket collection of inputs and outputs regardless of sensitivity. In this skill's self-improving multi-agent architecture, overcollection is more dangerous because captured data is likely to propagate across components and remain available over time.
The Finance Department Agent template explicitly grants autonomy to "send invoices and process payments" without any built-in warning, dual approval, or safety constraint around actions that can move funds or expose sensitive financial data. In a reusable prompt template, this omission can cause downstream agents to be deployed with authority over billing and payment operations without human verification, increasing the risk of fraudulent, erroneous, or privacy-impacting financial actions.
The file advocates retaining broad interaction and outcome data as part of a self-improvement loop, which creates over-collection and downstream leakage risk. Even if initially benign, persistent storage of comprehensive operational and customer-related history expands the attack surface and increases misuse potential.
Appending all captured entries to daily and long-term memory files creates durable retention of potentially sensitive information, including content that may no longer be necessary. Long-term persistence magnifies harm from accidental disclosure, unauthorized access, model reuse, or future context leakage.
The implementation checklist reinforces broad automated collection as a deployment goal, making over-collection a designed behavior rather than an incidental one. In practice, this can cause sensitive content to be captured at scale before privacy or security controls are added.
The checklist operationalizes auto-capture for all agent outputs, which may include prompts, private context, internal reasoning artifacts, secrets, or user-provided sensitive data. Turning this into automation increases the likelihood of broad, silent collection and persistence of information that should not be retained.
No suspicious patterns detected.