Back to skill

Security audit

Agent Memory for Commerce

Security checks for vulnerabilities and agentic risk

Overview

This non-executing commerce-memory guide is not malicious, but its production-style examples handle customer and payment data with under-scoped safeguards that warrant review before use.

Review this before installing or using its examples in a real commerce agent. Treat the code as instructional, not production-ready: pin and validate GreenHelix endpoints before sending API keys, never use session tokens as identity lookup values or Redis key material, minimize audit and handoff fields, require recipient authorization for agent handoffs, and verify any incoming payment-flow state against an authoritative source before persisting it.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:113
Finding

Bearer Credential Can Be Disclosed to an Attacker-Controlled Gateway

Content
View full analysis
dict: """Execute a GreenHelix tool via the GreenHelix REST API.""" response = requests.post( f"{GATEWAY_URL}/v1", headers=headers, json={"tool": tool, "input": params}, timeout=30, ) response.raise_for_status() return response.json() ``` ### Technical Analysis The gateway URL is read directly from the `GREENHELIX_API_URL` environment variable. The code does not validate that the destination: - Uses HTTPS. - Belongs to an approved GreenHelix hostname. - Uses an expected port. - Does not redirect the request to another origin. The same request attaches `GREENHELIX_API_KEY` as a bearer credential. Consequently, control over the environment variable is sufficient to redirect the credential and request payloads to an arbitrary network endpoint. Although the document is an educational guide rather than an executable package, it describes its examples as production-ready. Deploying this implementation without additional validation would create a credential-exfiltration primitive. ### Attack Path 1. An attacker gains the ability to influence deployment configuration, a container environment, a CI/CD variable, or a generated `.env` file. 2. The attacker sets `GREENHELIX_API_URL` to an endpoint under their control. 3. The application initializes `GATEWAY_URL` from the modified value. 4. Any call to `execute()` sends an `Authorization: Bearer ...` header to the attacker-controlled endpoint. 5. The attacker captures the API key and reuses it against the legitimate GreenH ...[truncated 543 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:1182
Finding

Raw Personal Identifiers and Session Tokens Are Transmitted and Embedded in Redis Keys

Content
View full analysis
str | None: """Resolve a set of identifiers to a canonical customer ID. Args: identifiers: Dict of identifier types to values, e.g., {"email": "alice@example.com", "session": "abc123"} Returns: Canonical agent_id or None if no match. """ # Check each identifier against the mapping cache for id_type, id_value in identifiers.items(): cache_key = f"identity_map:{id_type}:{id_value}" canonical_id = self.r.get(cache_key) if canonical_id: return canonical_id # No cache hit -- try GreenHelix identity lookup for id_type, id_value in identifiers.items(): try: result = execute("search_agents", { "query": id_value, "limit": 1, }) agents = result.get("agents", []) if agents: canonical_id = agents[0]["agent_id"] # Cache all provided identifiers for this customer for t, v in identifiers.items(): cache_key = f"identity_map:{t}:{v}" self.r.set(cache_key, canonical_id, ex=86400) return canonical_id except Exception: continue return None ``` The demonstrated call includes a session credential: ```python customer_id = resolver.resolve({ "email": "alice@example.com", "session_token": "sess_abc123", }) ``` ### Technical Analysis The resolver accepts an unrestricted dictionary of identifier types and values. Each raw value is: 1. Sent to the hosted gateway as a `search_agents` query. 2. Embedded verbatim in a Redis key. 3. Retained in Redis for 24 hours when a match is found. There is no allowlist limiting searches ...[truncated 1997 chars]
Remediation
View remediation
str: digest = hmac.new( secret, f"{id_type}:{value}".encode(), hashlib.sha256, ).hexdigest() return f"identity_map:{id_type}:{digest}" ``` 5. Do not transmit identifiers to third-party services unless necessary, authorized, and covered by an appropriate privacy policy. 6. Cache only the identifier that was independently verified; do not bind every supplied value after one successful lookup. 7. Apply short retention periods, encryption, access controls, deletion support, and audit logging. 8. Return a distinct security error for prohibited identifier classes instead of silently trying them. ]]>

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:1609
Finding

Agent Handoffs Disclose Full Customer and Payment Context Without Recipient Authorization

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:1737
Finding

Unauthenticated Handoff Content Can Poison Persistent Payment State

Content
View full analysis
HandoffContext | None: """Check for incoming handoff contexts.""" messages = execute("get_messages", { "agent_id": self.agent_id, "message_type": "handoff_context", "limit": 1, }) if not messages.get("messages"): return None msg = messages["messages"][0] context_data = json.loads(msg["content"]) # Reconstruct HandoffContext context = HandoffContext(**context_data) # Log receipt self.audit.log_event( event_type=AuditEventType.STATE_TRANSITION, payload={ "action": "handoff_received", "source_agent": context.source_agent, "reason": context.handoff_reason, }, customer_id=context.customer_id, flow_id=context.active_flow_id, ) # If there is an active payment flow, re-register it locally if context.flow_state and context.active_flow_id: r = redis.Redis.from_url( os.environ["REDIS_URL"], decode_responses=True ) r.set( f"payment_flow:{self.agent_id}:{context.active_flow_id}", json.dumps(context.flow_state), ex=7200, ) return context ``` ### Technical Analysis The receiving agent trusts fields parsed from the message body, including: - `source_agent` - `customer_id` - `active_flow_id` - `flow_state` - `conversation_summary` The implementation does not compare the authenticated transport sender with the body’s claimed `source_agent`. It also does not verify a digital signature, authorize the sender, detect replay, validate flow ownership, or enforce legal payment-state transitions. Most critically, attacker-controlled `flow_state` is written into Redis as an operational payment flow for two hours. This co ...[truncated 1654 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:1295
Finding

Generic Audit Logger Uploads Arbitrary Payloads Without Redaction or Schema Enforcement

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (11)

Ssd 3

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This section explicitly directs agents to package conversation summaries, transaction state, and customer profiles for transfer to other agents. That is a direct data-sharing pattern involving potentially sensitive financial and personal information, and without strict authorization and minimization it can become lateral data exposure across the agent ecosystem.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The introductory text broadly advocates persistent, cross-session, and cross-agent retention of customer and transaction context as a core design goal. In a commerce setting, this encouragement is dangerous because it can normalize collecting and sharing more conversational and identity data than necessary, increasing privacy, compliance, and insider-exposure risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 38)May include surrounding context.

md
Every 1.3 seconds, an AI agent somewhere drops a shopping cart because it forgot what the customer just said. Every 4.7 seconds, a payment agent creates a duplicate charge because it lost track of a transaction that was already confirmed. Every 18 seconds, a support agent asks a customer to repeat information they provided two turns ago. These are not hypothetical failure modes. They are measured production incidents from the first wave of commerce agents deployed in 2025 and early 2026. Forrester's Q1 2026 AI Commerce Infrastructure Report estimates that stateless agent failures cost enterprises $2.3 billion in the past twelve months -- abandoned transactions, duplicate charges, customer churn from repetitive interactions, and compliance violations from lost audit trails. The $67 billion in AI-agent-mediated transactions projected for 2026 will not survive on stateless architectures. Agents that handle money need memory. Not the vague, "store some embeddings in a vector database" kind of memory. They need structured, tiered, transaction-aware memory that persists across sessions, survives restarts, reconciles against ledger entries, and shares state across agent handoffs without data races or consistency violations.
This guide builds that memory system from scratch. Using the GreenHelix A2A Commerce Gateway's 128 tools accessible at `https://api.greenhelix.net/v1`, you will implement a three-tier memory architecture mapped to commerce data flows: immediate context for the current conversation, operational memory for active transaction state, and institutional knowledge for customer history and compliance records. Every chapter contains production-ready Python code, decision matrices for choosing your memory stack, and patterns extracted from teams that have already shipped memory-backed commerce agents at scale.
1. [Why Stateless Agents Fail at Commerce](#chapter-1-why-stateless-agents-fail-at-commerce)

## What You'll Learn

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 58)May include surrounding context.

md
Every 1.3 seconds, an AI agent somewhere drops a shopping cart because it forgot what the customer just said. Every 4.7 seconds, a payment agent creates a duplicate charge because it lost track of a transaction that was already confirmed. Every 18 seconds, a support agent asks a customer to repeat information they provided two turns ago. These are not hypothetical failure modes. They are measured production incidents from the first wave of commerce agents deployed in 2025 and early 2026. Forrester's Q1 2026 AI Commerce Infrastructure Report estimates that stateless agent failures cost enterprises $2.3 billion in the past twelve months -- abandoned transactions, duplicate charges, customer churn from repetitive interactions, and compliance violations from lost audit trails. The $67 billion in AI-agent-mediated transactions projected for 2026 will not survive on stateless architectures. Agents that handle money need memory. Not the vague, "store some embeddings in a vector database" kind of memory. They need structured, tiered, transaction-aware memory that persists across sessions, survives restarts, reconciles against ledger entries, and shares state across agent handoffs without data races or consistency violations.
This guide builds that memory system from scratch. Using the GreenHelix A2A Commerce Gateway's 128 tools accessible at `https://api.greenhelix.net/v1`, you will implement a three-tier memory architecture mapped to commerce data flows: immediate context for the current conversation, operational memory for active transaction state, and institutional knowledge for customer history and compliance records. Every chapter contains production-ready Python code, decision matrices for choosing your memory stack, and patterns extracted from teams that have already shipped memory-backed commerce agents at scale.
1. [Why Stateless Agents Fail at Commerce](#chapter-1-why-stateless-agents-fail-at-commerce)

## What You'll Learn

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The guide strongly promotes persistent storage of transaction state, customer history, and compliance records, but it does not give an explicit user-facing privacy warning or consent boundary near the initial description of data handling. In a commerce-memory skill, that omission increases the risk that implementers will persist personal and transactional data without transparency, minimization, or lawful-basis controls.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 124)May include surrounding context.

md
def execute(tool: str, params: dict) -> dict:
    """Execute a GreenHelix tool via the GreenHelix REST API."""
    response = requests.post(
        f"{GATEWAY_URL}/v1",
        headers=headers,
        json={"tool": tool, "input": params},

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a commerce memory guide focused on remembering customers, maintaining transaction state, and reconciling billing context. In the identity resolution example, the code performs broad identity lookup via search_agents using arbitrary identifier values and can create new identities with register_agent, which is an account/identity-management capability not clearly justified by a memory-layer skill.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The audit-trail guidance encourages persistent logging of customer interactions, state transitions, agent decisions, and operational context. While auditability is legitimate in payments, logging rich interaction and decision context without careful scoping can create a long-lived repository of sensitive personal and transactional data that is attractive for misuse or breach.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The handoff section instructs implementers to send conversation summaries, flow state, and customer profiles to other agents, but omits an explicit warning that this is a sensitive disclosure and may also be written to logs/events. That omission can cause over-sharing of customer data across internal agent boundaries without consent, need-to-know restrictions, or audit scoping.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The handoff code packages conversation summaries, customer profile data, and active transaction state, then transmits that bundle to another agent and also writes handoff metadata to persistent events. This creates a clear cross-agent data-sharing channel that can expose sensitive customer and billing context to unintended agents if routing, authorization, or minimization controls are weak.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 2142)May include surrounding context.

md
### Anti-Pattern 3: Trusting Memory Without Validation

**The mistake:** The agent reads a balance from memory and makes a purchase decision without checking whether the balance is current. The memory says $500. The actual balance is $12 (another agent spent the difference). The purchase fails at execution time, after the customer has already confirmed.

**The fix:** Always validate financial state against the source of truth before irrevocable actions. The `StaleStateDetector` pattern above handles this. For balance checks specifically, always call `get_balance` directly from GreenHelix before creating a payment intent -- never rely on cached balance for purchase decisions.

Static analysis

No suspicious patterns detected.