Back to skill

Security audit

Elevenlabs Twilio Memory Bridge

Security checks for vulnerabilities and agentic risk

Overview

The skill matches its stated memory-bridge purpose, but it needs Review because it persists sensitive caller context and injects stored text into agent system prompts, with encryption optional.

Install only after reviewing the data handling model. Use DATA_ENCRYPTION_KEY in production, keep ADMIN_API_KEY, WEBHOOK_SECRET, and PHONE_HASH_SALT strong and private, restrict who can add memories or global notes, review stored entries for prompt-injection text, define deletion and retention procedures, and make sure callers have appropriate notice or consent.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T02 · Agent Memory Poisoning

Warning
Location
app.py:204
Finding

Persistent System-Prompt Injection Through Stored Memories and Notes

Content
View full analysis

Vulnerability Details

File Location: app.py:204-257, app.py:496-516, memory.py:220-268
Vulnerability Type: Persistent prompt injection through unsanitized stored context
Risk Level: Medium

Vulnerable Code

python
def build_memory_context(phone_hash: str) -> str:
    """Assemble a memory context block for injection into the system prompt.

    Combines long-term facts and relevant daily notes for the caller.
    """
    sections: list[str] = []

    facts = get_memories(phone_hash)
    if facts:
        bullet_list = "\n".join(f"- {f}" for f in facts)
        sections.append(f"## Known Facts About This Caller\n{bullet_list}")

    notes = get_notes(phone_hash)
    if notes:
        note_lines = "\n".join(
            f"- [{datetime.fromtimestamp(n['timestamp'], tz=timezone.utc).strftime('%Y-%m-%d')}] {n['note']}"
            for n in notes[-10:]
        )
        sections.append(f"## Context Notes\n{note_lines}")

    return "\n\n".join(sections)


def build_system_prompt(session: Session, phone_hash: str) -> str:
    """Build the full system prompt from soul template + memory context."""
    parts: list[str] = [_soul_template]

    memory_ctx = build_memory_context(phone_hash)

    caller_name = session.get("caller_name", "")
    call_count = session.get("call_count", 1)

    caller_info_lines: list[str] = []
    if caller_name:
        caller_info_lines.append(f"- Caller name: {caller_name}")
    caller_info_lines.append(f"- This is call #{call_count} from this caller")
    if call_count > 1:
        caller_info_lines.append("- This is a returning caller — reference prior context naturally")

    caller_section = "## Caller Info\n" + "\n".join(caller_info_lines)

    context_block = "\n\n".join(filter(None, [caller_section, memory_ctx]))
    parts.append(f"---\n# Caller Context\n{context_block}")

    return "\n\n".join(parts)

Administrative endpoin ...[truncated 3583 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat all facts and notes as untrusted data, even when submitted through an authenticated administrative endpoint.
  2. Represent context using a structured data block rather than unrestricted prose whenever the downstream API permits it.
  3. Add an explicit higher-priority instruction stating that memory and note content is reference data, cannot modify policies, and must never be interpreted as instructions.
  4. Place stored content inside clearly delimited or serialized fields and escape delimiter characters.
  5. Apply strict maximum lengths and collection limits to facts, notes, and the final generated prompt.
  6. Validate content and reject known instruction-injection patterns where appropriate, while recognizing that filtering alone is not a complete defense.
  7. Restrict global-note creation to a separate, more privileged operation because global entries affect all callers.
  8. Use separate credentials and narrowly scoped authorization for caller-specific and global writes.
  9. Maintain audit logs for memory creation, modification, deletion, submitting identity, and affected caller scope.
  10. Add deletion and review workflows so poisoned entries can be identified and removed promptly.
  11. Ensure downstream tools independently enforce authorization and never rely on prompt instructions as a security boundary.

T09 · Insecure Skill Coding Practices

Warning
Location
memory.py:138
Finding

Sensitive Caller Memories and Notes Are Stored in Plaintext by Default

Content
View full analysis

Vulnerability Details

File Location: memory.py:25-28, memory.py:138-155, memory.py:220-268
Vulnerability Type: Optional encryption permits plaintext storage of sensitive caller data
Risk Level: Medium

Vulnerable Code

Encryption is disabled when DATA_ENCRYPTION_KEY is not configured:

python
DATA_DIR: Path = Path(os.getenv("DATA_DIR", "./data"))
_DATA_ENCRYPTION_KEY: str | None = os.getenv("DATA_ENCRYPTION_KEY")
_fernet: Fernet | None = Fernet(_DATA_ENCRYPTION_KEY.encode()) if _DATA_ENCRYPTION_KEY else None

The encryption helper returns the original plaintext in that configuration:

python
def _encrypt(data: str) -> str:
    """Encrypt *data* with Fernet if a key is configured."""
    if not _fernet:
        return data
    return _fernet.encrypt(data.encode("utf-8")).decode("utf-8")


def _decrypt(data: str) -> str:
    """Decrypt *data* with Fernet if a key is configured."""
    if not _fernet:
        return data
    try:
        return _fernet.decrypt(data.encode("utf-8")).decode("utf-8")
    except Exception:
        logger.error("Failed to decrypt data — key mismatch or corruption")
        return "[DECRYPTION_FAILED]"

Facts and notes are consequently persisted in their original form:

python
def add_memory(phone_hash: str, fact: str) -> list[str]:
    """Append *fact* to long-term memory for *phone_hash*.

    Returns the updated list of facts.
    """
    memories: dict[str, MemoryStore] = _read_json(_memories_path())
    if phone_hash not in memories:
        memories[phone_hash] = MemoryStore(phone_hash=phone_hash, facts=[])
    stored_fact = _encrypt(fact) if _fernet else fact
    memories[phone_hash].setdefault("facts", []).append(stored_fact)
    _write_json(_memories_path(), memories)
    logger.info("Stored fact for %s (total: %d)", phone_hash[:8], len(memories[phone_hash]["facts"]))
    return memories[phone_hash]["fact
...[truncated 2631 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require DATA_ENCRYPTION_KEY in production and fail startup when it is absent.
  2. Introduce an explicit development-only setting if plaintext storage is needed for local testing; do not make plaintext the implicit fallback.
  3. Encrypt sensitive session attributes in addition to facts and notes.
  4. Store encryption keys in a managed secret service rather than source files, images, or committed environment files.
  5. Define and test a key-rotation procedure that can decrypt existing records and re-encrypt them with a new key.
  6. Encrypt backups and snapshots independently of application-level encryption.
  7. Apply retention periods and deletion procedures so caller data is not stored indefinitely.
  8. Minimize collected caller information and document which fields may contain regulated or sensitive data.
  9. Preserve restrictive file and directory permissions as defense in depth.
  10. Add startup logging or health-state reporting that clearly indicates whether encryption at rest is active, without exposing the key.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (22)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 69)May include surrounding context.

2. Configure environment

bash
cp .env.example .env
# Edit .env with your actual values

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 70)May include surrounding context.

bash
cp .env.example .env
# Edit .env with your actual values

Required secrets (generate before first run):

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 261)May include surrounding context.

bash
cp .env.example .env
# Edit .env with your actual values

Required secrets (generate before first run):

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about a webhook service that personalizes conversational AI interactions by storing caller memory and injecting dynamic context. The supplied code chunk does not implement or demonstrate that functionality directly; instead, it contains automated tests for admin authentication rules and endpoint access behavior. While some endpoint names suggest relation to memory and webhooks, the primary behavior here is security testing of admin auth, not personalization, persistent caller memory logic, Twilio/ElevenLabs integration, or context injection. Therefore, the code chunk does not accurately represent the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for an application feature set centered on a personalization webhook integrating with ElevenLabs and Twilio, including persistent caller memory and dynamic context injection. The supplied code does none of that. It is a test module that constructs temporary FastAPI apps and validates CORS response header behavior using Hypothesis and TestClient. There is no webhook processing, no file-based persistence, no agent context injection, and no external service integration. This is a clear material mismatch in primary purpose and implemented capabilities, not merely a supporting detail.

Content

No source excerpt is available for this finding.

Possible Typosquatting: 'uvicorn' resembles popular package 'gunicorn'

High
Category
Supply Chain
Confidence
70% confidence
Finding

Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The README explicitly describes a workflow where caller metadata, memory, and personalized context are sent between Twilio, ElevenLabs, this service, and an external LLM backend, but it does not include a clear privacy/compliance warning to operators about handling personal data. In a telephony context, this can lead to inadvertent exposure of sensitive caller information to third-party processors without adequate user notice, consent, retention controls, or jurisdictional review.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README states that memories and notes are encrypted only when an optional key is configured, meaning deployments may store sensitive personal facts in plaintext by default or through operator omission. Because the project encourages storing caller preferences and potentially health-related or personal notes, weak warning language increases the risk of accidental insecure deployment and local data disclosure.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 139)May include surrounding context.

  1. Hang up, call again — the agent remembers your previous interaction
  2. Add a memory via the API:
    bash
    curl -X POST http://localhost:8000/api/memory/PHONE_HASH \
      -H "Authorization: Bearer <your-admin-api-key>" \
      -H "Content-Type: application/json" \
      -d '{"fact": "Prefers to be called Mike"}'
    

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly permits storing caller memories and notes as plain JSON when DATA_ENCRYPTION_KEY is unset, yet it handles persistent telephony-linked caller context that can contain sensitive personal data. In this context, unencrypted at-rest storage increases the risk of privacy breach from host compromise, backup exposure, misconfigured permissions, or operator access, especially because phone-derived identities and behavioral notes are being retained over time.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code exposes admin APIs that add persistent memory facts and notes, which are file-backed state changes affecting caller data. While the functions have developer-facing docstrings, there is no confirmation prompt or user-facing disclosure in this file indicating that caller-associated data will be written or globally applied.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module stores caller sessions, long-term facts, and notes on disk, and encryption is explicitly optional via DATA_ENCRYPTION_KEY. If the key is unset, sensitive conversational memory is written as plaintext JSON, so compromise of the host, backups, logs, or mounted volumes can expose personal data. In a voice-agent personalization bridge, this is more dangerous because the stored data is likely to contain persistent caller-specific information and may accumulate over time.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The docstring at L320 says a whitespace-only ADMIN_API_KEY should behave as unconfigured and return 403. However, the inline comment and assertion expect a 401 for a non-matching token because the code treats a whitespace string as configured. This is an active contradiction between the test's stated intent and the behavior it validates.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · app.py (reported line 74)May include surrounding context.

python
# ── Logging ─────────────────────────────────────────────────────────────────

logging.basicConfig(
    level=getattr(logging, LOG_LEVEL, logging.INFO),
    format="%(asctime)s | %(levelname)-7s | %(name)s | %(message)s",
    datefmt="%Y-%m-%dT%H:%M:%S",
)

Unverifiable Dependency: fastapi has 3 known advisory(ies) (CVE-2021-32677 (Cross-Site Request Forgery (CSRF) in FastAPI); CVE-2021-32677 (FastAPI is a web framework for building APIs with Python 3.6+ based on standard ); CVE-2024-24762 (FastAPI is a web framework for building APIs with Python 3.8+ based on standard )), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: uvicorn has 4 known advisory(ies) (CVE-2020-7694 (Log injection in uvicorn); CVE-2020-7695 (HTTP response splitting in uvicorn); CVE-2020-7694 (This affects all versions of package uvicorn. The request logger provided by the) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: python-dotenv has 2 known advisory(ies) (CVE-2026-28684 (python-dotenv: Symlink following in set_key allows arbitrary file overwrite via ); CVE-2026-28684 (python-dotenv reads key-value pairs from a .env file and can set them as environ)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: pydantic has 4 known advisory(ies) (CVE-2021-29510 (Use of "infinity" as an input to datetime and date fields causes infinite loop i); CVE-2024-3772 (Pydantic regular expression denial of service); CVE-2021-29510 (Pydantic is a data validation and settings management using Python type hinting.) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: cryptography has 16 known advisory(ies) (GHSA-39hc-v87j-747x (Vulnerable OpenSSL included in cryptography wheels); CVE-2023-50782 (Python Cryptography package vulnerable to Bleichenbacher timing oracle attack); GHSA-537c-gmf6-5ccf (Vulnerable OpenSSL included in cryptography wheels) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

This code posts to /api/memory/{phone_hash} and /api/notes, which are data-writing operations, but the helper and surrounding tests only describe authentication behavior. Within this file there is no user-facing warning, confirmation, or comment disclosing that these requests write memory/notes data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
71% confidence
Finding

The file repeatedly configures and sends ADMIN_API_KEY values via Authorization headers, which is handling a sensitive credential. While this is expected in auth tests, this file does not include any warning or note about secret handling, redaction, or safe use.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.