Back to skill

Security audit

Supercall

Security checks for vulnerabilities and agentic risk

Overview

SuperCall is a real phone-calling skill, but it exposes a high-impact calling and audio-processing surface with under-scoped controls around public media streams, recording, callbacks, and transcript storage.

Install only if you are comfortable operating a real outbound-calling service. Use a dedicated Twilio number and OpenAI key with spending limits, avoid public exposure unless necessary, prefer localhost or a tightly controlled tunnel, review consent and recording laws, and assume call audio/transcripts may be stored locally and by Twilio until you verify retention and deletion controls.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
src/webhook.ts:170
Finding

Unauthenticated Media WebSocket Allows Unauthorized OpenAI Realtime Sessions

Content
View full analysis
{ const url = new URL( request.url || "/", `http://${request.headers.host}`, ); if (url.pathname === streamPath) { console.log("[supercall] WebSocket upgrade for media stream"); this.mediaStreamHandler?.handleUpgrade(request, socket, head); } else { socket.destroy(); } }); ``` ```ts // src/media-stream.ts:61-72 handleUpgrade(request: IncomingMessage, socket: Duplex, head: Buffer): void { if (!this.wss) { this.wss = new WebSocketServer({ noServer: true }); this.wss.on("connection", (ws, req) => this.handleConnection(ws, req)); } this.wss.handleUpgrade(request, socket, head, (ws) => { this.wss?.emit("connection", ws, request); }); } ``` ```ts // src/media-stream.ts:149-178 private async handleStart( ws: WebSocket, message: TwilioMediaMessage, ): Promise { const streamSid = message.streamSid || ""; const callSid = message.start?.callSid || ""; // Guard against duplicate Twilio WebSocket connections for the same call. // Twilio sometimes sends two WS upgrades; the second would create a // competing OpenAI session and both end up dying. for (const existing of this.sessions.values()) { if (existing.callId === callSid) { console.log(`[MediaStream] Ignoring duplicate stream ${streamSid} for call ${callSid} (already have ${existing.streamSid})`); ws.close(); return null; } } console.log(`[MediaStream] Stream started: ${streamSid} (call: ${callSid})`); const instructions = this.config.getInstructionsForCall?.(callSid); const initialGreeting = this.config.getInitialGreetingForCall?.(callSid); const conversationSession = ...[truncated 2826 chars]
Remediation
View remediation
` URL or as a Twilio custom stream parameter. 3. Validate the token during or immediately after the WebSocket upgrade and bind it to the expected internal call ID and provider call SID. 4. Reject a `start` event unless: - The call SID maps to an active call. - The account SID matches the configured Twilio account. - The token is valid, unexpired, unused, and associated with that call. - The stream SID and media format are valid. 5. Do not create or connect an OpenAI session until all validation completes. 6. Add per-IP and global connection limits, handshake timeouts, maximum message sizes, and rate limits. 7. Close connections that send media before a valid `start` event or send duplicate/out-of-order events. 8. Consider using Twilio-supported request validation mechanisms where applicable, while retaining a call-bound nonce because ordinary WebSocket upgrade requests do not provide the same signed webhook body. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
src/manager.ts:636
Finding

Sensitive Call Transcripts and Metadata Are Persisted Without Restrictive File Permissions

Content
View full analysis
{ console.error("[supercall] Failed to persist call record:", err); }); } ``` The storage directory is initialized without an explicit restrictive mode: ```ts fs.mkdirSync(this.storePath, { recursive: true }); ``` ### Technical Analysis Call records are serialized in plaintext and appended to `calls.jsonl`. These records can include phone numbers, session keys, persona prompts, goals, opening messages, provider call identifiers, timestamps, call states, and complete user/assistant transcripts. Neither directory creation nor file creation specifies a security mode. Effective access therefore depends on the process umask and permissions inherited from parent directories. On systems with permissive defaults, other local users or services may be able to read the records. The implementation also has no retention period, rotation mechanism, size limit, content redaction, or encryption. Because every state update appends another complete call record, sensitive information is duplicated and the log grows indefinitely. ### Attack Path 1. A call is initiated and conversation content is added to its transcript. 2. The manager repeatedly serializes the complete call record into `calls.jsonl`. 3. A local user, compromised process, backup reader, or service account obtains read access through pe ...[truncated 896 chars]
Remediation
View remediation

other

Warning
Location
src/providers/twilio.ts:317
Finding

All Twilio Calls Are Recorded Without Explicit Opt-In or Adequate Disclosure

Content
View full analysis
= { To: input.to, From: input.from, Url: url.toString(), // TwiML serving endpoint StatusCallback: statusUrl.toString(), // Separate status callback endpoint StatusCallbackEvent: ["initiated", "ringing", "answered", "completed"], Record: "true", RecordingChannels: "dual", Timeout: "30", }; const result = await this.apiRequest( "/Calls.json", params, ); ``` ### Technical Analysis Every Twilio call is created with `Record: "true"` and dual-channel recording. Recording is therefore mandatory and cannot be disabled through the exposed configuration. Provider-side recording is not required for OpenAI Realtime audio streaming, transcript generation, DTMF navigation, webhook callbacks, or call-state management. It creates an additional copy of sensitive conversation data under the configured Twilio account. The Skill documentation describes local transcript persistence but does not adequately disclose that Twilio also records both call channels. It also does not provide recording-consent controls, retention configuration, deletion behavior, or a warning about jurisdiction-specific consent requirements. ### Attack Path 1. An authorized user or agent invokes `persona_call`. 2. The Skill creates the Twilio call with recording enabled automatically. 3. Twilio records both sides of the conversation. 4. The recording remains associated with the configured Twilio account according to provider settings and retention behavior. 5. Anyone who later obtains sufficient Twilio account access can retrieve or manage the recording. 6. The caller and recipient may not have been informed that the call is being recorded. ### Impac ...[truncated 672 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (32)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description describes a telephony/voice automation skill, but the supplied code chunk contains only a local filesystem path helper. This behavior is unrelated to making phone calls, handling audio, navigating IVRs, or interacting with OpenAI/Twilio services. While utility code can be a supporting detail, this specific utility concerns filesystem path expansion and resolution, which does not meaningfully support the stated phone-calling functionality based on the provided chunk alone. Therefore the code does not accurately represent the declared purpose.

Content

No source excerpt is available for this finding.

Context Leakage

High
Category
Data Exfiltration
Confidence
85% confidence
Finding

Code or instructions that leak agent conversation context to external services, potentially exposing sensitive user interactions.

Content

Scanner excerpt · src/providers/openai-realtime-conversation.ts (reported line 194)May include surrounding context.

ts
console.log("[RealtimeConversation] WebSocket connected");
        this.connected = true;

        // Send session configuration after short delay for stability
        setTimeout(() => this.sendSessionUpdate(), 250);
        resolve();
      });

Context Leakage

High
Category
Data Exfiltration
Confidence
85% confidence
Finding

Code or instructions that leak agent conversation context to external services, potentially exposing sensitive user interactions.

Content

Scanner excerpt · src/providers/openai-realtime-conversation.ts (reported line 307)May include surrounding context.

ts
break;

      case "session.updated":
        // Log the session config to verify tools were registered
        const sessionData = event.session as { tools?: unknown[] } | undefined;
        console.log(`[RealtimeConversation] Session updated - tools registered: ${sessionData?.tools?.length ?? 0}`);
        // Trigger initial greeting if configured

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill explicitly documents access to environment secrets, outbound network connectivity, inbound listeners, process spawning, and disk writes, yet it declares no tool scope such as permissions or allowed-tools. That omission weakens least-privilege controls and makes it harder for a host system or reviewer to bound what the skill may do at runtime.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The invocation guidance is broad enough to match many ordinary user requests involving calls, confirmations, deliveries, or phone-tree handling, which can cause over-selection of a high-impact external-action skill. Over-broad routing increases the chance of unintended real-world calls, charges, privacy exposure, or use without sufficient user confirmation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill promotes autonomous calling and IVR/DTMF navigation but does not present a clear upfront warning about consent, call recording/transcript retention, financial charges, or the impact of contacting third parties. In this context, the omission is significant because the skill can autonomously interact with real people and systems, potentially causing unauthorized contact, privacy violations, or costly actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The natural-language instruction "ONLY speak English" imposes a fixed language policy on all calls. This file does not offer user opt-in, language selection, or any documented reason for restricting conversations to English only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The tool can initiate real-world outbound phone calls immediately with arbitrary persona, goal, and opening line, but there is no built-in confirmation, policy gate, or user-visible warning. This increases the risk of unintended calls, social-engineering misuse, spam, or costly external actions being triggered by an agent without sufficient human approval.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill automatically packages full call transcripts and goal context into callback text that is sent back into the agent system. This creates a natural-language data exfiltration and prompt-injection channel from an external phone participant into the agent, and may also expose sensitive personal or business information collected during the call.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

On call completion, the skill posts transcript-derived text into the internal /hooks/wake endpoint, which can immediately trigger a new agent turn. Because the callback text is built from untrusted phone-recipient speech, this creates a cross-channel prompt-injection path where an external callee can influence downstream agent behavior beyond the phone-call scope.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description is broad enough to justify use in many situations without clearly constraining when autonomous calling should occur. For a skill that can place phone calls, adopt personas, and navigate IVR systems, vague activation scope increases the risk of unintended or unauthorized use, including social-engineering or nuisance-calling scenarios.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This manifest collects telephony credentials, can place real voice calls, and exposes a service on 0.0.0.0 with optional public tunneling, but it does not prominently warn users about those capabilities. Missing user-facing warnings can lead to unsafe deployment decisions, accidental internet exposure, unexpected billing, and privacy risks from handling live call audio and API keys.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The docblock at L19-L25 defines inbound-call policies such as allowlist, pairing, and open, while the inline comment at L26 states that inbound calls are not supported in this skill. This is an active contradiction in the file's own documentation about whether inbound call handling exists.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code persists complete call records, including transcripts, phone numbers, session keys, and metadata, to a local JSONL file without any minimization, encryption, retention control, or consent/disclosure mechanism visible here. In a voice-calling skill, this creates a meaningful privacy and data-exposure risk because sensitive spoken content and call metadata may be stored indefinitely and could be accessed by other local users, backups, or compromised processes.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The embedded IVR guidance explicitly instructs the assistant to 'ALWAYS stay in English' and to ignore offers for other languages. This is a natural-language locale policy constraint that removes user choice and is not accompanied by an opt-in mechanism or clear region-specific justification in the file.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code injects initialGreeting into a faux '[SYSTEM: ...]' directive inside a user-role message, which creates semantic privilege confusion. If initialGreeting is attacker-controlled or derived from untrusted input, it can smuggle prompt instructions that override conversation behavior, manipulate tool use, or alter what the model says at call start.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · src/providers/twilio.ts (reported line 56)May include surrounding context.

ts
this.accountSid = config.accountSid;
    this.authToken = config.authToken;
    this.baseUrl = `https://api.twilio.com/2010-04-01/Accounts/${this.accountSid}`;
    this.options = options;

    if (options.publicUrl) {

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code explicitly enables call recording and dual-channel recording for every outbound call, but the skill description does not disclose this behavior. In a voice-calling skill, undisclosed recording creates privacy, consent, and regulatory risk because conversations and potentially sensitive DTMF/speech content may be captured without user awareness.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The outbound call flow turns on recording without any warning, consent gate, or visible disclosure in this file. Given this skill autonomously calls people and navigates phone systems, silent recording increases privacy and compliance exposure and may capture sensitive personal or authentication information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Tailscale tunnel is started with --bg --yes, which automatically enables serving/funneling without any explicit user confirmation at the point of exposure. In this skill, that behavior can unintentionally publish a localhost webhook to a wider network or the public internet, increasing the attack surface for a telephony-facing service that may receive untrusted inbound requests.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The server logs raw user speech and assistant transcripts directly to console, which can capture sensitive personal data, authentication details spoken over IVR, or regulated information. In a voice-calling skill, transcript content is especially likely to include PII, making plaintext logging a meaningful privacy and security risk if logs are retained, aggregated, or viewed by unauthorized operators.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest describes a skill for AI-powered phone calls, webhook handling, realtime audio, and IVR navigation. In addition to that expected functionality, this file spawns the local tailscale binary and configures serve/funnel routes, which is a separate system/network administration capability not justified by the stated calling purpose itself.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code invokes Tailscale 'serve' or 'funnel' commands with '--bg' and '--yes', which can expose the local webhook endpoint over the network and make a persistent routing change. Although there is console logging after success or failure, there is no prior warning in comments or docstrings explaining that the function changes network accessibility.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill reads the host's hooks authentication token from shared configuration and uses it to make an authenticated internal callback. That expands the skill's privilege boundary: a phone-calling plugin gains the ability to invoke a more general agent-wake mechanism, increasing blast radius if the skill is abused or compromised.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · package.json (reported line 36)May include surrounding context.

json
"LICENSE"
  ],
  "dependencies": {
    "@sinclair/typebox": "^0.34.0",
    "ws": "^8.19.0",
    "zod": "^3.24.0"
  },

Static analysis

Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/tunnel.ts:70

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/webhook.ts:335

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/manager.ts:58

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/providers/twilio.ts:55

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/tunnel.ts:291