Back to skill

Security audit

OpenAI TTS

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward OpenAI text-to-speech skill that sends user-selected text to OpenAI and optionally saves the returned audio locally.

Install this only if you are comfortable sending the text you provide to OpenAI for speech generation. Avoid using it for secrets or confidential content unless that is allowed in your environment, and note that using --out writes the generated audio to the path you choose.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (5)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill sends user-provided text to OpenAI's /v1/audio/speech API, but the description does not clearly warn that prompt content leaves the local environment and is transmitted to a third-party service. This can cause inadvertent disclosure of sensitive or regulated data if users assume the operation is purely local text-to-speech.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script transmits the caller-provided text to OpenAI's remote TTS endpoint, but the script itself provides no runtime disclosure or confirmation that the input will leave the local environment. In a skill context, users may pass sensitive text assuming a local transformation, so this creates a real privacy and data-handling risk even though the network use is functionally intended.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The hardcoded OpenAI API endpoint confirms that the script sends data to a remote third-party domain. In this context that behavior is expected, but it still matters from a security/privacy perspective because the destination is off-host and receives the full text payload.

Content

Scanner excerpt · scripts/speak.sh (reported line 78)May include surrounding context.

sh
if [[ "$out" == "" ]]; then
  # Output to stdout
  curl -sS https://api.openai.com/v1/audio/speech \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$json"

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The hardcoded OpenAI API endpoint confirms that the script sends data to a remote third-party domain. In this context that behavior is expected, but it still matters from a security/privacy perspective because the destination is off-host and receives the full text payload.

Content

Scanner excerpt · scripts/speak.sh (reported line 78)May include surrounding context.

sh
if [[ "$out" == "" ]]; then
  # Output to stdout
  curl -sS https://api.openai.com/v1/audio/speech \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$json"

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This second curl path writes the response to a file, but it performs the same remote transmission of user-provided text to OpenAI. The output mode does not reduce the privacy risk associated with sending potentially sensitive input to an external API.

Content

Scanner excerpt · scripts/speak.sh (reported line 85)May include surrounding context.

sh
else
  # Output to file
  mkdir -p "$(dirname "$out")"
  curl -sS https://api.openai.com/v1/audio/speech \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$json" \

Static analysis

No suspicious patterns detected.