Back to skill

Security audit

Voice Ai Tts

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent Voice.ai text-to-speech integration, but it also exposes broader account voice management and a configurable SDK endpoint that can put API keys, submitted text, or voice assets at risk if used incautiously.

Review this skill before installing if your Voice.ai account contains private or valuable custom voices. Use it primarily through the documented CLI, keep VOICE_AI_API_KEY scoped appropriately, avoid passing untrusted baseUrl values to the SDK, and be careful with output paths because generated audio files may overwrite existing files.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
voice-ai-tts-sdk.js:178
Finding

Caller-Controlled HTTPS Base URL Can Expose the API Key and Submitted Text

Content
View full analysis
{ if (value !== undefined && value !== null) { url.searchParams.append(key, value); } }); } const requestOptions = { method, hostname: url.hostname, path: url.pathname + url.search, port: url.port || 443, headers: { 'Authorization': `Bearer ${this.apiKey}`, 'User-Agent': 'VoiceAI-SDK/1.1.4', ...options.headers }, timeout: this.timeout }; ``` ```javascript _streamRequest(method, endpoint, options = {}) { const url = new URL(`/api/${API_VERSION}${endpoint}`, this.baseUrl); if (url.protocol !== 'https:') { throw new ValidationError('Only https baseUrl is supported'); } const requestOptions = { method, hostname: url.hostname, path: url.pathname, port: url.port || 443, headers: { 'Authorization': `Bearer ${this.apiKey}`, 'Content-Type': 'application/json', 'User-Agent': 'VoiceAI-SDK/1.1.4', ...options.headers } }; ``` ### Technical Analysis The SDK accepts an arbitrary `options.baseUrl` and validates only that the resulting URL uses HTTPS. It does not verify that the destination hostname or origin is the intended Voice.ai service, `https://dev.voice.ai`. HTTPS protects a connection against in ...[truncated 2736 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The declared purpose emphasizes text-to-speech, but the documented SDK also supports broader operations such as listing/getting/deleting voices and writing arbitrary output files. This mismatch can cause over-trust by users or orchestrators, leading them to enable a skill with more capabilities than expected, especially destructive remote API actions like voice deletion.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
60% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SECURITY.md (reported line 26)May include surrounding context.

md
- Does not execute shell commands (`child_process` is not used).
- Does not download or install other software.
- Does not modify system configuration files.
- Does not run persistently in the background.

## Reporting

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
60% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 61)May include surrounding context.

md
- Does not execute shell commands (`child_process` is not used).
- Does not download or install other software.
- Does not modify system configuration files.
- Does not run persistently in the background.

## Reporting

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The SDK quick reference exposes voice-management functionality, including deleteVoice('voice-id'), even though the skill is presented as a TTS utility. Destructive account-level API operations hidden behind a synthesis-focused label increase the chance of accidental or unauthorized deletion of remote assets if the skill is invoked by an agent or user expecting read/generate-only behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a text-to-speech skill focused on voice synthesis, personas, languages, and streaming. This SDK also supports listing, fetching, updating, and deleting voices, which are separate voice-management capabilities not disclosed in the manifest description.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest markets the skill as text-to-speech, but the toolset also exposes state-changing account-management operations such as update_voice and delete_voice. This scope mismatch can mislead users or higher-level agents into granting or invoking destructive capabilities they did not expect, increasing the risk of accidental data loss or unauthorized modification of private voice assets.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 641)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 662)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 641)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 662)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech/stream",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 685)May include surrounding context.

yaml
-H "Authorization: Bearer YOUR_API_KEY"

    generate_speech: |
      curl -X POST "https://dev.voice.ai/api/v1/tts/speech" \
        -H "Authorization: Bearer YOUR_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{"text": "Hello world!", "audio_format": "mp3"}' \

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 693)May include surrounding context.

yaml
typescript:
    generate_speech: |
      const response = await fetch("https://dev.voice.ai/api/v1/tts/speech", {
        method: "POST",
        headers: {
          "Authorization": `Bearer ${API_KEY}`,

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
72% confidence
Finding

The changelog states the bundle was narrowed to reduce privacy risk by removing upload-related features. Yet the same document still presents SDK methods for non-TTS voice-resource management such as listing, fetching, and deleting voices, which undermines the claimed reduced-scope intent for a simple TTS skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The manifest describes high-quality voice synthesis and streaming via the Voice.ai API, but does not mention local filesystem output. The code adds file-writing behaviors through generateSpeechToFile and streamSpeechToFile, extending behavior beyond the described API-based synthesis/streaming scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The method writes generated audio directly to the filesystem with fs.writeFileSync(outputPath, audio) but provides no confirmation prompt, logging, or inline warning about overwriting or creating files. For code files, file-write operations should have some visible disclosure unless the action is clearly warned elsewhere, and this method's docstring only describes functionality, not the safety implication.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This method creates a write stream to outputPath and pipes remote audio content into it, but there is no prompt, logging, or warning comment/docstring indicating that a filesystem write will occur. The code performs a safety-relevant file write, and the current comments only state the behavior descriptively without disclosing potential impact such as overwriting files.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This manifest includes fixed natural-language gender assignments such as "female" and "male" for voices, and the pattern continues across the file. While not a security flaw, it is a natural-language policy concern because the file imposes a specific gender categorization scheme without any visible opt-in, alternative labels, or documented need for that constraint.

Content

No source excerpt is available for this finding.