Back to skill

Security audit

Voice.ai Voices

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real Voice.ai text-to-speech skill, but it needs review because it can overwrite local files and includes account-level voice update/delete capabilities beyond basic speech generation.

Install only if you are comfortable sending TTS text and your Voice.ai API key to Voice.ai, and avoid using it on sensitive text unless you have reviewed Voice.ai's handling policies. Treat the bundled SDK/spec as broader than a simple TTS command: do not expose update/delete voice actions to an agent without explicit user approval, and use safe, dedicated output paths because existing writable files can be overwritten.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
voice-ai-tts-sdk.js:173
Finding

Voice.ai API Credential Can Be Redirected to an Arbitrary HTTPS Host

Content
View full analysis
{ if (value !== undefined && value !== null) { url.searchParams.append(key, value); } }); } const requestOptions = { method, hostname: url.hostname, path: url.pathname + url.search, port: url.port || 443, headers: { 'Authorization': `Bearer ${this.apiKey}`, 'User-Agent': 'VoiceAI-SDK/1.1.5', ...options.headers }, timeout: this.timeout }; ``` The streaming transport follows the same pattern: ```javascript _streamRequest(method, endpoint, options = {}) { const url = new URL(`/api/${API_VERSION}${endpoint}`, this.baseUrl); if (url.protocol !== 'https:') { throw new ValidationError('Only https baseUrl is supported'); } const requestOptions = { method, hostname: url.hostname, path: url.pathname, port: url.port || 443, headers: { 'Authorization': `Bearer ${this.apiKey}`, ...[truncated 2639 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/tts.js:28
Finding

Unrestricted Output Path Allows Overwriting Arbitrary Writable Files

Content
View full analysis
{ const writeStream = fs.createWriteStream(outputPath); stream.pipe(writeStream); writeStream.on('finish', () => resolve(outputPath)); writeStream.on('error', reject); stream.on('error', reject); }); } ``` ### Technical Analysis The `--output` argument is used directly as a filesystem destination without ...[truncated 2037 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documentation presents the skill as a bounded TTS tool, but it also advertises broader voice-management functions including listing, retrieving, and deleting voices, plus writing files locally. This mismatch is dangerous because users or orchestrators may grant trust and invoke the skill under assumptions that are narrower than its actual capabilities, enabling destructive remote actions or unexpected local writes.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
60% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SECURITY.md (reported line 26)May include surrounding context.

md
- Does not execute shell commands (`child_process` is not used).
- Does not download or install other software.
- Does not modify system configuration files.
- Does not run persistently in the background.

## Reporting

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
60% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 54)May include surrounding context.

md
- Does not execute shell commands (`child_process` is not used).
- Does not download or install other software.
- Does not modify system configuration files.
- Does not run persistently in the background.

## Reporting

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill declares required environment variables in metadata but does not declare an explicit tool scope such as permissions or allowed-tools. This creates a transparency and policy gap: hosts and reviewers may not fully understand that the skill requires secret material and can access it at runtime, increasing the chance of over-broad execution in environments with sensitive credentials.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The changelog states that risky voice-sample upload capabilities were removed, but nearby documentation continues to market non-TTS voice-management features. This inconsistency undermines trust in the stated security boundary and may cause reviewers to miss that the package still exposes account-affecting operations beyond simple speech generation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The documentation claims privacy-risky voice-sample upload features were removed, yet it still exposes broader management operations such as deleting voices via the SDK reference. Conflicting security documentation can mislead operators into believing the bundle is narrower and safer than it is, which weakens review and can lead to accidental destructive use of remote account resources.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The SDK sends requests to the Voice.ai API using an Authorization header containing the API key, and methods such as generateSpeech pass user-provided text to the remote service. Although network transmission is core to an SDK like this, the code itself provides no confirmation prompt, user-facing log, or warning that input text and credentials are sent to an external service.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a text-to-speech skill focused on voice synthesis, personas, languages, and streaming. However, the code also lists, retrieves, updates, and deletes voices, which extends beyond straightforward TTS generation into voice resource management, including destructive account-level operations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

generateSpeechToFile writes returned audio directly to the caller-supplied outputPath using fs.writeFileSync, which can overwrite local files. The method lacks any confirmation prompt, visible warning, or user-facing disclosure about local file modification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

streamSpeechToFile creates a write stream to the provided outputPath and pipes remote audio data into it, modifying the local filesystem. There is no confirmation prompt, visible warning, or disclosure that this operation writes a file and may replace existing content.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill explicitly sends user-supplied text and a bearer token to a third-party service, but it does not clearly warn users that their prompts/content leave the local environment and are processed by Voice.ai. In an agent ecosystem, missing disclosure can lead to unintentional transmission of sensitive data, credentials-adjacent metadata, or regulated content to an external provider.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 641)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 662)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 641)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 662)May include surrounding context.

yaml
API_KEY = "your_api_key_here"
      
      response = requests.post(
          "https://dev.voice.ai/api/v1/tts/speech/stream",
          headers={
              "Authorization": f"Bearer {API_KEY}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 685)May include surrounding context.

yaml
-H "Authorization: Bearer YOUR_API_KEY"

    generate_speech: |
      curl -X POST "https://dev.voice.ai/api/v1/tts/speech" \
        -H "Authorization: Bearer YOUR_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{"text": "Hello world!", "audio_format": "mp3"}' \

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · voice-ai-tts.yaml (reported line 693)May include surrounding context.

yaml
typescript:
    generate_speech: |
      const response = await fetch("https://dev.voice.ai/api/v1/tts/speech", {
        method: "POST",
        headers: {
          "Authorization": `Bearer ${API_KEY}`,

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

Both generateSpeech and streamSpeech default the language parameter to 'en', which imposes a specific locale choice when the caller does not specify one. Under the policy, forcing a language without user opt-in can be a natural-language policy concern unless clearly justified or user-selectable.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The manifest frames the skill as providing synthesis via the Voice.ai API with personas, languages, and streaming. The code additionally writes generated or streamed audio directly to arbitrary local file paths, which is extra behavior not disclosed in the description and not necessary to describe API synthesis itself.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The bundled examples write audio to local files such as output.mp3 and stream_output.mp3 without warning about filesystem side effects or possible overwrite behavior. While this is not remote code execution, it can still cause unintended local data loss or create artifacts in sensitive environments if users copy-paste examples blindly.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This manifest sets a fixed default voice while separately listing supported languages, but provides no natural-language indication that users can choose their preferred language or locale. Because the configuration appears to privilege English-oriented defaults without explicit opt-in, it may conflict with language/locale choice expectations.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.