Back to skill

Security audit

it will help you to send voice messages to your AI Assistant and also can make it talk

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward ElevenLabs text-to-speech and speech-to-text helper, with expected third-party API use and no hidden persistence or deceptive behavior found.

Install only if you are comfortable sending the text or audio you choose to process to ElevenLabs. Use a dedicated ElevenLabs API key, avoid placing unrelated secrets in .env files used with this skill, and do not submit highly sensitive recordings or confidential text unless that external processing is intended.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (14)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description claims both Text-to-Speech and Speech-to-Text functionality, but the supplied code chunk only performs audio transcription. It reads an audio file, sends it to ElevenLabs' speech-to-text endpoint, and returns transcription-related metadata. There is no code for synthesizing speech from text, selecting voices for generation, or any TTS endpoint usage. The multilingual/transcription aspect is broadly consistent, but the overall description overstates the implemented capabilities, so this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code's primary behavior is limited to text-to-speech generation and listing available voices via the ElevenLabs API. There is no code to upload audio, transcribe speech, process voice messages, or call any speech-to-text endpoint. The declared description therefore materially overstates the skill's capabilities by advertising STT/transcription features that are absent. The environment variable loading and file output are supporting details and not the basis for the mismatch.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/elevenlabs_speech.py (reported line 7)May include surrounding context.

python
from pathlib import Path
from dotenv import load_dotenv

# Load environment variables from workspace .env
load_dotenv(dotenv_path=os.path.join(os.path.dirname(__file__), '..', '..', '..', '.env'))

class ElevenLabsClient:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/elevenlabs_speech.py (reported line 8)May include surrounding context.

python
from dotenv import load_dotenv

# Load environment variables from workspace .env
load_dotenv(dotenv_path=os.path.join(os.path.dirname(__file__), '..', '..', '..', '.env'))

class ElevenLabsClient:
    """Client for ElevenLabs Text-to-Speech API"""

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documents use of environment variables and external ElevenLabs APIs, but it does not declare any explicit tool scope such as permissions or allowed-tools. That omission weakens policy transparency and can allow broader-than-expected access to network and secret-bearing environment data when the skill is invoked.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill encourages users to send text and audio to ElevenLabs but does not warn that this content is transmitted to a third-party service for processing. In a voice workflow, that can expose sensitive speech, personal data, or confidential text to an external provider without adequate user awareness or consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The transcription examples process user voice messages through ElevenLabs Scribe without any privacy warning or consent step. Because voice messages may contain biometric, personal, or confidential content, normalizing this workflow without disclosure increases the risk of improper third-party data sharing.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/elevenlabs_scribe.py (reported line 12)May include surrounding context.

python
def __init__(self, api_key=None):
        self.api_key = api_key or os.getenv('ELEVENLABS_API_KEY')
        self.base_url = "https://api.elevenlabs.io/v1"
    
    def transcribe(self, audio_file_path, language_code=None, tag_audio_events=True, 
                   num_speakers=None, timestamps_granularity="word"):

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/elevenlabs_speech.py (reported line 15)May include surrounding context.

python
def __init__(self, api_key=None):
        self.api_key = api_key or os.getenv('ELEVENLABS_API_KEY')
        self.base_url = "https://api.elevenlabs.io/v1"
    
    def transcribe(self, audio_file_path, language_code=None, tag_audio_events=True, 
                   num_speakers=None, timestamps_granularity="word"):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code uploads the full contents of a local audio file to ElevenLabs' external speech-to-text API, but there is no built-in user notice, consent check, or data-sensitivity gating before transmission. In a voice-processing skill, this behavior is expected functionally, but it still creates a real privacy and data-handling risk because audio may contain sensitive personal, financial, or confidential information.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest description says the skill can convert text to speech, transcribe voice messages, and work with speech in multiple languages. In this file, the implemented capabilities are limited to text-to-speech and fetching available voices; there is no speech-to-text or transcription logic anywhere in the code.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This request transmits user-supplied content to an external service, which is expected for TTS but still constitutes a privacy and data-handling risk. The danger is increased because the code path has no safeguards around sensitive content, no minimization, and no transparency to the end user.

Content

Scanner excerpt · scripts/elevenlabs_speech.py (reported line 43)May include surrounding context.

python
}
        
        try:
            response = requests.post(url, headers=headers, json=payload, timeout=60)
            response.raise_for_status()
            
            # Save audio file

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The function sends arbitrary user-provided text to a third-party API without any disclosure, consent check, or sensitivity filtering. In a speech skill context, users may provide private or regulated content, so silent external transmission creates a real data exposure risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The skill's stated purpose is speech synthesis/transcription via ElevenLabs, but the code explicitly reads environment variables from a .env file located several directories above the script. Accessing a workspace-wide secrets file is broader than the user-facing speech functionality and is not mentioned in the manifest description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.