Back to skill

Security audit

Whisper Piper Voice

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its stated local voice-server purpose, but it needs Review because it exposes an unauthenticated network service by default and recommends unverified downloads plus persistent autostart.

Install only if you are comfortable hardening a local HTTP service. Pin and verify downloaded executables, bind the server to localhost unless you add authentication and TLS, run it as a dedicated unprivileged user, add request size/rate limits, and avoid enabling the systemd autostart until those controls are in place.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
Findings (5)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:23
Finding

Mutable and Unverified External Executables and Dependencies

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:23-35; duplicated in references/setup-guide.md:13-18, 37-46
Vulnerability Type: Unverified remote executable retrieval and unpinned dependencies
Risk Level: High

Affected code:

bash
python3 -m venv ~/whisper-env && source ~/whisper-env/bin/activate
pip install faster-whisper
apt install ffmpeg  # or brew install ffmpeg on macOS
bash
mkdir -p ~/piper && cd ~/piper
wget https://github.com/rhasspy/piper/releases/latest/download/piper_linux_x86_64.tar.gz
tar xzf piper_linux_x86_64.tar.gz
mkdir voices && cd voices
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/de/de_DE/thorsten_emotional/medium/de_DE-thorsten_emotional-medium.onnx
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/de/de_DE/thorsten_emotional/medium/de_DE-thorsten_emotional-medium.onnx.json

Technical Analysis

The setup retrieves a native Piper executable through a mutable latest release URL, extracts it, and later instructs the user to execute it. No expected checksum, cryptographic signature, fixed release version, or provenance verification is provided. The Python dependency is also installed without a pinned version or lock file.

GitHub and Hugging Face are recognizable upstream hosting services, and the project does not pipe downloaded data directly into a shell. Nevertheless, the effective native payload can change after this Skill has been reviewed. A compromised upstream account, release pipeline, package index, dependency, or mutable release could therefore introduce arbitrary executable code.

Attack Path

  1. An attacker compromises an upstream release, package, account, or distribution channel.
  2. The mutable latest asset or unpinned Python dependency is replaced with a malicious version.
  3. A user follows the documented installation commands without verifying artifact integrity.
  4. The downloa ...[truncated 596 chars]
Remediation
View remediation

Remediation Suggestions

  • Replace latest with an explicitly reviewed release version.
  • Publish and verify an expected SHA-256 or stronger digest before extraction.
  • Verify upstream signatures or attestations where available.
  • Pin Python packages to reviewed versions and use a hash-locked requirements file.
  • Install dependencies only from explicitly configured, trusted indexes.
  • Fail installation when integrity verification does not succeed.
  • Document the expected archive layout and validate extracted paths to prevent archive path traversal.
  • Run the resulting executable as a dedicated, unprivileged account inside a restricted service sandbox.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/voice-server.py:93
Finding

Unauthenticated Voice API Listens on Every Network Interface

Content
View full analysis

Vulnerability Details

File Location: scripts/voice-server.py:93-101
Vulnerability Type: Unrestricted network exposure and missing access control
Risk Level: High

Affected code:

python
handler = create_handler(model, args.piper_bin, args.piper_model, args.piper_speaker, args.speed)
print(f'Voice Server on http://0.0.0.0:{args.port}')
print(f'  POST /transcribe  (audio → text)')
print(f'  POST /speak       (text → audio/ogg)')
HTTPServer(('0.0.0.0', args.port), handler).serve_forever()

The request handler exposes both operations without authentication:

python
if self.path == '/transcribe':
    self._transcribe(data)
elif self.path == '/speak':
    self._speak(data)
else:
    self.send_response(404)
    self.end_headers()

Technical Analysis

Binding to 0.0.0.0 exposes the service through every available network interface. Neither endpoint authenticates callers, checks authorization, limits source addresses, or requires transport encryption.

Although the Skill describes the pipeline as local and offline, this binding allows any host with network reachability to submit transcription or speech-generation jobs. Plain HTTP also allows submitted audio, text, and generated responses to be observed or modified by an attacker with access to the network path.

Attack Path

  1. An attacker identifies a host exposing TCP port 9998, or another configured server port.
  2. The attacker sends arbitrary POST requests to /transcribe or /speak.
  3. The server accepts the requests without credentials or authorization checks.
  4. Whisper, Piper, and ffmpeg process attacker-controlled jobs using the victim's CPU or GPU.
  5. Repeated or expensive requests consume resources and can prevent legitimate use.

Impact Assessment

A remote attacker with network access can use the service without authorization and consume the host's processing, memory, and storage resources. Subm ...[truncated 348 chars]

Remediation
View remediation

Remediation Suggestions

  • Bind to 127.0.0.1 by default.
  • Add an explicit --host option and require conscious opt-in before accepting remote connections.
  • Require authenticated requests, using strong API credentials or mutual TLS.
  • Place remote deployments behind a TLS-enabled reverse proxy with authorization and rate limiting.
  • Restrict access with host firewall rules or a private network.
  • Apply per-client concurrency, request-rate, and resource quotas.
  • Avoid logging credentials or sensitive request contents.
  • Clearly document that 0.0.0.0 is not a local-only configuration.

T09 · Insecure Skill Coding Practices

Error
Location
scripts/voice-server.py:19
Finding

Unbounded HTTP Request Body Can Exhaust Server Memory

Content
View full analysis

Vulnerability Details

File Location: scripts/voice-server.py:19-29
Vulnerability Type: Unbounded attacker-controlled allocation and blocking read
Risk Level: High

Affected code:

python
def do_POST(self):
    length = int(self.headers.get('Content-Length', 0))
    data = self.rfile.read(length)

    if self.path == '/transcribe':
        self._transcribe(data)
    elif self.path == '/speak':
        self._speak(data)
    else:
        self.send_response(404)
        self.end_headers()

Technical Analysis

Content-Length is controlled by the client and is converted to an integer without validation or an upper bound. The implementation then reads the entire body into memory before determining how it will be processed.

The server is based on the single-threaded HTTPServer. An oversized body can consume substantial memory, while a client that advertises a large body and sends it slowly can occupy the only request-processing thread. Invalid Content-Length values can also raise an uncaught conversion exception.

Attack Path

  1. An attacker connects to the exposed HTTP port.
  2. The attacker supplies a very large Content-Length value.
  3. The server attempts to read the complete body into the data object.
  4. The attacker either transmits enough data to exhaust memory or sends it slowly to hold the single server thread.
  5. Legitimate transcription and speech requests become unavailable, and the process may be terminated by the operating system.

Impact Assessment

Exploitation can cause remote denial of service through memory exhaustion, prolonged blocking, or process termination. Because the service also performs computationally expensive inference, even bodies that fit in memory can consume significant CPU or GPU capacity.

Remediation
View remediation

Remediation Suggestions

  • Validate Content-Length and reject missing, negative, malformed, or oversized values.
  • Define endpoint-specific maximum sizes and return HTTP 413 Payload Too Large when exceeded.
  • Stream audio into a securely created temporary file instead of loading it entirely into memory.
  • Configure socket read timeouts to mitigate slow-client attacks.
  • Limit concurrent requests and computational jobs.
  • Apply rate limiting at the application or reverse-proxy layer.
  • Return controlled 400 responses for malformed headers rather than allowing uncaught exceptions.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/voice-server.py:43
Finding

Predictable Temporary Output Files Permit Local Race Attacks

Content
View full analysis

Vulnerability Details

File Location: scripts/voice-server.py:43-68
Vulnerability Type: Insecure temporary-file creation
Risk Level: Medium

Affected code:

python
def _speak(self, data):
    wav_file = ogg_file = None
    try:
        text = json.loads(data).get('text', '')
        speaker = json.loads(data).get('speaker', piper_speaker)
        wav_file = tempfile.mktemp(suffix='.wav')
        ogg_file = tempfile.mktemp(suffix='.ogg')
        subprocess.run(
            [piper_bin, '--model', piper_model, '--speaker', str(speaker),
             '--length_scale', str(speed), '--output_file', wav_file],
            input=text.encode(), capture_output=True, timeout=30
        )
        subprocess.run(
            ['ffmpeg', '-i', wav_file, '-c:a', 'libopus', '-b:a', '32k', ogg_file, '-y'],
            capture_output=True, timeout=15
        )
        with open(ogg_file, 'rb') as f:
            audio = f.read()

Technical Analysis

tempfile.mktemp() only produces a candidate pathname; it does not atomically create and reserve the file. A race exists between pathname selection and the subsequent Piper or ffmpeg operation.

A local attacker who can access the shared temporary directory may attempt to pre-create the generated path as a symbolic link or manipulate it before a subprocess opens it. The -y ffmpeg option permits overwriting an existing destination. Exploitability depends on the attacker's ability to predict or discover the temporary name and win the race, but Python explicitly considers mktemp() unsafe for this reason.

Attack Path

  1. A local attacker monitors or predicts temporary filenames used by the service.
  2. After mktemp() returns but before Piper or ffmpeg creates the file, the attacker places a file or symbolic link at that path.
  3. The subprocess opens or overwrites the attacker-selected target using the service account's privileges.
  4. Th ...[truncated 580 chars]
Remediation
View remediation

Remediation Suggestions

  • Replace tempfile.mktemp() with tempfile.NamedTemporaryFile(), tempfile.mkstemp(), or a private TemporaryDirectory.
  • Atomically create files with restrictive permissions before invoking subprocesses.
  • Store both outputs inside a per-request private directory accessible only to the service account.
  • Verify subprocess return codes with check=True before reading output.
  • Remove partial files through a scoped cleanup mechanism.
  • Run the service as a dedicated unprivileged user with narrowly restricted filesystem access.

T06 · System Persistence

Warning
Location
references/setup-guide.md:84
Finding

Persistent Auto-Start of an Unauthenticated Network Service

Content
View full analysis

Vulnerability Details

File Location: references/setup-guide.md:84-114
Vulnerability Type: System-wide persistent service registration
Risk Level: Medium

Affected code:

ini
# /etc/systemd/system/voice-server.service
[Unit]
Description=Whisper STT + Piper TTS Server
After=network.target

[Service]
Type=simple
User=your-user
WorkingDirectory=/home/your-user
ExecStart=/home/your-user/whisper-env/bin/python3 /path/to/voice-server.py \
  --port 9998 \
  --whisper-model small \
  --piper-bin /home/your-user/piper/piper/piper \
  --piper-model /home/your-user/piper/voices/your-voice.onnx
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
bash
sudo systemctl enable voice-server
sudo systemctl start voice-server

Technical Analysis

Auto-start is a legitimate optional operational feature for a server and is openly documented rather than concealed. It is not required for basic transcription or speech-generation functionality, however, and the guide uses administrator privileges to install and enable a system-wide service.

Restart=always and WantedBy=multi-user.target cause the process to return after failures and system reboots. Because the supplied application binds to every interface and does not authenticate requests, persistence materially extends the duration of that exposure. The unit also omits common systemd sandboxing and resource-control directives.

Attack Path

  1. A user creates the documented unit in /etc/systemd/system using administrator privileges.
  2. The user enables and starts the service.
  3. The unauthenticated server starts on subsequent boots and is restarted after termination.
  4. Any attacker with network reachability can repeatedly access the endpoints.
  5. The absence of service hardening allows compromised dependencies or runtime flaws to operate with the full authority of the configured user.

Impact Asse

...[truncated 350 chars]

Remediation
View remediation

Remediation Suggestions

  • Present auto-start as an explicit, optional deployment step rather than part of the minimum setup.
  • Prefer a user-level systemd service when system-wide registration is unnecessary.
  • Change the application to bind to localhost by default before recommending persistence.
  • Use a dedicated account with no interactive login and minimal filesystem permissions.
  • Add systemd hardening directives such as NoNewPrivileges=true, PrivateTmp=true, ProtectSystem=strict, ProtectHome=true, and narrowly scoped ReadWritePaths.
  • Add MemoryMax, CPUQuota, task limits, and restart-rate controls.
  • Apply firewall rules, authentication, and TLS before exposing the persistent service remotely.
  • Document disablement and removal procedures for the service.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (12)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill clearly instructs shell-based installation and execution steps, but it does not declare any tool scope or allowed-tools boundaries. In an agent environment, that mismatch can let the skill be invoked for broad requests and then perform local command execution without explicit upfront restriction, increasing the risk of unintended package installation, downloads, or system modification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description contains broad trigger phrases like setting up voice capabilities, transcribing audio, generating speech, configuring STT/TTS, and building a voice assistant pipeline. Overbroad matching can cause the skill to activate in many common contexts, which is dangerous here because the skill includes shell commands, dependency installation, external downloads, and server exposure steps.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

Transcribe (audio → text):

bash
curl -X POST -F "file=@message.ogg" http://HOST:9998/transcribe
# {"text": "Hallo Welt", "language": "de"}

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/setup-guide.md (reported line 112)May include surrounding context.

text

```bash
sudo systemctl enable voice-server
sudo systemctl start voice-server

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/setup-guide.md (reported line 113)May include surrounding context.

text

```bash
sudo systemctl enable voice-server
sudo systemctl start voice-server

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · references/setup-guide.md (reported line 112)May include surrounding context.

text

```bash
sudo systemctl enable voice-server
sudo systemctl start voice-server

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/setup-guide.md (reported line 120)May include surrounding context.

Transcribe Audio → Text

bash
curl -X POST -F "file=@audio.ogg" http://localhost:9998/transcribe
# → {"text": "Hello world", "language": "en"}

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The transcribe endpoint writes user-supplied audio to a named temporary file on disk before processing. Even though the file is deleted in a finally block, audio may persist briefly on disk and could remain after crashes, exposing sensitive speech content on multi-user systems or systems with disk forensics, backups, or monitoring.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/voice-server.py (reported line 50)May include surrounding context.

python
speaker = json.loads(data).get('speaker', piper_speaker)
                wav_file = tempfile.mktemp(suffix='.wav')
                ogg_file = tempfile.mktemp(suffix='.ogg')
                subprocess.run(
                    [piper_bin, '--model', piper_model, '--speaker', str(speaker),
                     '--length_scale', str(speed), '--output_file', wav_file],
                    input=text.encode(), capture_output=True, timeout=30

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The /speak handler invokes subprocesses for both Piper and ffmpeg, which is safety-relevant shell execution in a network-facing service. Although TTS generation is the feature's purpose, the code provides no explicit disclosure in comments, logs, or endpoint description that user input is passed to external executables for processing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/voice-server.py (reported line 55)May include surrounding context.

python
'--length_scale', str(speed), '--output_file', wav_file],
                    input=text.encode(), capture_output=True, timeout=30
                )
                subprocess.run(
                    ['ffmpeg', '-i', wav_file, '-c:a', 'libopus', '-b:a', '32k', ogg_file, '-y'],
                    capture_output=True, timeout=15
                )

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The setup and API examples consistently use a German Piper voice and German sample output ('Hallo Welt', language 'de') without noting that this is only an example or giving the user a language/locale choice. This can be interpreted as a locale preference embedded in the skill's natural-language guidance.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.