Back to skill

Security audit

Clonev

Security checks for vulnerabilities and agentic risk

Overview

This voice-cloning skill is functional and not malicious, but it needs Review because it handles sensitive voice data with weak consent guidance, persistent sample storage, and a mutable Docker runtime.

Install only if you are comfortable handling voice recordings as sensitive biometric data. Use it only with your own voice or explicit permission, avoid celebrity or third-party impersonation, delete retained files under the configured voice-samples directory, and consider pinning and hardening the Docker image before use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T08 · Insecure Dependencies

Error
Location
scripts/clonev.sh:34
Finding

Mutable Remote Container Image Is Executed with Access to Host Data

Content
View full analysis
&2 ``` ### Technical Analysis The script executes `ghcr.io/coqui-ai/tts:latest`, which is identified only by a mutable tag. A mutable tag can resolve to different container contents in future executions without any change to this reviewed project. There is no immutable digest, signature verification, provenance check, or version lock. The container receives writable host mounts for the model cache, all retained voice samples, and generated output. Consequently, code in the image can read every voice recording in the shared sample directory and modify files in all three mounted host directories. The container is also not configured with network isolation, a read-only root filesystem, dropped capabilities, or an explicit non-root user. This is a supply-chain exposure. The reviewed repository does not itself prove that the current upstream image is malicious, but its effective executable payload can change after review. ### Attack Path 1. An attacker compromises the upstream image publisher, registry account, build pipeline, or mutable `latest` tag. 2. The `latest` tag is changed to reference an altered image. 3. A user invokes `scripts/clonev.sh`. 4. Docker retrieves or runs the altered image. 5. The altered image reads voice recordings from `/samples` and can modify the writable model and output mounts. 6. If outbound networking ...[truncated 677 chars]
Remediation
View remediation
" ``` 2. Verify image provenance and signatures before deployment, and update the digest only through a controlled review process. 3. Mount the model directory read-only when runtime modification is unnecessary: ```bash -v "${COQUI_DIR}/models-xtts:/root/.local/share/tts:ro" ``` 4. Expose only the selected voice sample through a private per-run directory rather than mounting the shared sample collection. 5. Use a separate, permission-restricted output directory for each invocation. 6. Disable networking when it is not required: ```bash --network none ``` 7. Harden the container with an explicit non-root user, dropped Linux capabilities, `no-new-privileges`, resource limits, and a read-only root filesystem where compatible. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/clonev.sh:15
Finding

Shared Sample and Output Filenames Allow Stale-Voice Substitution and Cross-Request Races

Content
View full analysis
&2 # Convert to OGG for Telegram ffmpeg -y -i "$OUTPUT_WAV" -c:a libopus -b:a 24k -vn "$OUTPUT_OGG" 2>/dev/null # Cleanup WAV rm -f "$OUTPUT_WAV" # Return OGG path echo "$OUTPUT_OGG" ``` ### Technical Analysis The sample cache is indexed only by the source file's basename. If a file with that basename already exists, the script does not compare its contents or replace it. A caller supplying a different recording with the same basename therefore causes the preexisting recording to be used. All executions also share the fixed paths `clonev_output.wav` and `clonev_output.ogg`. There is no per-run directory, unique identifier, file lock, or atomic output publication. Concurrent processes can overwrite, convert, delete, or return another process's output. Shell quoting prevents direct command injection through the supplied arguments, but it does not address the shared-state flaw. ### Attack Path #### Stale-sample substitution 1. An earlier user or attacker causes a voice recording named, for example, `sample.wav` to be stored in the shared `voice-samples` directory. 2. A victim later invokes the script with a different source file that is also named `sample.wav`. 3. The exis ...[truncated 1137 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/clonev.sh:20
Finding

Sensitive Voice Samples Are Persistently Retained Without Cleanup

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The markdown repeatedly promotes cloning any voice, including 'someone else's voice,' 'a celebrity,' or 'a character,' yet provides no warning, restriction, or verification around consent, privacy, or impersonation abuse. In this context, the lack of safeguards materially increases the chance of non-consensual voice cloning for scams, social engineering, reputational harm, or bypassing voice-based trust mechanisms.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill description is intentionally broad and explicitly supports cloning 'your voice or someone else's voice' and generating speech that sounds like a specific person. In an agent setting, this can cause the skill to be invoked for impersonation, fraud, harassment, or non-consensual synthetic media generation without any gating, consent checks, or safety screening.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This skill is explicitly designed to clone voices and generate impersonation-capable audio, creating elevated risks of non-consensual cloning, fraud, social engineering, and privacy violations. Although the guide includes a brief ethics section later, it does not present a prominent, operational warning near the start or embed consent, provenance, retention, and misuse safeguards throughout the workflow.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/complete-guide.md (reported line 157)May include surrounding context.

md
- ✅ **Do**: Clone your own voice
- ✅ **Do**: Clone with explicit permission
- ✅ **Do**: Use for personal productivity
- ❌ **Don't**: Clone without consent
- ❌ **Don't**: Use for deception/fraud
- ❌ **Don't**: Impersonate others maliciously

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script copies the provided voice sample into a persistent voice-samples directory and does not remove it afterward or warn the user that retention occurs. Because voice samples are sensitive biometric data, retaining them beyond the immediate task increases privacy risk, unauthorized reuse risk, and exposure if the host or container environment is later accessed by another user or process.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
91% confidence
Finding

The container image is pulled as ghcr.io/coqui-ai/tts:latest, which is mutable and can change over time without review. In a skill that processes arbitrary user text and local voice samples, a compromised or unexpectedly updated image could execute untrusted code with access to mounted host directories containing sensitive biometric data and outputs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.