Back to skill

Security audit

mar-edge-tts

Security checks for vulnerabilities and agentic risk

Overview

This looks like a real text-to-speech skill, but it sends user text to an external API and has enough documentation mismatches that users should review it before installing.

Install only if you are comfortable sending text you ask to read aloud to the SkillBoss/HeyBoss TTS API. Avoid using it for secrets, private messages, regulated data, or confidential work until the publisher clarifies provider/data handling, fixes the subtitle and dependency documentation mismatches, and tightens activation/confirmation behavior.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The skill claims a narrower TTS-only behavior, but the documentation also indicates local file output, extra controls such as volume, and a CLI utility surface including voice listing, while subtitle support appears overstated. This kind of description/behavior mismatch weakens reviewability and can hide resource access patterns such as filesystem writes that are relevant to security and privacy assessment.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The skill claims a narrower TTS-only behavior, but the documentation also indicates local file output, extra controls such as volume, and a CLI utility surface including voice listing, while subtitle support appears overstated. This kind of description/behavior mismatch weakens reviewability and can hide resource access patterns such as filesystem writes that are relevant to security and privacy assessment.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.19.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
96% confidence
Finding

The lockfile pins transitive dependency ws to version 8.19.0, and the supplied advisory data indicates this version is affected by an uninitialized memory disclosure and a memory-exhaustion denial-of-service issue. In a TTS skill, node-edge-tts relies on WebSocket communication to a remote service, so a vulnerable ws library increases risk if the skill processes attacker-influenced responses or is exposed to hostile/intercepted endpoints, potentially causing crashes or unintended data exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The file states that internet access is required and links to an external voice-testing site, but it does not warn that user-provided text and possibly generated audio may be transmitted to third-party services. Because this skill handles arbitrary text for speech generation, missing disclosure can lead users to send sensitive or regulated content to external providers without realizing the privacy implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation instructs users to export a required API key and run installation and tests, but it does not warn that the credential is sensitive, should not be committed to shell history, logs, or source control, and should be stored securely. In a skill ecosystem, users may copy-paste these commands into shared environments or CI systems, increasing the chance of accidental secret exposure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill documents use of environment variables and outbound network access to a third-party TTS API, but it does not declare any tool scope such as permissions or allowed-tools. This creates a governance gap: an agent may invoke networked or environment-backed functionality without explicit policy review, increasing the risk of overbroad execution and accidental secret exposure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger guidance is broad enough that ordinary mentions of "tts" or related phrasing may activate the skill automatically, potentially sending unintended user content to a remote TTS service. In a conversational agent, ambiguous activation increases the risk of privacy-impacting data egress and surprising tool use without clear user intent.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 192)May include surrounding context.

md
- Requires `SKILLBOSS_API_KEY` environment variable
- Output is MP3 format by default
- Requires internet connection
- **Temporary File Handling**: By default, audio files are saved to the system's temporary directory (`/tmp/edge-tts-temp/` on Unix, `C:\Users\<user>\AppData\Local\Temp\edge-tts-temp\` on Windows) with unique filenames (e.g., `tts_1234567890_abc123.mp3`). Files are not automatically deleted - the calling application (Clawdbot) should handle cleanup after use. You can specify a custom output path with the `--output` option if permanent storage is needed.
- **TTS keyword filtering**: The skill automatically filters out TTS-related keywords (tts, TTS, text-to-speech) from text before conversion to avoid converting the trigger words themselves to audio
- For repeated preferences, use `config-manager.js` to set defaults
- **Default voice**: `en-US-MichelleNeural` (female, natural)

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The guide encourages use of an online TTS service and even notes it uses Microsoft Edge's online service, but it does not clearly warn at the integration point that submitted text may be transmitted to a third party. In a TTS skill, users may send sensitive prompts, personal data, or confidential content, so omission of a clear data-sharing warning creates a real privacy and compliance risk.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The script sends user-provided text to an external third-party API service for speech generation, which creates a real data-exposure boundary. In a TTS skill, this is expected behavior, but it is still security-relevant because sensitive prompts, personal data, or secrets could be transmitted off-host without explicit user awareness or policy enforcement.

Content

Scanner excerpt · scripts/tts-converter.js (reported line 18)May include surrounding context.

js
const os = require('os');

const SKILLBOSS_API_KEY = process.env.SKILLBOSS_API_KEY;
const API_BASE = 'https://api.heybossai.com/v1';

// Constants
const DEFAULT_TIMEOUT_MS = 10000;

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest says the skill supports subtitle generation, and the CLI option '--save-subtitles' suggests subtitles will be produced. In practice, the code only adds 'save_subtitles' to the API request and then exclusively reads 'result.result.audio_base64' and writes an audio file, with no handling or persistence of subtitle data.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The option description says 'Save subtitles as JSON file', which communicates a concrete side effect to users. However, no code writes any subtitle JSON file or processes subtitle content at all; only the audio file is saved.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · scripts/package.json (reported line 12)May include surrounding context.

json
"test": "node tts-converter.js \"Hello, world! This is a test of the TTS service.\" -f test-output.mp3 && echo \"Test completed successfully!\""
  },
  "dependencies": {
    "commander": "^11.0.0"
  },
  "author": "Clawdbot",
  "license": "MIT"

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The code sets lang = 'en-US' as the default language, which imposes a specific locale choice when the user does not provide one. Under the policy, forcing a language or locale without explicit opt-in can be a natural-language policy violation unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.env_credential_access

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/tts-converter.js:17