Back to skill

Security audit

qui-edge-tts

Security checks for vulnerabilities and agentic risk

Overview

This TTS skill is mostly purpose-aligned, but it needs Review because it can send user text to an external service, persist generated audio/config files, and has underspecified triggers and dependency installation behavior.

Before installing, use this only for text you are comfortable sending to the SkillBoss/HeyBossAI TTS service, keep the API key out of logs and shared shells, and avoid running the unpinned npx examples. Prefer a locked dependency install, review or remove the stale node-edge-tts lockfile entries, and clean up generated audio files, especially on shared machines.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
install.sh:10
Finding

Non-reproducible and lifecycle-capable dependency installation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/tts-converter.js:43
Finding

Generated speech is stored using unsafe temporary-file handling

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description says the skill performs text-to-speech conversion via a SkillBoss API Hub TTS service. However, this code chunk does not synthesize audio, process input text, contact any external API, or generate subtitles/audio output. Its actual role is limited to managing persistent TTS configuration on disk and exposing a CLI for getting/setting preferences. While these settings align with TTS-related options and could support a larger TTS system, the primary behavior of this code is materially different from the declared purpose. Therefore this chunk is a mismatch with the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The code generally aligns with the declared TTS purpose: it performs text-to-speech via the SkillBoss API and supports voices, languages, pitch, speed/rate, and subtitle-related behavior. However, there are notable undeclared behaviors/capabilities. It saves generated audio to the local filesystem, including temp directories or arbitrary output paths, which is a meaningful resource interaction absent from the declaration. It also supports volume control, which is not listed. Most importantly, it alters user input by removing TTS-related keywords before synthesis; this behavior is not implied by the description and could materially affect output fidelity. The primary purpose is still TTS, but these undeclared behaviors make the description not fully accurate.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.19.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
93% confidence
Finding

The lockfile pins transitive dependency ws to version 8.19.0, and the provided finding indicates this version is affected by memory disclosure and memory-exhaustion denial-of-service advisories. In a TTS skill, text and network responses may traverse WebSocket-based code paths via node-edge-tts, so a vulnerable ws library can expose the service or client process to crashes or unintended data leakage if an attacker can influence or trigger those connections.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill invokes networked TTS functionality and relies on an API key in the environment, but it does not declare an explicit tool scope such as allowed tools or permissions. This weakens least-privilege boundaries and makes it harder for a host agent to constrain network and environment access safely.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The trigger guidance tells the agent to act not only on a precise keyword but also on a broad 'user request' interpretation. Overbroad activation can cause the skill to run unexpectedly on unrelated content, sending user text to an external TTS service and producing files without sufficiently explicit intent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Treating 'tts' as a generic keyword without exclusions is ambiguous and can trigger on incidental text rather than an intentional command. In this skill's context, accidental invocation matters because it can transmit content over the network, create local files, and alter input text before synthesis.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 192)May include surrounding context.

md
- Requires `SKILLBOSS_API_KEY` environment variable
- Output is MP3 format by default
- Requires internet connection
- **Temporary File Handling**: By default, audio files are saved to the system's temporary directory (`/tmp/edge-tts-temp/` on Unix, `C:\Users\<user>\AppData\Local\Temp\edge-tts-temp\` on Windows) with unique filenames (e.g., `tts_1234567890_abc123.mp3`). Files are not automatically deleted - the calling application (Clawdbot) should handle cleanup after use. You can specify a custom output path with the `--output` option if permanent storage is needed.
- **TTS keyword filtering**: The skill automatically filters out TTS-related keywords (tts, TTS, text-to-speech) from text before conversion to avoid converting the trigger words themselves to audio
- For repeated preferences, use `config-manager.js` to set defaults
- **Default voice**: `en-US-MichelleNeural` (female, natural)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guide states that the module uses Microsoft Edge's online TTS service, but it does not clearly warn that input text is transmitted to a third-party remote service. In a TTS skill context, users may submit sensitive prompts, personal data, or confidential content, so omission of a privacy warning can lead to unintended data disclosure.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The documentation instructs users to run npx node-edge-tts without pinning a version, which causes execution of whatever package version is current at install time. If the upstream package is compromised, typo-squatted, or a future release becomes malicious, users may execute attacker-controlled code on their system.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

This example uses npx node-edge-tts without a fixed version, creating a supply-chain risk because the command will fetch and execute the latest available package. In skill documentation, copy-pasteable shell commands are especially risky because users may run them directly with little scrutiny.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The unpinned npx command exposes users to execution of an unreviewed future package version. Because this is presented as a normal usage pattern, it may normalize unsafe installation behavior and increase exposure to package-registry compromise.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

Running npx node-edge-tts without version pinning means the executed code is not reproducible and may change over time. An attacker controlling the package or one of its distribution paths could turn a documentation example into arbitrary code execution on user machines.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

This copy-pasteable npx example is vulnerable to package supply-chain drift because it does not constrain which code will be downloaded and run. The skill context increases practical risk since agents or users may automate TTS generation and execute examples verbatim.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The subtitle example again recommends npx node-edge-tts without a version, preserving the same supply-chain execution risk. Repetition across the document increases the chance that a user copies at least one unsafe command.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill hard-codes both a U.S. English neural voice and the 'en-US' locale as defaults. Under the policy, forcing a specific language or locale without user opt-in can be a natural-language policy violation unless the constraint is clearly justified or alternatives are offered.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This skill is explicitly designed to transmit text to an external TTS API, so external transmission is expected in context; however, it is still a real privacy/security-relevant behavior because any supplied text leaves the local environment. The danger is elevated by the skill purpose: users may submit arbitrary content for speech generation, including secrets or sensitive material, and the code provides no consent gating, minimization, or response validation.

Content

Scanner excerpt · scripts/tts-converter.js (reported line 18)May include surrounding context.

js
const os = require('os');

const SKILLBOSS_API_KEY = process.env.SKILLBOSS_API_KEY;
const API_BASE = 'https://api.heybossai.com/v1';

// Constants
const DEFAULT_TIMEOUT_MS = 10000;

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest explicitly advertises subtitle generation as a supported capability. In the implementation, enabling --save-subtitles only adds save_subtitles to the API request, while the code then reads only result.result.audio_base64 and writes only an audio file, with no handling of subtitle content or subtitle file creation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script sends arbitrary user-supplied text to a third-party TTS service without any explicit notice, consent flow, or privacy warning. In a skill context, users may provide sensitive data assuming local processing, so silent transmission can expose personal, confidential, or regulated information to an external provider.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file instructs users to set the SKILLBOSS_API_KEY environment variable and notes that an internet connection is required, but it does not warn users that the skill uses credentials and transmits data to an external TTS service. For markdown files, omission of warnings about behaviors affecting privacy or sensitive data handling is in scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file includes repeated examples that create output.mp3 and output.json files, but it does not explicitly warn users that running the commands will write files into the current workspace. For markdown files, user-facing documentation should disclose behaviors that affect user data or system state, even when the writes are expected.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · scripts/package.json (reported line 12)May include surrounding context.

json
"test": "node tts-converter.js \"Hello, world! This is a test of the TTS service.\" -f test-output.mp3 && echo \"Test completed successfully!\""
  },
  "dependencies": {
    "commander": "^11.0.0"
  },
  "author": "Clawdbot",
  "license": "MIT"

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The docstring promises successful generation semantics, yet the function does not validate that result.result.audio_base64 exists before decoding and writing. If the upstream response shape differs or subtitle-only metadata is returned, the behavior diverges from the documented intent rather than reliably producing the documented artifact.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The function defaults lang to en-US, and the CLI also defaults the language option to en-US, which imposes a specific locale unless the user overrides it. Under the stated policy, forcing a language or locale without user opt-in can be a natural-language policy violation when not clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The command-line option definition sets 'en-US' as the default language, so the tool operates in a specific locale even if the user does not actively choose one. This can conflict with the policy against forcing a language or locale without opt-in unless the constraint is documented and justified.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.env_credential_access

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/tts-converter.js:17