T08 · Insecure Dependencies
- Location
install.sh:10- Finding
Non-reproducible and lifecycle-capable dependency installation
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This TTS skill is mostly purpose-aligned, but it needs Review because it can send user text to an external service, persist generated audio/config files, and has underspecified triggers and dependency installation behavior.
Before installing, use this only for text you are comfortable sending to the SkillBoss/HeyBossAI TTS service, keep the API key out of logs and shared shells, and avoid running the unpinned npx examples. Prefer a locked dependency install, review or remove the stale node-edge-tts lockfile entries, and clean up generated audio files, especially on shared machines.
install.sh:10Non-reproducible and lifecycle-capable dependency installation
scripts/tts-converter.js:43Generated speech is stored using unsafe temporary-file handling
The declared description says the skill performs text-to-speech conversion via a SkillBoss API Hub TTS service. However, this code chunk does not synthesize audio, process input text, contact any external API, or generate subtitles/audio output. Its actual role is limited to managing persistent TTS configuration on disk and exposing a CLI for getting/setting preferences. While these settings align with TTS-related options and could support a larger TTS system, the primary behavior of this code is materially different from the declared purpose. Therefore this chunk is a mismatch with the declared description.
The code generally aligns with the declared TTS purpose: it performs text-to-speech via the SkillBoss API and supports voices, languages, pitch, speed/rate, and subtitle-related behavior. However, there are notable undeclared behaviors/capabilities. It saves generated audio to the local filesystem, including temp directories or arbitrary output paths, which is a meaningful resource interaction absent from the declaration. It also supports volume control, which is not listed. Most importantly, it alters user input by removing TTS-related keywords before synthesis; this behavior is not implied by the description and could materially affect output fidelity. The primary purpose is still TTS, but these undeclared behaviors make the description not fully accurate.
The lockfile pins transitive dependency ws to version 8.19.0, and the provided finding indicates this version is affected by memory disclosure and memory-exhaustion denial-of-service advisories. In a TTS skill, text and network responses may traverse WebSocket-based code paths via node-edge-tts, so a vulnerable ws library can expose the service or client process to crashes or unintended data leakage if an attacker can influence or trigger those connections.
The skill invokes networked TTS functionality and relies on an API key in the environment, but it does not declare an explicit tool scope such as allowed tools or permissions. This weakens least-privilege boundaries and makes it harder for a host agent to constrain network and environment access safely.
The trigger guidance tells the agent to act not only on a precise keyword but also on a broad 'user request' interpretation. Overbroad activation can cause the skill to run unexpectedly on unrelated content, sending user text to an external TTS service and producing files without sufficiently explicit intent.
Treating 'tts' as a generic keyword without exclusions is ambiguous and can trigger on incidental text rather than an intentional command. In this skill's context, accidental invocation matters because it can transmit content over the network, create local files, and alter input text before synthesis.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
- Requires `SKILLBOSS_API_KEY` environment variable
- Output is MP3 format by default
- Requires internet connection
- **Temporary File Handling**: By default, audio files are saved to the system's temporary directory (`/tmp/edge-tts-temp/` on Unix, `C:\Users\<user>\AppData\Local\Temp\edge-tts-temp\` on Windows) with unique filenames (e.g., `tts_1234567890_abc123.mp3`). Files are not automatically deleted - the calling application (Clawdbot) should handle cleanup after use. You can specify a custom output path with the `--output` option if permanent storage is needed.
- **TTS keyword filtering**: The skill automatically filters out TTS-related keywords (tts, TTS, text-to-speech) from text before conversion to avoid converting the trigger words themselves to audio
- For repeated preferences, use `config-manager.js` to set defaults
- **Default voice**: `en-US-MichelleNeural` (female, natural)
The guide states that the module uses Microsoft Edge's online TTS service, but it does not clearly warn that input text is transmitted to a third-party remote service. In a TTS skill context, users may submit sensitive prompts, personal data, or confidential content, so omission of a privacy warning can lead to unintended data disclosure.
The documentation instructs users to run npx node-edge-tts without pinning a version, which causes execution of whatever package version is current at install time. If the upstream package is compromised, typo-squatted, or a future release becomes malicious, users may execute attacker-controlled code on their system.
This example uses npx node-edge-tts without a fixed version, creating a supply-chain risk because the command will fetch and execute the latest available package. In skill documentation, copy-pasteable shell commands are especially risky because users may run them directly with little scrutiny.
The unpinned npx command exposes users to execution of an unreviewed future package version. Because this is presented as a normal usage pattern, it may normalize unsafe installation behavior and increase exposure to package-registry compromise.
Running npx node-edge-tts without version pinning means the executed code is not reproducible and may change over time. An attacker controlling the package or one of its distribution paths could turn a documentation example into arbitrary code execution on user machines.
This copy-pasteable npx example is vulnerable to package supply-chain drift because it does not constrain which code will be downloaded and run. The skill context increases practical risk since agents or users may automate TTS generation and execute examples verbatim.
The subtitle example again recommends npx node-edge-tts without a version, preserving the same supply-chain execution risk. Repetition across the document increases the chance that a user copies at least one unsafe command.
The skill hard-codes both a U.S. English neural voice and the 'en-US' locale as defaults. Under the policy, forcing a specific language or locale without user opt-in can be a natural-language policy violation unless the constraint is clearly justified or alternatives are offered.
This skill is explicitly designed to transmit text to an external TTS API, so external transmission is expected in context; however, it is still a real privacy/security-relevant behavior because any supplied text leaves the local environment. The danger is elevated by the skill purpose: users may submit arbitrary content for speech generation, including secrets or sensitive material, and the code provides no consent gating, minimization, or response validation.
const os = require('os');
const SKILLBOSS_API_KEY = process.env.SKILLBOSS_API_KEY;
const API_BASE = 'https://api.heybossai.com/v1';
// Constants
const DEFAULT_TIMEOUT_MS = 10000;
The manifest explicitly advertises subtitle generation as a supported capability. In the implementation, enabling --save-subtitles only adds save_subtitles to the API request, while the code then reads only result.result.audio_base64 and writes only an audio file, with no handling of subtitle content or subtitle file creation.
The script sends arbitrary user-supplied text to a third-party TTS service without any explicit notice, consent flow, or privacy warning. In a skill context, users may provide sensitive data assuming local processing, so silent transmission can expose personal, confidential, or regulated information to an external provider.
This markdown file instructs users to set the SKILLBOSS_API_KEY environment variable and notes that an internet connection is required, but it does not warn users that the skill uses credentials and transmits data to an external TTS service. For markdown files, omission of warnings about behaviors affecting privacy or sensitive data handling is in scope.
This markdown file includes repeated examples that create output.mp3 and output.json files, but it does not explicitly warn users that running the commands will write files into the current workspace. For markdown files, user-facing documentation should disclose behaviors that affect user data or system state, even when the writes are expected.
Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.
"test": "node tts-converter.js \"Hello, world! This is a test of the TTS service.\" -f test-output.mp3 && echo \"Test completed successfully!\""
},
"dependencies": {
"commander": "^11.0.0"
},
"author": "Clawdbot",
"license": "MIT"
The docstring promises successful generation semantics, yet the function does not validate that result.result.audio_base64 exists before decoding and writing. If the upstream response shape differs or subtitle-only metadata is returned, the behavior diverges from the documented intent rather than reliably producing the documented artifact.
The function defaults lang to en-US, and the CLI also defaults the language option to en-US, which imposes a specific locale unless the user overrides it. Under the stated policy, forcing a language or locale without user opt-in can be a natural-language policy violation when not clearly justified as region-specific.
The command-line option definition sets 'en-US' as the default language, so the tool operates in a specific locale even if the user does not actively choose one. This can conflict with the policy against forcing a language or locale without opt-in unless the constraint is documented and justified.
Detected: suspicious.env_credential_access