Back to skill

Security audit

Agentvibes Voice Skill

Security checks across malware telemetry and agentic risk

Overview

The skill is a coherent TTS helper, but it under-scopes sensitive audio/history behavior, external text handling, and automatic download/install side effects.

Review this skill before installing if you use agents with private prompts, credentials, or confidential work. Avoid high verbosity, treat replayed audio as sensitive history, and confirm where voices, provider components, translations, and cache files are stored or sent before enabling automatic downloads, installation, or translation features.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Description-Behavior Mismatch

Medium
Confidence
90% confidence
Finding
The skill markets itself as free, offline, and no-account, but elsewhere documents automatic voice downloads from HuggingFace and translation/API-dependent behavior. This mismatch can mislead users and agents into assuming no network or privacy impact, causing unintended outbound requests or data handling in supposedly offline workflows.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The skill describes automatic voice downloads, automatic installation, and cleanup of cached files without prominent warnings about filesystem changes, package installation, or network side effects. In an agent context with exec capability, this can lead to unexpected modification of the host environment or outbound traffic without informed user consent.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The translate-and-play feature implies user text may be sent to translation or model services, yet the skill provides no privacy warning. Users may submit sensitive prompts believing the tool is local, creating a data exposure path to external services or logs.

Ssd 3

Medium
Confidence
95% confidence
Finding
Offering a verbosity mode that speaks the agent's complete reasoning process creates a direct leakage channel for sensitive prompts, internal context, credentials, hidden instructions, or derived secrets. Spoken disclosure is especially risky in shared physical spaces and may bypass protections that only consider text output channels.

Ssd 3

Medium
Confidence
94% confidence
Finding
Persistent mute/replay behavior across sessions and retention of recent audio creates a natural-language history channel that can re-expose previously spoken sensitive content. On shared devices or multi-user environments, another user or process could trigger replay and recover private information without needing access to original text logs.

VirusTotal

64/64 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.