Back to skill

Security audit

speech-to-text-api

Security checks across malware telemetry and agentic risk

Overview

This is an instruction-only skill, but it presents itself as speech-to-text while steering users toward a broad third-party API gateway and non-speech capabilities.

Review before installing. Use this only if you trust SkillBoss with the audio or other data sent through its API, and if you are comfortable that the required API key may enable many paid non-speech APIs. Prefer a narrower STT-only integration if you only need transcription.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
The skill is marketed as narrowly for speech-to-text, but the setup text expands its scope to hundreds of unrelated APIs and capabilities. This creates deceptive scope and can cause an agent or user to grant credentials and install a much broader integration surface than expected, increasing the chance of unintended tool use and data exposure.

Description-Behavior Mismatch

Medium
Confidence
96% confidence
Finding
Later sections advertise chat, image, video, scraping, and social-data capabilities that do not match the declared STT purpose. This mismatch undermines least-privilege expectations and may induce users or agents to enable a generalized API broker under the guise of a single-purpose speech tool.

Intent-Code Divergence

Low
Confidence
97% confidence
Finding
The examples claim to demonstrate speech-to-text but send image-style text prompts to a Whisper model instead of audio input. Misleading examples can cause agents to build incorrect integrations, mishandle user data, or silently fail in production while users believe audio transcription is occurring.

Intent-Code Divergence

Medium
Confidence
95% confidence
Finding
The agent instructions recommend unrelated chat/reasoning models as cheaper or higher-quality alternatives in a speech-to-text skill. This can steer agents into using the wrong model class entirely, causing inappropriate data routing and tool substitution outside the user's requested task.

Vague Triggers

Medium
Confidence
84% confidence
Finding
The activation phrase 'USE THIS when the user needs speech to text api' is broad and lacks limiting conditions, making accidental invocation more likely. In an agent context, overly broad routing language can trigger external API use for loosely related requests without sufficient user intent or review.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The 'When To Use This Skill' section is ambiguous and contains only positive triggers without boundaries or negative examples. This increases the risk that an agent routes general audio, AI, or model-selection tasks into this third-party service unnecessarily, expanding data sharing beyond user expectations.

VirusTotal

65/65 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.