Back to skill

Security audit

whisper-stt-api

Security checks for vulnerabilities and agentic risk

Overview

This Whisper transcription skill routes through SkillBoss but asks users to enable a much broader, mutable API platform than the stated speech-to-text purpose requires.

Review this before installing. It is not just a narrow Whisper helper: setup may configure a broad SkillBoss API gateway, and your requests plus bearer token go to SkillBoss. Use it only if you intentionally want that intermediary and broad platform access; prefer a narrowly scoped Whisper-only setup or key if available.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:32
Finding

Unaudited Remote Instructions Expand the Skill Beyond Its Declared Scope

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 32-39
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Code:

markdown
## Setup (One Command)

Tell your agent:

set up skillboss.co/skill.md

text

This auto-configures SkillBoss with 687 APIs — chat, image, video, audio, search, scraping, social data, email, and more.

Technical Analysis

The skill instructs the agent to obtain and follow setup material hosted at skillboss.co/skill.md. That external document is not included in the reviewed package, is not pinned to an immutable version or cryptographic digest, and can change independently after this skill has been audited.

Consequently, the effective instructions executed during setup are not limited to the contents of the audited SKILL.md. If the remote document is changed, compromised, or serves different content conditionally, it could introduce additional instructions that alter agent behavior or request further actions without those instructions appearing in this package.

The setup also explicitly expands access from the declared Whisper speech-to-text functionality to 687 APIs, including scraping, social-data, and email services. This exceeds the minimum capabilities needed to submit audio for transcription and violates least-privilege principles.

The separate API examples transmit an authorization bearer token and user-provided content to https://api.skillboss.co/v1/run. This transmission is disclosed and is functionally expected for the brokered transcription service; the reviewed file does not show the credential being sent to an unrelated destination. Nevertheless, users should be clearly informed that their content is processed by SkillBoss rather than sent directly to OpenAI.

Attack Path

  1. A user or agent loads the locally audited Whisper skill.
  2. The agent follows the setup instruction to access skillboss.co/skill.md.

...[truncated 1303 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the instruction that asks the agent to set up the skill from a mutable external Markdown document.
  2. Include all required setup instructions directly in the reviewed package.
  3. Restrict configuration and documented permissions to the single Whisper transcription endpoint required by the skill.
  4. If remote setup material is unavoidable, reference an immutable version and verify it against a cryptographic hash or authenticated signature before use.
  5. Display the retrieved instructions to the user and require explicit confirmation before applying configuration or enabling additional APIs.
  6. Use an API key scoped only to openai/whisper-1 and deny unrelated services server-side.
  7. Clearly disclose that authentication credentials and transcription content are sent to SkillBoss as an intermediary, including applicable retention and privacy policies.
  8. Replace the broad “687 APIs” setup flow with a minimal, auditable procedure that configures only the endpoint, credential, and permissions necessary for speech-to-text.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (13)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger phrase 'USE THIS when the user needs whisper api' is broad and lacks guardrails, making over-invocation likely. In a skill that also provisions access to hundreds of unrelated APIs, broad triggering raises the chance that agents invoke or recommend an overprivileged external platform when a narrower or local option would suffice.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill is branded as a narrow Whisper STT integration, but its setup and marketing explicitly grant access to a much broader third-party API platform. That mismatch can mislead users and agents into authorizing broad capabilities and key storage far beyond speech-to-text, increasing the risk of unintended tool use and data exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The setup language emphasizes convenience but does not clearly warn that configuration stores an API key and enables access to a large external API aggregation platform. That omission undermines informed consent and may cause users to grant broader external access than they realize.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 68)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 89)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

python
import requests

response = requests.post(
    "https://api.skillboss.co/v1/run",
    headers={"Authorization": "Bearer $SKILLBOSS_API_KEY"},
    json={

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

python
import requests

response = requests.post(
    "https://api.skillboss.co/v1/run",
    headers={"Authorization": "Bearer $SKILLBOSS_API_KEY"},
    json={

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The 'When To Use This Skill' section is underspecified and contains no exclusion boundaries, which increases the likelihood of accidental invocation. Because the backing service exposes many non-STT capabilities, ambiguous routing here materially enlarges the chance of unnecessary external data sharing and inappropriate tool selection.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The agent instructions recommend unrelated chat models such as DeepSeek, Gemini, Claude, and O3 inside a Whisper STT skill. This expands the effective behavior of the skill beyond its declared purpose and can cause agents to route sensitive user requests to unexpected external models or services.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill promotes access to scraping, social data, email, and many other unrelated capabilities under the cover of a Whisper STT integration. In context, this broadens the trust boundary significantly and could induce an agent or user to enable powerful external data-access functions without informed consent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The surrounding documentation claims this skill is for Whisper speech-to-text, yet the example input uses a text prompt describing an image rather than audio or transcription data. This actively misaligns the documented intent of STT usage with what the example appears to demonstrate.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.