Back to skill

Security audit

elevenlabs-tts-api

Security checks for vulnerabilities and agentic risk

Overview

This skill is presented as ElevenLabs text-to-speech, but its setup and instructions expand to a broad third-party API gateway beyond that purpose.

Review before installing. Use only if you intend to use SkillBoss as a broker and are comfortable giving an agent a SkillBoss key that may cover many API categories. Avoid the remote one-command setup unless you can inspect and approve what it changes, and do not send sensitive text through the API unless SkillBoss's data handling is acceptable to you.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:24
Finding

Unpinned External Setup Instructions Allow Remote Skill Instruction Hijacking

Content
View full analysis
**One API key. ElevenLabs TTS API. Zero markup.** USE THIS when the user needs elevenlabs api. SkillBoss provides ElevenLabs text-to-speech via OpenAI-compatible API with 0% markup. ## Setup (One Command) Tell your agent: ``` set up skillboss.co/skill.md ``` This auto-configures SkillBoss with 687 APIs — chat, image, video, audio, search, scraping, social data, email, and more. ``` ### Technical Analysis The Skill tells the agent to set itself up using `skillboss.co/skill.md`, an externally hosted and mutable instruction document that is not included in the audited package. Consequently, the effective setup behavior cannot be determined from the reviewed `SKILL.md` file and may change after publication or audit. When an agent interprets the remote document as instructions, the operator controlling that document can potentially introduce new directives that alter setup actions, tool use, credential handling, or the agent's current objectives. This is an instruction-channel trust-boundary violation: unaudited remote content is treated as authoritative Skill configuration rather than untrusted external data. The requested setup also exceeds the minimum capabilities required for the declared ElevenLabs text-to-speech function. The Skill states that this command configures 687 APIs, including scraping, social-data, email, search, image, video, and chat services. A TTS integration only requires access to the relevant audio endpoint and credential. Enabling unrelated services unnecessarily broadens the available capability and compromise scope. This finding ...[truncated 1927 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (13)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Advertising scraping, social data, search, email, and many other capabilities inside a TTS-specific skill materially broadens the operational surface area. In an agent environment, this can cause capability escalation, where installation of a seemingly harmless audio skill implicitly introduces powerful unrelated actions and external data flows.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Discover More' section promotes broad use of unrelated chat, image, video, social-data, and scraping APIs in the context of a TTS skill. This encourages post-install capability expansion beyond the user’s original intent and increases the likelihood of inappropriate data access or tool selection by the agent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The broad instruction to 'USE THIS' when the user needs 'elevenlabs api' lacks scope limits and can trigger the skill for loosely related requests. In an agent setting, ambiguous invocation rules increase the chance of accidental activation and unintended external calls to a third-party gateway.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill is advertised as a narrowly scoped ElevenLabs TTS integration, but it immediately expands into a gateway for hundreds of unrelated APIs and model categories. That scope mismatch can mislead an agent into granting or using much broader capabilities than the user intended, increasing the chance of unauthorized tool use and data exposure through unrelated services.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This finding identifies the same curl example endpoint itself as an external transmission sink. The risk is amplified by the skill’s scope confusion: a user installing a TTS skill may not realize they are enabling communication with a broad API aggregation platform rather than a single-purpose provider.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This finding identifies the same curl example endpoint itself as an external transmission sink. The risk is amplified by the skill’s scope confusion: a user installing a TTS skill may not realize they are enabling communication with a broad API aggregation platform rather than a single-purpose provider.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

Quick Test

bash
curl -X POST "https://api.skillboss.co/v1/run" \
  -H "Authorization: Bearer $SKILLBOSS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This finding also reflects external transmission through the documented Python request, including bearer-token authenticated traffic to a non-local service. In the context of a narrowly branded TTS skill, the intermediary routing makes the transmission more sensitive because users may assume a direct provider integration rather than a multi-API broker.

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

python
import requests

response = requests.post(
    "https://api.skillboss.co/v1/run",
    headers={"Authorization": "Bearer $SKILLBOSS_API_KEY"},
    json={

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This finding also reflects external transmission through the documented Python request, including bearer-token authenticated traffic to a non-local service. In the context of a narrowly branded TTS skill, the intermediary routing makes the transmission more sensitive because users may assume a direct provider integration rather than a multi-API broker.

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

python
import requests

response = requests.post(
    "https://api.skillboss.co/v1/run",
    headers={"Authorization": "Bearer $SKILLBOSS_API_KEY"},
    json={

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The Python example’s endpoint field again confirms outbound transmission to SkillBoss. This is not inherently malicious, but it is a true exposure point because any text submitted for synthesis and associated authentication metadata leave the local environment and may be processed by an intermediary.

Content

Scanner excerpt · SKILL.md (reported line 68)May include surrounding context.

md
import requests

response = requests.post(
    "https://api.skillboss.co/v1/run",
    headers={"Authorization": "Bearer $SKILLBOSS_API_KEY"},
    json={
        "model": "elevenlabs/eleven_multilingual_v2",

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The 'When To Use This Skill' section is ambiguous and lacks negative examples, so an agent may over-apply the skill whenever a user mentions ElevenLabs or APIs generally. That ambiguity is more dangerous here because the backend is a broad multi-service gateway rather than a narrowly scoped TTS endpoint.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The API reference documents a remote endpoint that requires a bearer token and sends requests off-platform. In a skill context, that is a real security-relevant behavior because agents may forward user content automatically; the main risk comes from insufficient disclosure and the mismatch between the claimed provider-specific function and the broader brokered service.

Content

Scanner excerpt · SKILL.md (reported line 89)May include surrounding context.

Endpoint

text
POST https://api.skillboss.co/v1/run
Authorization: Bearer $SKILLBOSS_API_KEY

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The agent instructions recommend unrelated chat models and providers even though the skill is presented as an ElevenLabs TTS skill. This creates a scope-confusion risk where an agent may route requests to non-TTS services unexpectedly, violating least privilege and user expectations.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The documentation says this is an ElevenLabs text-to-speech skill, yet the example input prompt is 'A beautiful sunset over mountains, cinematic lighting', which is characteristic of image-generation prompts rather than TTS text. This actively misaligns the example's intent with the documented TTS purpose and may mislead agents about what the call actually does.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.